Why there is no “normalize everything” button
Normalization is not universal spelling correction. Converting Alef Maqsura to Ya can improve database matching while changing the written form. Simplifying Hamza carriers can increase recall in approximate search but is unsuitable for publication. The studio therefore exposes every decision and starts with conservative defaults: remove optional marks and tatweel, normalize Alef, convert digits and clean repeated spacing.
Verifiable privacy
Every transformation runs locally in JavaScript. No text upload or account is required.
Independent controls
Each rule can be enabled or disabled instead of applying an opaque, irreversible cleanup.
Open test corpus
We publish inputs, expected outputs and limitations so the behavior can be checked and reproduced.
A safe workflow for search and comparison
- 1Keep the original text as an unchanged authoritative record.
- 2Define the purpose: search, deduplication, comparison or indexing.
- 3Enable the smallest set of transformations needed.
- 4Test names, digits, Hamza forms and marked text from the real corpus.
- 5Store the normalized copy separately and document the enabled rules.
Methodology and downloadable tests
Review the definition of every rule and the regression corpus used to catch unexpected changes. Developers can download the JSON cases and reuse them in their own tests.
Frequently asked questions
Does the studio modify my original text?
No. Input stays in the browser and the studio creates a separate result that you can copy or download.
Is removing Arabic diacritics always appropriate?
No. Marks can be essential for pronunciation, study, quotation and religious text. Always retain the authoritative original.
Is Alef, Ya or Hamza normalization spelling correction?
No. It is an optional technical transformation for matching or comparison and can remove meaningful linguistic distinctions.