General Information

Abstract

This document establishes a system for the transliteration of Perso-Arabic characters into Latin characters. This modification of the stringent rules established by ISO 233 [1] is specifically intended to facilitate the processing of bibliographic information (e.g. catalogues, indices, citations, etc.).

Status
Published
Publication Date
26-Aug-2026
Current Stage
6060 - International Standard published
Start Date
27-Aug-2026
Due Date
20-Apr-2027
Completion Date
27-Aug-2026

Buy Documents

Standard

ISO 233-3:2026 - Information and documentation — Transliteration of Perso-Arabic characters into Latin characters — Part 3: Persian language

Release Date:27-Aug-2026
English language (14 pages)
sale 15% off
Preview
sale 15% off
Preview

Overview

ISO 233-3:2026, titled Information and documentation - Transliteration of Perso-Arabic characters into Latin characters - Part 3: Persian language, is an international standard developed by ISO for the accurate and standardized transliteration of Persian (Farsi) text from Perso-Arabic script into Latin characters. This part specifically addresses modifications to the strict transliteration rules of ISO 233, focusing on optimizing the process for bibliographic data management, including catalogues, citations, and indices.

The Persian language, written with the Perso-Arabic script, poses unique challenges for information processing and international communication due to script differences and unmarked diacritics. ISO 233-3:2026 provides a universal, reversible system to ensure consistency in information exchange, retrieval, and records management across borders and languages.

Key Topics

  • Strict Transliteration System: Defines a one-to-one, fully reversible mapping between Perso-Arabic and Latin characters, including treatment for vowels, consonants, Hamze (ء), and special Persian suffixes like ez̤āfe.
  • Modified Transliteration System: Offers a flexible, context-sensitive approach prioritizing readability and pronunciation, especially useful in educational and general publishing where reversibility is not essential.
  • Handling of Diacritical Marks: Details the treatment of vowels and diacritics, which may not always be written in Persian manuscripts, ensuring accurate representation and disambiguation in the Latin script.
  • Processing Arabic Elements in Persian: Sets rules for handling Arabic loanwords and specific characters, ensuring consistency with related ISO standards.
  • Punctuation and Numerals: Specifies the conversion of Persian punctuation marks and numerals into their Latin equivalents for seamless information exchange.
  • General Principles: Emphasizes univocity, reversibility, and adherence to international best practices for script conversion and romanization.

Applications

ISO 233-3:2026 is particularly valuable for:

  • Bibliographic Management: Libraries, archives, and information centers can apply this standard for cataloguing Persian language materials, enhancing searchability and interoperability.
  • Academic Research and Publishing: Scholars can ensure accuracy in citations, references, and indices when working with Persian sources.
  • Linguistic Data Processing: Software developers and database managers can use the standard to automate text transformation, information retrieval, and indexing in multilingual environments.
  • Cross-Border Information Exchange: International organizations and governmental bodies can rely on standardized transliteration to enable seamless communication and data sharing involving Persian-language texts.
  • Educational Resources: The modified system supports educational tools by providing more readable and pronounceable Latin forms for learners of Persian.

Related Standards

ISO 233-3:2026 is part of the broader ISO 233 series, which covers transliteration from Arabic-derived scripts. Key related standards include:

  • ISO 233:1984: The foundational standard for transliteration of Arabic characters into Latin characters, providing the strictest rules.
  • ISO 233-1: Covers general principles and systems for the transliteration of Arabic.
  • ISO/IEC 10646: Specifies the Universal Coded Character Set (UCS), crucial for character encoding and interoperability.
  • ISIRI 6219: The Iranian national standard addressing keyboarding and representation of Persian script.

Adoption of ISO 233-3:2026 ensures compatibility and interoperability with other national and international script conversion systems, contributing to more unified global information management for Persian and other Perso-Arabic scripts.


Keywords: ISO 233-3:2026, Persian transliteration, Perso-Arabic to Latin, bibliographic information, script conversion, romanization, international standard, information and documentation, library cataloguing, Unicode, ISO standards, diacritical marks, Persian data processing.

Relations

Effective Date
20-Oct-2025

Buy Documents

Standard

ISO 233-3:2026 - Information and documentation — Transliteration of Perso-Arabic characters into Latin characters — Part 3: Persian language

Release Date:27-Aug-2026
English language (14 pages)
sale 15% off
Preview
sale 15% off
Preview

Frequently Asked Questions

ISO 233-3:2026 is a standard published by the International Organization for Standardization (ISO). Its full title is "Information and documentation — Transliteration of Perso-Arabic characters into Latin characters — Part 3: Persian language". This standard covers: This document establishes a system for the transliteration of Perso-Arabic characters into Latin characters. This modification of the stringent rules established by ISO 233 [1] is specifically intended to facilitate the processing of bibliographic information (e.g. catalogues, indices, citations, etc.).

This document establishes a system for the transliteration of Perso-Arabic characters into Latin characters. This modification of the stringent rules established by ISO 233 [1] is specifically intended to facilitate the processing of bibliographic information (e.g. catalogues, indices, citations, etc.).

ISO 233-3:2026 is classified under the following ICS (International Classification for Standards) categories: 01.140.10 - Writing and transliteration. The ICS classification helps identify the subject area and facilitates finding related standards.

ISO 233-3:2026 has the following relationships with other standards: It is inter standard links to ISO 233-3:2023. Understanding these relationships helps ensure you are using the most current and applicable version of the standard.

ISO 233-3:2026 is available in PDF format for immediate download after purchase. The document can be added to your cart and obtained through the secure checkout process. Digital delivery ensures instant access to the complete standard document.

Standards Content (Sample)


International
Standard
ISO 233-3
Third edition
Information and documentation —
2026-08
Transliteration of Perso-Arabic
characters into Latin characters —
Part 3:
Persian language
Information et documentation — Translittération des caractères
perso-arabes en caractères latins —
Partie 3: Langue persane
Reference number
© ISO 2026
All rights reserved. Unless otherwise specified, or required in the context of its implementation, no part of this publication may
be reproduced or utilized otherwise in any form or by any means, electronic or mechanical, including photocopying, or posting on
the internet or an intranet, without prior written permission. Permission can be requested from either ISO at the address below
or ISO’s member body in the country of the requester.
ISO copyright office
CP 401 • Ch. de Blandonnet 8
CH-1214 Vernier, Geneva
Phone: +41 22 749 01 11
Email: copyright@iso.org
Website: www.iso.org
Published in Switzerland
ii
Contents Page
Foreword .iv
Introduction .v
1 Scope . 1
2 Normative references . 1
3 Terms and definitions . 1
4 Strict transliteration . 2
4.1 General .2
4.2 Consonants .2
4.3 Vowels .4
4.4 Arabic elements in the Persian language .4
4.5 Hamze .5
4.6 Persian relational suffix (ez̤̤āfe) .5
4.7 Punctuation marks .5
4.8 Persian numerals .6
5 Modified transliteration . 6
5.1 General .6
5.2 Vowels and consonants .6
5.3 Arabic elements in the Persian language .8
5.4 Hamze .8
5.5 Persian relational suffix (ez̤̤āfe) .8
6 General principles of transliteration . 9
Annex A (informative) Different positional forms of characters .10
Annex B (normative) General principles .11
Bibliography .13

iii
Foreword
ISO (the International Organization for Standardization) is a worldwide federation of national standards
bodies (ISO member bodies). The work of preparing International Standards is normally carried out through
ISO technical committees. Each member body interested in a subject for which a technical committee
has been established has the right to be represented on that committee. International organizations,
governmental and non-governmental, in liaison with ISO, also take part in the work. ISO collaborates closely
with the International Electrotechnical Commission (IEC) on all matters of electrotechnical standardization.
The procedures used to develop this document and those intended for its further maintenance are described
in the ISO/IEC Directives, Part 1. In particular, the different approval criteria needed for the different types
of ISO document should be noted. This document was drafted in accordance with the editorial rules of the
ISO/IEC Directives, Part 2 (see www.iso.org/directives).
ISO draws attention to the possibility that the implementation of this document may involve the use of (a)
patent(s). ISO takes no position concerning the evidence, validity or applicability of any claimed patent
rights in respect thereof. As of the date of publication of this document, ISO had not received notice of (a)
patent(s) which may be required to implement this document. However, implementers are cautioned that
this may not represent the latest information, which may be obtained from the patent database available at
www.iso.org/patents. ISO shall not be held responsible for identifying any or all such patent rights.
Any trade name used in this document is information given for the convenience of users and does not
constitute an endorsement.
For an explanation of the voluntary nature of standards, the meaning of ISO specific terms and expressions
related to conformity assessment, as well as information about ISO's adherence to the World Trade
Organization (WTO) principles in the Technical Barriers to Trade (TBT), see www.iso.org/iso/foreword.html.
This document was prepared by Technical Committee ISO/TC 46, Information and documentation.
This third edition cancels and replaces the second edition (ISO 233-3:2023), of which it constitutes a minor
revision.
The changes are as follows:
— the title has been revised in line with the draft being developed for ISO 233-1;
— the Scope has been modified to align with the terminology used in the title.
A list of all parts in the ISO 233 series can be found on the ISO website.
Any feedback or questions on this document should be directed to the user’s national standards body. A
complete listing of these bodies can be found at www.iso.org/members.html.

iv
Introduction
This document is one of a series of International Standards, dealing with the conversion of systems of writing.
The aim of the ISO 233 series is to provide a means for international communication of written messages in
a form which permits the automatic transmission and reconstitution of these, by humans or machines. The
system of conversion, in this case, must be univocal and entirely reversible to allow for retransliteration.
This means that consideration to phonetic and aesthetic matters or to certain national customs is not a
priority: all these considerations are, indeed, ignored by the machine performing the function.
This document can be used by anyone who has a clear understanding of the system and is certain that it can
be applied without ambiguity. The result obtained will not give a correct pronunciation of the original text
in a person’s own language, but it will serve as a means of finding automatically the original graphism, and
thus allow anyone who has knowledge of the original language to pronounce it correctly. Similarly, one can
only correctly pronounce a text written in, for example, English or Polish, if one has a knowledge of English
or Polish.
The existence in Perso-Arabic script of vowel signs and other diacritical marks, which are pronounced but
often not written somewhat complicates reading of the text, but as those with knowledge of the language
can read and mentally fill in the missing signs/sounds when reading the original script, so can they with
the transliterated version, for example, the word رَپَِسِ, which consists of three consonants and two diacritical
vowel signs (transliteration: separ) when written without vowel signs would be رپَسِ (transliteration: spr).
To address the issue of diacritical vowels and other signs that are unwritten, and the fact that some characters
perform more than one function (e.g. characters that can function as either a vowel or a consonant), this
document incorporates three levels of transliteration:
1) strict and fully reversible, univocal, with diacritical vowels and other signs only transliterated if written
in the source text;
2) strict and fully reversible, univocal, with diacritical vowels and other signs included for clarity,
regardless of their presence or absence in the source text;
3) a modified version of the system that while not fully reversible includes the diacritical vowels and other
signs and takes account of the different functions performed by some characters (see 4.1 and 5.1 for
further details).
The adoption of this document for international communication leaves every country free to adopt for its
own use a national standard which can be different, on condition that it is compatible with this document.
The system proposed herein will make this possible and be acceptable to international use if the graphisms
it creates are such that they can be converted automatically into the graphisms used in any strict national
systems.
The adoption of national standards compatible with this document permits the representation, in an
international publication, of the morphemes of each language according to the customs of the country where
it is spoken. It is possible to simplify this representation in order to take into account the number of the
character sets available on different kinds of machines.

v
International Standard ISO 233-3:2026(en)
Information and documentation — Transliteration of Perso-
Arabic characters into Latin characters —
Part 3:
Persian language
1 Scope
This document establishes a system for the transliteration of Perso-Arabic characters into Latin characters.
[1]
This modification of the stringent rules established by ISO 233 is specifically intended to facilitate the
processing of bibliographic information (e.g. catalogues, indices, citations, etc.).
2 Normative references
The following documents are referred to in the text in such a way that some or all of their content constitutes
requirements of this document. For dated references, only the edition cited applies. For undated references,
the latest edition of the referenced document (including any amendments) applies.
ISO/IEC 10646, Information technology — Universal coded character set (UCS)
3 Terms and definitions
For the purposes of this document, the following terms and definitions apply.
ISO and IEC maintain terminological databases for use in standardization at the following addresses:
— ISO Online browsing platform: available at https:// www .iso .org/ obp
— IEC Electropedia: available at https:// www .electropedia .org/
3.1
character
element of an alphabetical or other type of writing system that graphically represents a phoneme, a syllable,
a word or even a prosodical characteristic of a given language
Note 1 to entry: It is used either alone (for example, a letter, a syllabic sign, an ideographical character, a digit, a
punctuation mark) or in combination (such as an accent or a diacritical mark).
Note 2 to entry: A letter having an accent or a diacritical mark, for example â, è, ö, is therefore a character in the same
way as a basic letter.
3.2
vowel
speech sound produced by unobstructed flow of air through the mouth
3.3
consonant
speech sound produced by complete or partial closure of the vocal tract

3.4
transliteration
process which consists of representing the characters (3.1) of an alphabetical or syllabic system of writing
by the characters of a conversion alphabet
3.5
retransliteration
process whereby the characters (3.1) of a conversion alphabet are transformed back into those of the
converted writing system
3.6
transcription
process whereby the sounds of a given language are noted by the system of signs of a conversion language
3.7
romanization
conversion of non-Latin writing systems to the Latin alphabet
4 Strict transliteration
4.1 General
4.1.1 “Hex” values in the following tables shall be interpreted as character codes in ISO/IEC 10646
(Universal Character Set).
4.1.2 The strict transliteration is intended to be a one-to-one reversible transliteration system allowing
for a simple rule-based machine transliteration.
4.1.3 Persian script does not distinguish between upper and lower case. In Latin script, Persian names
may be written using upper or lower case according to the conventions of the target language. This is
optional. Capitalization rules are not part of this document. A system transliterating from Latin into Persian
script shall therefore be case insensitive. Some of the characters, both Latin and Persian, have canonical
decompositions in Unicode. Any transliteration system should treat precomposed and decomposed
characters equally on input.
4.2 Consonants
4.2.1 For a fully reversible transliteration, Table 1 should be used with Table 2 for any vowels included in
the original text.
4.2.2 Different positional forms of Persian characters (initial, medial, final and isolated) are shown in
Table A.1, Annex A.
e
Table 1 — Consonants
No. Persian character Persian name Hex Latin transliteration Hex
a
1 ا alef 0627 ā 0101
2 ب be 0628 b 0062
3 پ pe 067E p 0070
4 ت te 062A t 0074
5 ث s̱ e 062B s̱ 0073+0331
6 ج jīm 062C j 006A
7 چ če 0686 č 010D
8 ح ḥe 062D ḥ 1E25
9 خ xe 062E x 0078
10 د dāl 062F d 0064
11 ذ ẕāl 0630 ẕ 1E95
12 ر re 0631 r 0072
13 ز ze 0632 z 007A
14 ژ že 0698 ž 017E
15 س sīn 0633 s 0073
16 ش šīn 0634 š 0161
17 ص ṣād 0635 ṣ 1E63
18 ض z̤̤ād 0636 z̤ 007A+0324
19 ط ṭā 0637 ṭ 1E6D
20 ظ ẓā 0638 ẓ 1E93
b
21 ع ʻeyn 0639 ʻ 02BB
22 غ ġeyn 063A ġ 0121
23 ف fe 0641 f 0066
24 ق qāf 0642 q 0071
25 ک kāf 06A9 k 006B
26 گ gāf 06AF g 0067
27 ل lām 0644 l 006C
28 م mīm 0645 m 006D
29 ن nūn 0646 n 006E
30 و vāv 0648 v 0076
31 ه he 0647 h 0068
cd
32 ی ye 06CC y 0079
a
For transliteration of âye maddī see 4.3; and for hamze see 4.5. Initial alef may function as the bearer of a short vowel (see 4.3)
or a hamze. In the strict transliteration alef is always transliterated as ‘a�’, hence alef carrying the short vowel pīš ( ) would be
transliterated as ‘a�o’. For example, the name ديما would be ‘a�omyd’ according to this system, or, if the short vowel was unwritten,
‘a�myd’.
b
Implementations may encounter a single left quotation mark (hex 2018) in existing text.
c
Implementations may encounter the Arabic yeh (hex 064A) in existing text.
d
Alef maqṣūre (Arabic hex 0649) is a feature of loan words and names of Arabic origin. In Persian, it is usually written as ی
(hex 06CC) or, for clarity, ٰی (hex 06CC+0670). In the strict transliteration the latter variant is transliterated ‘y�’ (hex 00FD).
e
For the transliteration of hamze, see 4.5.

4.3 Vowels
4.3.1 Generally, Persian words are written without diacritical vowel signs. However, as the change of vowel
sign can bring about a different meaning (for example: رَپَ par = feather; رُپَ por = full), vowel signs may be
used intentionally whenever a difference in meaning must be emphasized. In Table 2 and Table 3, both cases
are represented. However, both the aforementioned examples will be often written رپَ, transliterated ‘pr’ in
the strict univocal system. In transliteration, the diacritical vowel can be included to clarify the meaning.
This would then be included in the Perso-Arabic script if the word were to be reverse transliterated.
Table 2 — Vowels for fully reversible system
Example
Persian Persian Latin
No. Hex Hex
With diacritical Without diacritical
character name transliteration
vowel signs vowel signs
1 آ âye maddī 0622 â 00E2 âẕar رَذآ âẕr رذآ
2 zebar 064E a 0061 sam مَسِ sm مسِ
3 pīš 064F o 006F por رُپَ pr رپَ
4 z̤īr 0650 e 0065 separ رَپَِسِ spr رپَسِ
4.4 Arabic elements in the Persian language
4.4.1 Persian contains many loan words from Arabic. Arabic elements occurring in Persian texts are
treated as follows. Where an Arabic element is present in the text but not mentioned in this document,
[2]
ISO/DIS 233-1:— should be followed.
4.4.2 As with the diacritical vowel markings, these signs, of Arabic origin, are often not written in Persian
script, usually being used only when a difference in meaning is to be emphasized.
Table 3 — Conventional signs
Example
Persian Persian Latin
With
No. Hex Hex
Without diacritical
character name transliteration
diacritical
vowel signs
vowel signs
عَبََّرُم عّبَّرم
1 tašdīd 0651 ʺ 02BA
morabʺaʻ mrbʺʻ
ًًلاًَثََم لاًثَم
064B ã 00E3
mas̱ alāã ms̱ lāã
یرخُا ٍتَرابِعِِبَّ یرخا ٍترابعِبَّ
2 tanvīn 064D ẽ 1EBD
be‘ebāratẽ āoxry b‘bārtẽ āxry
هيَلَِاٌراشُم هيلَاٌراشم
064C õ 00F5
mošārõāelayh mšārõālyh
4.4.3 Tāʼ marbūṭaẗ is not part of the Persian language. Where it occurs on an Arabic word found in a Persian
text, it should be treated according to ISO 233 Arabic transliteration: ة (hex 0629) should be transliterated as
ẗ (hex 1E97).
4.4.4 The Arabic word for God [(الله) hex 0627, 0644, 0644, 0651, 0670, 0647] should be transliterated as
Alla�h.
4.5 Hamze
Hamze ( ء ) is not regarded as a character of the Persian alphabet, but as a diacritical mark, and as such is not
always expressed in writing. In fully-pointed words, however, it appears in several graphic forms, standing
alone or written in conjunction with alef ( أ ), vāv ( ؤ ) and ye ( ئ ). In strict, fully-reversible transliteration,
hamze should be transliterated with apostrophe ( ʼ ) and the character bearing it is transliterated according
to Table 4.
Table 4 — Different forms of hamze in strict transliteration
Example
Persian Latin
No. Persian name Hex Hex
Without
character transliteration
Persian script
diacritical vowels
a
1 ء 0621 ʼ 02BC jzʼ ءزج
a
2 أ 0623 āʼ 0101+02BC rā’s سأر
hamze
a
3 ؤ 0624 vʼ 0076+02BC svʼāl لاؤسِ
a
4 ئ 0626 yʼ 0079+02BC pāyʼyn نيئاپَ
a
Implementations may encounter a single right quotation mark (hex 2019) in existing text.
4.6 Persian relation
...