Qortora · Search · Indexed page

en.wikipedia.orgFetched 2026-08-15T00:53:26Z

Unicode - Wikipedia

Unicode - Wikipedia Jump to content Main menu Main menu move to sidebar hide Navigation Main page Contents Current events Random article About Wikipedia Contact us Contribute Help Learn to edit Community portal Recent changes Upload file Special pages Search Search Appearance Don…

Open original source · Full cached text

Unicode - Wikipedia Jump to content Main menu Main menu move to sidebar hide Navigation Main page Contents Current events Random article About Wikipedia Contact us Contribute Help Learn to edit Community portal Recent changes Upload file Special pages Search Search Appearance Donate Create account Log in Personal tools Donate Create account Log in Contents move to sidebar hide (Top) 1 Origin and development Toggle Origin and development subsection 1.1 History 1.2 Unicode Consortium 1.3 Scripts covered 1.4 Proposals for adding scripts 1.5 Versions 2 Architecture and terminology Toggle Architecture and terminology subsection 2.1 Codespace and code points 2.2 Code planes and blocks 2.3 General Category property 2.4 Abstract characters 2.5 Precomposed vis-à-vis composite characters 2.6 Ligatures 2.7 Standardized subsets 2.8 Mapping and encodings 3 Adoption Toggle Adoption subsection 3.1 Operating systems 3.2 Input methods 3.3 Email 3.4 Web 3.5 Fonts 3.6 Newlines 4 Issues Toggle Issues subsection 4.1 Character unification 4.1.1 Han unification 4.1.2 Italic or cursive characters in Cyrillic 4.1.3 Localised case pairs 4.1.4 Diacritics on lowercase I 4.2 Security 4.3 Mapping to legacy character sets 4.4 Indic scripts 4.5 Combining characters 4.6 Anomalies 5 See also 6 Notes 7 References 8 Further reading 9 External links Toggle the table of contents Unicode 122 languages Afrikaans Alemannisch አማርኛ العربية অসমীয়া Asturianu Azərbaycanca Boarisch Беларуская (тарашкевіца) Беларуская Български भोजपुरी ပအိုဝ်ႏဘာႏသာႏ বাংলা Brezhoneg Bosanski Català 閩東語 / Mìng-dĕ̤ng-ngṳ̄ ᏣᎳᎩ کوردی Čeština Чӑвашла Cymraeg Dansk Deutsch Ελληνικά Esperanto Español Eesti Euskara فارسی Suomi Français Gaeilge Galego ગુજરાતી 客家語 / Hak-kâ-ngî עברית हिन्दी Hrvatski Magyar Հայերեն Interlingua Bahasa Indonesia Ilokano Íslenska Medžuslovjansky Italiano 日本語 Jawa ქართული Қазақша ಕನ್ನಡ 한국어 کٲشُر Kurdî Кыргызча Latina Lingua Franca Nova Lietuvių Latviešu मैथिली Олык марий Македонски മലയാളം Монгол ဘာသာမန် मराठी Bahasa Melayu မြန်မာဘာသာ Plattdüütsch नेपाली नेपाल भाषा Nederlands Norsk nynorsk Norsk bokmål Occitan ਪੰਜਾਬੀ Papiamentu Polski پنجابی پښتو Português Română Русский संस्कृतम् Саха тыла ᱥᱟᱱᱛᱟᱲᱤ Scots سنڌي Srpskohrvatski / српскохрватски සිංහල Simple English Slovenčina Slovenščina Shqip Српски / srpski Sunda Svenska Kiswahili தமிழ் తెలుగు Тоҷикӣ ไทย Tagalog Toki pona Türkçe Тыва дыл ئۇيغۇرچە / Uyghurche Українська اردو Oʻzbekcha / ўзбекча Tiếng Việt Walon 吴语 მარგალური ייִדיש Yorùbá 文言 閩南語 / Bân-lâm-gí 粵語 中文 Edit links Article Talk English Read Edit View history Tools Tools move to sidebar hide Actions Read Edit View history General What links here Related changes Upload file Permanent link Page information Cite this page Get shortened URL Switch to legacy parser Print/export Download as PDF Printable version In other projects Wikimedia Commons Wikibooks Wikidata item Appearance move to sidebar hide From Wikipedia, the free encyclopedia Character encoding standard UnicodeLogo of the Unicode Consortium Alias(es) Universal Coded Character Set (UCS) ISO/IEC 10646 Languages 172 scripts (list) Standard Unicode Standard Encoding formats UTF-8 UTF-16 GB18030 UTF-32 BOCU SCSU UTF-EBCDIC (uncommon) UTF-7 UTF-1 (obsolete) Preceded by ISO/IEC 8859, among others Official website Technical website This article contains uncommon Unicode characters. Without proper rendering support, you may see question marks, boxes, or other symbols. Unicode (also known as The Unicode Standard and TUS[1][2]) is a character encoding standard maintained by the Unicode Consortium designed to support the use of text in all of the world's writing systems that can be digitized. Version 17.0[A] defines 159,801 characters and 172 scripts[3] used in various ordinary, literary, academic and technical contexts. Unicode has largely supplanted the previous environment of myriad incompatible character sets used within different locales and on different computer architectures. The entire repertoire of these sets, plus many additional characters, were merged into the single Unicode set. Unicode is used to encode the vast majority of text on the Internet, including most web pages, and relevant Unicode support has become a common consideration in contemporary software development. Unicode is ultimately capable of encoding more than 1.1 million characters. The Unicode character repertoire is synchronized with ISO/IEC 10646, each being code-for-code identical with one another. However, The Unicode Standard is more than just a repertoire within which characters are assigned. To aid developers and designers, the standard also provides charts and reference data, as well as annexes explaining concepts germane to various scripts, providing guidance for their implementation. Topics covered by these annexes include character normalization, character composition and decomposition, collation, and directionality.[4] Unicode encodes 3,790 emoji, with the continued development thereof conducted by the Consortium as a part of the standard.[5] The widespread adoption of Unicode was in large part responsible for the initial popularization of emoji outside of Japan.[citation needed] Unicode text is processed and stored as binary data using one of several encodings, which define how to translate the standard's abstracted codes for characters into sequences of bytes. The Unicode Standard itself defines three encodings: UTF-8, UTF-16,[a] and UTF-32, though several others exist. UTF-8 is the most widely used by a large margin, in part due to its backwards-compatibility with ASCII. Origin and development [edit] Unicode was originally designed with the intent of transcending limitations present in all text encodings designed up to that point: each encoding was relied upon for use in its own context, but with no particular expectation of compatibility with any other. Indeed, any two encodings chosen were often totally unworkable when used together, with text encoded in one interpreted as garbage characters by the other. Most encodings had only been designed to facilitate interoperation between a handful of scripts—often primarily between a given script and Latin characters—not between a large number of scripts, and not with all of the scripts supported being treated in a consistent manner. The philosophy that underpins Unicode seeks to encode the underlying characters—graphemes and grapheme-like units—rather than graphical distinctions considered mere variant glyphs thereof, that are instead best handled by the typeface, through the use of markup, or by some other means. In particularly complex cases, such as the treatment of orthographical variants in Han characters, there is considerable disagreement regarding which differences justify their own encodings, and which are only graphical variants of other characters. At the most abstract level, Unicode assigns a unique number called a code point to each character. Many issues of visual representation—including size, shape, and style—are intended to be up to the discretion of the software actually rendering the text, such as a web browser or word processor. However, partially with the intent of encouraging rapid adoption, the simplicity of this original model has become somewhat more elaborate over time, and various pragmatic concessions have been made over the course of the standard's development. The first 256 code points mirror the ISO/IEC 8859-1 standard, with the intent of trivializing the conversion of text already written in Western European scripts. To preserve the distinctions made by different legacy encodings, therefore allowing for conversion between them and Unicode without any loss of information, many characters nearly identical to others, in both appearance and intended function, were given distinct code points. For example, the Halfwidth and Fullwidth Forms block encompasses a full semantic duplicate of the Latin alphabet, because legacy CJK encodings contained both "fullwidth" (matching the width of CJK characters) and "halfwidth" (matching ordinary Latin script) characters. History [edit] The origins of Unicode can be traced back to the 1980s, to a group of individuals with connections to Xerox's Character Code Standard (XCCS).[6] In 1987, Xerox employee Joe Becker, along with Apple employees Lee Collins and Mark Davis, started investigating the practicalities of creating a universal character set.[7] With additional input from Peter Fenwick and Dave Opstad,[6] Becker published a draft proposal for an "international/multilingual text character encoding system in August 1988, tentatively called Unicode". He explained that "the name 'Unicode' is intended to suggest a unique, unified, universal encoding".[6] In this document, entitled Unicode 88, Becker outlined a scheme using 16-bit characters:[6] Unicode is intended to address the need for a workable, reliable world text encoding. Unicode could be roughly described as "wide-body ASCII" that has been stretched to 16 bits to encompass the characters of all the world's living languages. In a properly engineered design, 16 bits per character are more than sufficient for this purpose. This design decision was made based on the assumption that only scripts and characters in "modern" use would require encoding:[6] Unicode gives higher priority to ensuring utility for the future than to preserving past antiquities. Unicode aims in the first instance at the characters published in the modern text (e.g. in the union of all newspapers and magazines printed in the world in 1988), whose number is undoubtedly far below 214 = 16,384. Beyond those modern-use characters, all others may be defined to be obsolete or rare; these are better candidates for private use registration than for congesting the public list of generally useful Unicode. In early 1989, the Unicode working group expanded to include Ken Whistler and Mike Kernaghan of Metaphor, Karen Smith-Yoshimura and Joan Aliprand of Research Libraries Group, and Glenn Wright of Sun Microsystems. The Research Libraries Group had an existing solution for East Asian character sets, which became one of the inputs to the Unicode character set.[7] In 1990, Michel Suignard and Asmus Freytag of Microsoft and NeXT's Rick McGowan had also joined the group. By the end of 1990, most of the work of remapping existing standards had been completed, and a final review draft of Unicode was ready. The Unicode Consortium was incorporated in California on 3 January 1991,[8] and the first volume of The Unicode Standard was published that October. The second volume, now adding Han ideographs, was published in June 1992. In 1996, a surrogate character mechanism was implemented in Unicode 2.0, so that Unicode was no longer restricted to 16 bits. This increased the Unicode codespace to over a million code points, which allowed for the encoding of many historic scripts, such as Egyptian hieroglyphs, and thousands of rarely used or obsolete characters that had not been anticipated for inclusion in the standard. Among these characters are various rarely used CJK characters—many mainly being used in proper names, making them far more necessary for a universal encoding than the original Unicode architecture envisioned.[9] Unicode Consortium [edit] Main article: Unicode Consortium The Unicode Consortium is a non-profit orga…