In a writing system , a letter is a grapheme that generally corresponds to a phoneme —the smallest functional unit of speech—though there is rarely total one-to-one correspondence between the two. An alphabet is a writing system that uses letters.
59-660: Alpha / ˈ æ l f ə / (uppercase Α , lowercase α ) is the first letter of the Greek alphabet . In the system of Greek numerals , it has a value of one. Alpha is derived from the Phoenician letter aleph [REDACTED] , which is the West Semitic word for " ox ". Letters that arose from alpha include the Latin letter A and the Cyrillic letter А . In Ancient Greek , alpha
118-405: A compound in physical chemistry . It is also commonly used in mathematics in algebraic solutions representing quantities such as angles. Furthermore, in mathematics, the letter alpha is used to denote the area underneath a normal curve in statistics to denote significance level when proving null and alternative hypotheses . In ethology , it is used to name the dominant individual in
177-509: A lowercase form (also called minuscule ). Upper- and lowercase letters represent the same sound, but serve different functions in writing. Capital letters are most often used at the beginning of a sentence, as the first letter of a proper name or title, or in headers or inscriptions. They may also serve other functions, such as in the German language where all nouns begin with capital letters. The terms uppercase and lowercase originated in
236-539: A variety of modern uses in mathematics, science, and engineering . People and objects are sometimes named after letters, for one of these reasons: The word letter entered Middle English c. 1200 , borrowed from the Old French letre . It eventually displaced the previous Old English term bōcstæf ' bookstaff '. Letter ultimately descends from the Latin littera , which may have been derived from
295-498: A block is always a multiple of 16, and is often a multiple of 128, but is otherwise arbitrary. Characters required for a given script may be spread out over several different, potentially disjunct blocks within the codespace. Each code point is assigned a classification, listed as the code point's General Category property. Here, at the uppermost level code points are categorized as one of Letter, Mark, Number, Punctuation, Symbol, Separator, or Other. Under each category, each code point
354-710: A calendar year and with rare cases where the scheduled release had to be postponed. For instance, in April 2020, a month after version 13.0 was published, the Unicode Consortium announced they had changed the intended release date for version 14.0, pushing it back six months to September 2021 due to the COVID-19 pandemic . Unicode 16.0, the latest version, was released on 10 September 2024. It added 5,185 characters and seven new scripts: Garay , Gurung Khema , Kirat Rai , Ol Onal , Sunuwar , Todhri , and Tulu-Tigalari . Thus far,
413-432: A comprehensive catalog of character properties, including those needed for supporting bidirectional text , as well as visual charts and reference data sets to aid implementers. Previously, The Unicode Standard was sold as a print volume containing the complete core specification, standard annexes, and code charts. However, version 5.0, published in 2006, was the last version printed this way. Starting with version 5.2, only
472-518: A full semantic duplicate of the Latin alphabet, because legacy CJK encodings contained both "fullwidth" (matching the width of CJK characters) and "halfwidth" (matching ordinary Latin script) characters. The Unicode Bulldog Award is given to people deemed to be influential in Unicode's development, with recipients including Tatsuo Kobayashi , Thomas Milo, Roozbeh Pournader , Ken Lunde , and Michael Everson . The origins of Unicode can be traced back to
531-450: A group of animals. In aerodynamics, the letter is used as a symbol for the angle of attack of an aircraft and the word "alpha" is used as a synonym for this property. In mathematical logic , α is sometimes used as a placeholder for ordinal numbers . The proportionality operator " ∝ " (in Unicode : U+221D) is sometimes mistaken for alpha. The uppercase letter alpha is not generally used as
590-429: A handful of scripts—often primarily between a given script and Latin characters —not between a large number of scripts, and not with all of the scripts supported being treated in a consistent manner. The philosophy that underpins Unicode seeks to encode the underlying characters— graphemes and grapheme-like units—rather than graphical distinctions considered mere variant glyphs thereof, that are instead best handled by
649-530: A low-surrogate code point forms a surrogate pair in UTF-16 in order to represent code points greater than U+FFFF . In principle, these code points cannot otherwise be used, though in practice this rule is often ignored, especially when not using UTF-16. A small set of code points are guaranteed never to be assigned to characters, although third-parties may make independent use of them at their discretion. There are 66 of these noncharacters : U+FDD0 – U+FDEF and
SECTION 10
#1732837702246708-526: A project run by Deborah Anderson at the University of California, Berkeley was founded in 2002 with the goal of funding proposals for scripts not yet encoded in the standard. The project has become a major source of proposed additions to the standard in recent years. The Unicode Consortium together with the ISO have developed a shared repertoire following the initial publication of The Unicode Standard : Unicode and
767-399: A properly engineered design, 16 bits per character are more than sufficient for this purpose. This design decision was made based on the assumption that only scripts and characters in "modern" use would require encoding: Unicode gives higher priority to ensuring utility for the future than to preserving past antiquities. Unicode aims in the first instance at the characters published in
826-591: A symbol because it tends to be rendered identically to the uppercase Latin A . In the International Phonetic Alphabet , the letter ɑ, which looks similar to the lower-case alpha, represents the open back unrounded vowel . The Phoenician alphabet was adopted for Greek in the early 8th century BC, perhaps in Euboea . The majority of the letters of the Phoenician alphabet were adopted into Greek with much
885-558: A total of 168 scripts are included in the latest version of Unicode (covering alphabets , abugidas and syllabaries ), although there are still scripts that are not yet encoded, particularly those mainly used in historical, liturgical, and academic contexts. Further additions of characters to the already encoded scripts, as well as symbols, in particular for mathematics and music (in the form of notes and rhythmic symbols), also occur. The Unicode Roadmap Committee ( Michael Everson , Rick McGowan, Ken Whistler, V.S. Umamaheswaran) maintain
944-648: A universal encoding than the original Unicode architecture envisioned. Version 1.0 of Microsoft's TrueType specification, published in 1992, used the name "Apple Unicode" instead of "Unicode" for the Platform ID in the naming table. The Unicode Consortium is a nonprofit organization that coordinates Unicode's development. Full members include most of the main computer software and hardware companies (and few others) with any interest in text-processing standards, including Adobe , Apple , Google , IBM , Meta (previously as Facebook), Microsoft , Netflix , and SAP . Over
1003-475: Is a text encoding standard maintained by the Unicode Consortium designed to support the use of text in all of the world's writing systems that can be digitized. Version 16.0 of the standard defines 154 998 characters and 168 scripts used in various ordinary, literary, academic, and technical contexts. Many common characters, including numerals, punctuation, and other symbols, are unified within
1062-803: Is a type of grapheme , the smallest functional unit within a writing system. Letters are graphemes that broadly correspond to phonemes , the smallest functional units of sound in speech. Similarly to how phonemes are combined to form spoken words, letters may be combined to form written words. A single phoneme may also be represented by multiple letters in sequence, collectively called a multigraph . Multigraphs include digraphs of two letters (e.g. English ch , sh , th ), and trigraphs of three letters (e.g. English tch ). The same letterform may be used in different alphabets while representing different phonemic categories. The Latin H , Greek eta ⟨Η⟩ , and Cyrillic en ⟨Н⟩ are homoglyphs , but represent different phonemes. Conversely,
1121-506: Is considered to be a separate letter from ⟨n⟩ , though this distinction is not usually recognised in English dictionaries. In computer systems, each has its own code point , U+006E n LATIN SMALL LETTER N and U+00F1 ñ LATIN SMALL LETTER N WITH TILDE , respectively. Letters may also function as numerals with assigned numerical values, for example with Roman numerals . Greek and Latin letters have
1180-426: Is indicated by the existence of precomposed characters for use with computer systems (for example, ⟨á⟩ , ⟨à⟩ , ⟨ä⟩ , ⟨â⟩ , ⟨ã⟩ .) In the following table, letters from multiple different writing systems are shown, to demonstrate the variety of letters used throughout the world. Unicode Unicode , formally The Unicode Standard ,
1239-413: Is intended to suggest a unique, unified, universal encoding". In this document, entitled Unicode 88 , Becker outlined a scheme using 16-bit characters: Unicode is intended to address the need for a workable, reliable world text encoding. Unicode could be roughly described as "wide-body ASCII " that has been stretched to 16 bits to encompass the characters of all the world's living languages. In
SECTION 20
#17328377022461298-428: Is more than just a repertoire within which characters are assigned. To aid developers and designers, the standard also provides charts and reference data, as well as annexes explaining concepts germane to various scripts, providing guidance for their implementation. Topics covered by these annexes include character normalization , character composition and decomposition, collation , and directionality . Unicode text
1357-453: Is not padded. There are a total of 2 + (2 − 2 ) = 1 112 064 valid code points within the codespace. (This number arises from the limitations of the UTF-16 character encoding, which can encode the 2 code points in the range U+0000 through U+FFFF except for the 2 code points in the range U+D800 through U+DFFF , which are used as surrogate pairs to encode the 2 code points in
1416-417: Is processed and stored as binary data using one of several encodings , which define how to translate the standard's abstracted codes for characters into sequences of bytes. The Unicode Standard itself defines three encodings: UTF-8 , UTF-16 , and UTF-32 , though several others exist. Of these, UTF-8 is the most widely used by a large margin, in part due to its backwards-compatibility with ASCII . Unicode
1475-480: Is projected to include 4301 new unified CJK characters . The Unicode Standard defines a codespace : a sequence of integers called code points in the range from 0 to 1 114 111 , notated according to the standard as U+0000 – U+10FFFF . The codespace is a systematic, architecture-independent representation of The Unicode Standard ; actual text is processed as binary data via one of several Unicode encodings, such as UTF-8 . In this normative notation,
1534-400: Is then further subcategorized. In most cases, other properties must be used to adequately describe all the characteristics of any given code point. The 1024 points in the range U+D800 – U+DBFF are known as high-surrogate code points, and code points in the range U+DC00 – U+DFFF ( 1024 code points) are known as low-surrogate code points. A high-surrogate code point followed by
1593-615: The Proto-Indo-European * n̥- ( syllabic nasal) and is cognate with English un- . Copulative a is the Greek prefix ἁ- or ἀ- ha-, a- . It comes from Proto-Indo-European * sm̥ . The letter alpha represents various concepts in physics and chemistry , including alpha radiation , angular acceleration , alpha particles , alpha carbon and strength of electromagnetic interaction (as fine-structure constant ). Alpha also stands for thermal expansion coefficient of
1652-579: The iota subscript ( ᾳ ). In the Attic – Ionic dialect of Ancient Greek, long alpha [aː] fronted to [ ɛː ] ( eta ). In Ionic, the shift took place in all positions. In Attic, the shift did not take place after epsilon , iota , and rho ( ε, ι, ρ ; e, i, r ). In Doric and Aeolic , long alpha is preserved in all positions. Privative a is the Ancient Greek prefix ἀ- or ἀν- a-, an- , added to words to negate them. It originates from
1711-568: The typeface , through the use of markup , or by some other means. In particularly complex cases, such as the treatment of orthographical variants in Han characters , there is considerable disagreement regarding which differences justify their own encodings, and which are only graphical variants of other characters. At the most abstract level, Unicode assigns a unique number called a code point to each character. Many issues of visual representation—including size, shape, and style—are intended to be up to
1770-574: The 1980s, to a group of individuals with connections to Xerox 's Character Code Standard (XCCS). In 1987, Xerox employee Joe Becker , along with Apple employees Lee Collins and Mark Davis , started investigating the practicalities of creating a universal character set. With additional input from Peter Fenwick and Dave Opstad , Becker published a draft proposal for an "international/multilingual text character encoding system in August 1988, tentatively called Unicode". He explained that "the name 'Unicode'
1829-593: The Greek diphthera 'writing tablet' via Etruscan . Until the 19th century, letter was also used interchangeably to refer to a speech segment . Before alphabets, phonograms , graphic symbols of sounds, were used. There were three kinds of phonograms: verbal, pictures for entire words, syllabic, which stood for articulations of words, and alphabetic, which represented signs or letters. The earliest examples of which are from Ancient Egypt and Ancient China, dating to c. 3000 BCE . The first consonantal alphabet emerged around c. 1800 BCE , representing
Alpha - Misplaced Pages Continue
1888-564: The ISO's Universal Coded Character Set (UCS) use identical character names and code points. However, the Unicode versions do differ from their ISO equivalents in two significant ways. While the UCS is a simple character map, Unicode specifies the rules, algorithms, and properties necessary to achieve interoperability between different platforms and languages. Thus, The Unicode Standard includes more information, covering in-depth topics such as bitwise encoding, collation , and rendering. It also provides
1947-653: The Phoenicians, Semitic workers in Egypt. Their script was originally written and read from right to left. From the Phoenician alphabet came the Etruscan and Greek alphabets. From there, the most widely used alphabet today emerged, Latin, which is written and read from left to right. The Phoenician alphabet had 22 letters, nineteen of which the Latin alphabet used, and the Greek alphabet, adapted c. 900 BCE , added four letters to those used in Phoenician. This Greek alphabet
2006-595: The alphabet. Ammonius asks Plutarch what he, being a Boeotian , has to say for Cadmus , the Phoenician who reputedly settled in Thebes and introduced the alphabet to Greece, placing alpha first because it is the Phoenician name for ox —which, unlike Hesiod , the Phoenicians considered not the second or third, but the first of all necessities. "Nothing at all," Plutarch replied. He then added that he would rather be assisted by Lamprias , his own grandfather, than by Dionysus ' grandfather, i.e. Cadmus. For Lamprias had said that
2065-416: The concept of dominant "alpha" members in groups of animals. All code points with ALPHA or ALFA but without WITH (for accented Greek characters, see Greek diacritics: Computer encoding ): These characters are used only as mathematical symbols. Stylized Greek text should be encoded using normal Greek letters, with markup and formatting to indicate text style: Letter (alphabet) A letter
2124-496: The core specification, published as a print-on-demand paperback, may be purchased. The full text, on the other hand, is published as a free PDF on the Unicode website. A practical reason for this publication method highlights the second significant difference between the UCS and Unicode—the frequency with which updated versions are released and new characters added. The Unicode Standard has regularly released annual expanded versions, occasionally with more than one version released in
2183-438: The days of handset type for printing presses. Individual letter blocks were kept in specific compartments of drawers in a type case. Capital letters were stored in a higher drawer or upper case. In most alphabetic scripts, diacritics (or accents) are a routinely used. English is unusual in not using them except for loanwords from other languages or personal names (for example, naïve , Brontë ). The ubiquity of this usage
2242-470: The discretion of the software actually rendering the text, such as a web browser or word processor . However, partially with the intent of encouraging rapid adoption, the simplicity of this original model has become somewhat more elaborate over time, and various pragmatic concessions have been made over the course of the standard's development. The first 256 code points mirror the ISO/IEC 8859-1 standard, with
2301-622: The distinct forms of ⟨S⟩ , the Greek sigma ⟨Σ⟩ , and Cyrillic es ⟨С⟩ each represent analogous /s/ phonemes. Letters are associated with specific names, which may differ between languages and dialects. Z , for example, is usually called zed outside of the United States, where it is named zee . Both ultimately derive from the name of the parent Greek letter zeta ⟨Ζ⟩ . In alphabets, letters are arranged in alphabetical order , which also may vary by language. In Spanish, ⟨ñ⟩
2360-464: The first articulate sound made is "alpha", because it is very plain and simple—the air coming off the mouth does not require any motion of the tongue—and therefore this is the first sound that children make. According to Plutarch's natural order of attribution of the vowels to the planets , alpha was connected with the Moon . As the first letter of the alphabet, Alpha as a Greek numeral came to represent
2419-401: The following versions of The Unicode Standard have been published. Update versions, which do not include any changes to character repertoire, are signified by the third number (e.g., "version 4.0.1") and are omitted in the table below. The Unicode Consortium normally releases a new version of The Unicode Standard once a year. Version 17.0, the next major version,
Alpha - Misplaced Pages Continue
2478-516: The group. By the end of 1990, most of the work of remapping existing standards had been completed, and a final review draft of Unicode was ready. The Unicode Consortium was incorporated in California on 3 January 1991, and the first volume of The Unicode Standard was published that October. The second volume, now adding Han ideographs, was published in June 1992. In 1996, a surrogate character mechanism
2537-549: The intent of trivializing the conversion of text already written in Western European scripts. To preserve the distinctions made by different legacy encodings, therefore allowing for conversion between them and Unicode without any loss of information, many characters nearly identical to others , in both appearance and intended function, were given distinct code points. For example, the Halfwidth and Fullwidth Forms block encompasses
2596-403: The last two code points in each of the 17 planes (e.g. U+FFFE , U+FFFF , U+1FFFE , U+1FFFF , ..., U+10FFFE , U+10FFFF ). The set of noncharacters is stable, and no new noncharacters will ever be defined. Like surrogates, the rule that these cannot be used is often ignored, although the operation of the byte order mark assumes that U+FFFE will never be the first code point in
2655-571: The late 7th and early 8th centuries. Finally, many slight letter additions and drops were made to the common alphabet used in the western world. Minor changes were made such as the removal of certain letters, such as thorn ⟨Þ þ⟩ , wynn ⟨Ƿ ƿ⟩ , and eth ⟨Ð ð⟩ . A letter can have multiple variants, or allographs , related to variation in style of handwriting or printing . Some writing systems have two major types of allographs for each letter: an uppercase form (also called capital or majuscule ) and
2714-625: The list of scripts that are candidates or potential candidates for encoding and their tentative code block assignments on the Unicode Roadmap page of the Unicode Consortium website. For some scripts on the Roadmap, such as Jurchen and Khitan large script , encoding proposals have been made and they are working their way through the approval process. For other scripts, such as Numidian and Rongorongo , no proposal has yet been made, and they await agreement on character repertoire and other details from
2773-675: The modern text (e.g. in the union of all newspapers and magazines printed in the world in 1988), whose number is undoubtedly far below 2 = 16,384. Beyond those modern-use characters, all others may be defined to be obsolete or rare; these are better candidates for private-use registration than for congesting the public list of generally useful Unicode. In early 1989, the Unicode working group expanded to include Ken Whistler and Mike Kernaghan of Metaphor, Karen Smith-Yoshimura and Joan Aliprand of Research Libraries Group , and Glenn Wright of Sun Microsystems . In 1990, Michel Suignard and Asmus Freytag of Microsoft and NeXT 's Rick McGowan had also joined
2832-470: The number 1 . Therefore, Alpha, both as a symbol and term, is used to refer to the "first", or "primary", or "principal" (most significant) occurrence or status of a thing. The New Testament has God declaring himself to be the "Alpha and Omega, the beginning and the end, the first and the last." ( Revelation 22:13 , KJV, and see also 1:8 ). Consequently, the term "alpha" has also come to be used to denote "primary" position in social hierarchy, examples being
2891-554: The previous environment of a myriad of incompatible character sets , each used within different locales and on different computer architectures. Unicode is used to encode the vast majority of text on the Internet, including most web pages , and relevant Unicode support has become a common consideration in contemporary software development. The Unicode character repertoire is synchronized with ISO/IEC 10646 , each being code-for-code identical with one another. However, The Unicode Standard
2950-807: The range U+10000 through U+10FFFF .) The Unicode codespace is divided into 17 planes , numbered 0 to 16. Plane 0 is the Basic Multilingual Plane (BMP), and contains the most commonly used characters. All code points in the BMP are accessed as a single code unit in UTF-16 encoding and can be encoded in one, two or three bytes in UTF-8. Code points in planes 1 through 16 (the supplementary planes ) are accessed as surrogate pairs in UTF-16 and encoded in four bytes in UTF-8 . Within each plane, characters are allocated within named blocks of related characters. The size of
3009-455: The same sounds as they had had in Phoenician, but ʼāleph , the Phoenician letter representing the glottal stop [ʔ] , was adopted as representing the vowel [a] ; similarly, hē [h] and ʽayin [ʕ] are Phoenician consonants that became Greek vowels, epsilon [e] and omicron [o] , respectively. Plutarch , in Moralia , presents a discussion on why the letter alpha stands first in
SECTION 50
#17328377022463068-493: The standard and are not treated as specific to any given writing system. Unicode encodes 3790 emoji , with the continued development thereof conducted by the Consortium as a part of the standard. Moreover, the widespread adoption of Unicode was in large part responsible for the initial popularization of emoji outside of Japan. Unicode is ultimately capable of encoding more than 1.1 million characters. Unicode has largely supplanted
3127-418: The two-character prefix U+ always precedes a written code point, and the code points themselves are written as hexadecimal numbers. At least four hexadecimal digits are always written, with leading zeros prepended as needed. For example, the code point U+00F7 ÷ DIVISION SIGN is padded with two leading zeros, but U+13254 𓉔 EGYPTIAN HIEROGLYPH O004 ( [REDACTED] )
3186-607: The user communities involved. Some modern invented scripts which have not yet been included in Unicode (e.g., Tengwar ) or which do not qualify for inclusion in Unicode due to lack of real-world use (e.g., Klingon ) are listed in the ConScript Unicode Registry , along with unofficial but widely used Private Use Areas code assignments. There is also a Medieval Unicode Font Initiative focused on special Latin medieval characters. Part of these proposals has been already included in Unicode. The Script Encoding Initiative,
3245-635: The years several countries or government agencies have been members of the Unicode Consortium. Presently only the Ministry of Endowments and Religious Affairs (Oman) is a full member with voting rights. The Consortium has the ambitious goal of eventually replacing existing character encoding schemes with Unicode and its standard Unicode Transformation Format (UTF) schemes, as many of the existing schemes are limited in size and scope and are incompatible with multilingual environments. Unicode currently covers most major writing systems in use today. As of 2024 ,
3304-491: Was implemented in Unicode 2.0, so that Unicode was no longer restricted to 16 bits. This increased the Unicode codespace to over a million code points, which allowed for the encoding of many historic scripts, such as Egyptian hieroglyphs , and thousands of rarely used or obsolete characters that had not been anticipated for inclusion in the standard. Among these characters are various rarely used CJK characters—many mainly being used in proper names, making them far more necessary for
3363-483: Was originally designed with the intent of transcending limitations present in all text encodings designed up to that point: each encoding was relied upon for use in its own context, but with no particular expectation of compatibility with any other. Indeed, any two encodings chosen were often totally unworkable when used together, with text encoded in one interpreted as garbage characters by the other. Most encodings had only been designed to facilitate interoperation between
3422-626: Was pronounced [ a ] and could be either phonemically long ([aː]) or short ([a]). Where there is ambiguity, long and short alpha are sometimes written with a macron and breve today: Ᾱᾱ, Ᾰᾰ . In Modern Greek , vowel length has been lost, and all instances of alpha simply represent the open front unrounded vowel IPA: [a] . In the polytonic orthography of Greek, alpha, like other vowel letters, can occur with several diacritic marks: any of three accent symbols ( ά, ὰ, ᾶ ), and either of two breathing marks ( ἁ, ἀ ), as well as combinations of these. It can also combine with
3481-504: Was the first to assign letters not only to consonant sounds, but also to vowels . The Roman Empire further developed and refined the Latin alphabet, beginning around 500 BCE. During the fifth and sixth centuries, the development of lowercase letters began to emerge in Roman writing. At this point, paragraphs, uppercase and lowercase letters, and the concept of sentences and clauses still had not emerged; these final bits of development emerged in
#245754