Unicode

E3674

Unicode is a universal character encoding standard that assigns unique code points to virtually all written scripts, symbols, and emojis used in modern computing.

All labels observed (10)

How this entity was disambiguated

Statements (60)

Predicate Object
instanceOf character encoding standard
international standard
abbreviation Unicode self-link
alignedWith ISO/IEC 10646
basicMultilingualPlaneRange U+0000 to U+FFFF
codeSpaceRange U+0000 to U+10FFFF
compatibleWith ISO/IEC 10646
defines bidirectional text behavior
character properties
code points
collation rules
encoding forms
grapheme cluster boundaries
line breaking rules
normalization forms
developedBy Unicode Consortium
documentationFormat multi-volume standard text and data files
firstPublished 1991
fullName Unicode self-linksurface differs
surface form: The Unicode Standard
goal interoperability across platforms and languages
universal character set
hasEncodingForm UTF-16
UTF-32
UTF-8
hasGoverningBody Unicode Consortium
hasTechnicalReport Unicode Technical Report #29
hasTechnicalStandard Unicode Technical Standard #10
includesDatabase Unicode Character Database
initialVersion Unicode self-linksurface differs
surface form: Unicode 1.0
latestVersion Unicode 15.1
license freely available standard
maintainedBy Unicode Consortium
mostCommonEncodingOnWeb UTF-8
numberOfPlanes 17
organizesInto planes
provides Unicode Scalar Values
replaces many legacy character encodings
supplementaryPlanesRange U+10000 to U+10FFFF
supports Arabic script
Chinese characters
Cyrillic script
Devanagari script
Greek script
Hebrew script
Japanese scripts
Korean Hangul
Latin script
currency symbols
emoji
historic scripts
mathematical symbols
musical notation symbols
punctuation
technical symbols
usedIn databases
modern operating systems
modern programming languages
web technologies
usesBitWidth 21-bit code space
versioningScheme major.minor

How these facts were elicited

Referenced by (156)

Full triples — surface form annotated when it differs from this entity's canonical label.

Sentence_Break definedIn Unicode
this entity surface form: Unicode Standard
Telu encodingContext Unicode
Unicode Standard Annexes associatedWith Unicode
this entity surface form: Unicode Standard
Unicode, Inc. develops Unicode
this entity surface form: Unicode Standard
Unicode, Inc. maintains Unicode
this entity surface form: Unicode Standard
Unicode, Inc. shortName Unicode
Unicode ICU supports Unicode
UCS compatibleWith Unicode
ISO/IEC 8859 relatedStandard Unicode
RFC 3629 relatesTo Unicode
Tfng relatedStandard Unicode
Kawa supportsStandard Unicode
UTF-8 designedFor Unicode
MODS supportsEncoding Unicode
Nbat unicodeStandard Unicode
ISO 15924 relatedTo Unicode
Buhid hasScriptCodeStandard Unicode
Kana encodingStandard Unicode
Tagbanwa hasDigitalEncoding Unicode
Joe Becker fieldOfWork Unicode
Joe Becker notableWork Unicode
this entity surface form: Unicode character encoding standard
iTerm2 supportsFeature Unicode
RFC 4518 uses Unicode
TEI compatibleWith Unicode
Zhuyin encodingStandard Unicode
QuarkXPress supportsStandard Unicode
Karen script hasEncoding Unicode
XeTeX supports Unicode
Vai script encodedInStandard Unicode
LuaTeX supports Unicode
upTeX supportsEncoding Unicode
LyX supports Unicode
Vim supports Unicode
Hebr relatedStandard Unicode
encodingStandard Unicode
Armn usedInStandard Unicode