Unicode

E3674

Unicode is a universal character encoding standard that assigns unique code points to virtually all written scripts, symbols, and emojis used in modern computing.

All labels observed (10)

How this entity was disambiguated

Statements (60)

Predicate Object
instanceOf character encoding standard
international standard
abbreviation Unicode self-link
alignedWith ISO/IEC 10646
basicMultilingualPlaneRange U+0000 to U+FFFF
codeSpaceRange U+0000 to U+10FFFF
compatibleWith ISO/IEC 10646
defines bidirectional text behavior
character properties
code points
collation rules
encoding forms
grapheme cluster boundaries
line breaking rules
normalization forms
developedBy Unicode Consortium
documentationFormat multi-volume standard text and data files
firstPublished 1991
fullName Unicode self-linksurface differs
surface form: The Unicode Standard
goal interoperability across platforms and languages
universal character set
hasEncodingForm UTF-16
UTF-32
UTF-8
hasGoverningBody Unicode Consortium
hasTechnicalReport Unicode Technical Report #29
hasTechnicalStandard Unicode Technical Standard #10
includesDatabase Unicode Character Database
initialVersion Unicode self-linksurface differs
surface form: Unicode 1.0
latestVersion Unicode 15.1
license freely available standard
maintainedBy Unicode Consortium
mostCommonEncodingOnWeb UTF-8
numberOfPlanes 17
organizesInto planes
provides Unicode Scalar Values
replaces many legacy character encodings
supplementaryPlanesRange U+10000 to U+10FFFF
supports Arabic script
Chinese characters
Cyrillic script
Devanagari script
Greek script
Hebrew script
Japanese scripts
Korean Hangul
Latin script
currency symbols
emoji
historic scripts
mathematical symbols
musical notation symbols
punctuation
technical symbols
usedIn databases
modern operating systems
modern programming languages
web technologies
usesBitWidth 21-bit code space
versioningScheme major.minor

How these facts were elicited

Referenced by (156)

Full triples — surface form annotated when it differs from this entity's canonical label.

Georgian Supplement assignedInStandard Unicode
this entity surface form: Unicode Standard
Unicode 15.0 partOf Unicode
this entity surface form: Unicode Standard
Icelandic hasDigitalEncoding Unicode
Perl supports Unicode
BMP partOf Unicode
subject surface form: Basic Multilingual Plane
BMP partOf Unicode
subject surface form: Basic Multilingual Plane
this entity surface form: Unicode Standard
Plane 0 partOf Unicode
this entity surface form: Unicode Standard
Plane 0 isPrimaryBlockOf Unicode
Unicode Core Specification defines Unicode
this entity surface form: Unicode character set
Unicode Core Specification hasComponent Unicode
this entity surface form: Unicode normalization
UTF-7 supports Unicode
ASA influenced Unicode
subject surface form: ASCII
Thai script standardizedIn Unicode
this entity surface form: Unicode Standard
UTS #10 relatedTo Unicode
this entity surface form: Unicode Standard
UCA partOf Unicode
subject surface form: Unicode Collation Algorithm
this entity surface form: Unicode Standard
DUCET relatedTo Unicode
this entity surface form: Unicode Standard
OpenType font technology supports Unicode
subject surface form: OpenType
Default Unicode Collation Element Table definedInStandard Unicode
this entity surface form: Unicode Standard
XML technology stack basedOn Unicode
Supplementary Ideographic Plane partOf Unicode
this entity surface form: Unicode Standard
Unicode CLDR relatedTo Unicode
this entity surface form: Unicode Standard
Emoji Subcommittee usesStandard Unicode
this entity surface form: Unicode Standard
Unicode 5.1 partOf Unicode
this entity surface form: Unicode Standard
Unicode 5.1 standardizes Unicode
this entity surface form: Unicode character set
Cyrillic Extended-A partOfStandard Unicode
this entity surface form: Unicode Standard
Cyrillic Extended-C standard Unicode
this entity surface form: Unicode Standard
Mark relatedTo Unicode
Mark Davis contributedTo Unicode
this entity surface form: Unicode Standard
Mark Davis hasExpertise Unicode
Mark Davis helpedStandardize Unicode
this entity surface form: Unicode character encoding system
Hangul Jamo Extended-A standard Unicode
this entity surface form: Unicode Standard
Hangul Jamo Extended-B belongsTo Unicode
this entity surface form: Unicode Standard
Grantha (U+11300–U+1137F) standard Unicode
subject surface form: Grantha (Unicode block)
Unicode Hebrew block partOf Unicode
this entity surface form: Unicode Standard
UTR #29 relatedTo Unicode
this entity surface form: Unicode Standard
Grapheme_Cluster_Break definedIn Unicode
this entity surface form: Unicode Standard