Triple

T135511
Position Surface form Disambiguated ID Type / Status
Subject Turkish language E2737 entity
Predicate belongsToLanguageGroup P3349 FINISHED
Object Southwestern Turkic
Southwestern Turkic is a major branch of the Turkic language family that includes languages such as Turkish, Azerbaijani, and Turkmen, primarily spoken across Anatolia, the Caucasus, and parts of Central and Western Asia.
E16295 NE FINISHED

How this triple was built (5 steps)

Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.

NER Named-entity recognition gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: Southwestern Turkic | Statement: [Turkish language, belongsToLanguageGroup, Southwestern Turkic]
NED1 Entity disambiguation (via context triple) gpt-5-mini-2025-08-07
Target entity: Southwestern Turkic
Context triple: [Turkish language, belongsToLanguageGroup, Southwestern Turkic]
  • A. Tocharian languages
    The Tocharian languages were an extinct branch of the Indo-European family once spoken in the Tarim Basin of Central Asia, known from early medieval manuscripts and notable for their archaic linguistic features.
  • B. Nuristani languages
    Nuristani languages are a small, distinct group of Indo-Iranian languages spoken primarily in the remote Nuristan region of eastern Afghanistan.
  • C. Anatolian languages
    Anatolian languages are an extinct branch of the Indo-European language family once spoken in ancient Anatolia, including languages such as Hittite and Luwian.
  • D. Hokan languages
    Hokan languages are a proposed but controversial grouping of several Native American language families of the western United States and Mexico that share certain typological and lexical similarities.
  • E. Altaic languages (proposed)
    Altaic languages (proposed) is a controversial hypothetical language family that groups together Turkic, Mongolic, Tungusic, and sometimes Koreanic and Japonic languages, primarily spoken across northern and central Asia.
  • F. None of above. chosen
  • G. Unsure - the case is ambiguous/there is not enough information to decide.
NEDg Description generation gpt-5.1
Instruction
Generate a one-sentence description of the target entity. 
You are given a context triple in the form (subject, predicate, object), where the object is the target entity. 
# Instructions
Use the triple to infer relevant information about the entity. Describe the entity based on what is most defining, well-known. 
Avoid repeating the information from the triple, unless really essential.
# Response Format
Return only the sentence: "Description: [one-sentence description of the target entity]"
Input
Entity: Southwestern Turkic
Triple: [Turkish language, belongsToLanguageGroup, Southwestern Turkic]
Generated description
Southwestern Turkic is a major branch of the Turkic language family that includes languages such as Turkish, Azerbaijani, and Turkmen, primarily spoken across Anatolia, the Caucasus, and parts of Central and Western Asia.
NED2 Entity disambiguation (via description) gpt-5-mini-2025-08-07
Target entity: Southwestern Turkic
Target entity description: Southwestern Turkic is a major branch of the Turkic language family that includes languages such as Turkish, Azerbaijani, and Turkmen, primarily spoken across Anatolia, the Caucasus, and parts of Central and Western Asia.
  • A. Tocharian languages
    The Tocharian languages were an extinct branch of the Indo-European family once spoken in the Tarim Basin of Central Asia, known from early medieval manuscripts and notable for their archaic linguistic features.
  • B. Nuristani languages
    Nuristani languages are a small, distinct group of Indo-Iranian languages spoken primarily in the remote Nuristan region of eastern Afghanistan.
  • C. Anatolian languages
    Anatolian languages are an extinct branch of the Indo-European language family once spoken in ancient Anatolia, including languages such as Hittite and Luwian.
  • D. Hokan languages
    Hokan languages are a proposed but controversial grouping of several Native American language families of the western United States and Mexico that share certain typological and lexical similarities.
  • E. Altaic languages (proposed)
    Altaic languages (proposed) is a controversial hypothetical language family that groups together Turkic, Mongolic, Tungusic, and sometimes Koreanic and Japonic languages, primarily spoken across northern and central Asia.
  • F. None of above. chosen
PD Predicate disambiguation gpt-5-mini-2025-08-07
Target predicate: belongsToLanguageGroup
Context triple: [Turkish language, belongsToLanguageGroup, Southwestern Turkic]
  • A. hasLanguageGroup chosen
    Indicates that an entity belongs to, is associated with, or is categorized under a particular language group.
  • B. includesLanguage
    Indicates that one entity contains, supports, or makes use of a specified language as part of its content, functionality, or representation.
  • C. isLanguageOf
    Indicates that a particular language is used as the official or primary language associated with a given entity (such as a person, document, or region).
  • D. correspondsToLanguageClass
    Indicates that something is associated with, mapped to, or classified under a particular language class.
  • E. languageBranch
    Indicates that one language belongs to, or is classified under, a broader linguistic branch or subgroup.
  • F. None of above.

Provenance (6 batches)

The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.

Step Stage Batch ID Status When
creating Elicitation batch_69a2520c0f3481908b0ed054a2fca8d0 completed Feb. 28, 2026, 2:25 a.m.
NER Named-entity recognition batch_69a257a3ad908190b6a8652f09ae0cbb completed Feb. 28, 2026, 2:49 a.m.
NED1 Entity disambiguation (via context triple) batch_69a2b4bad1a0819098459e2a9d6b8d2a completed Feb. 28, 2026, 9:26 a.m.
NEDg Description generation batch_69a2b55279a08190b9e73f6f9faf1ebe completed Feb. 28, 2026, 9:28 a.m.
NED2 Entity disambiguation (via description) batch_69a2b5e5c0d08190a1da63663e5c932d completed Feb. 28, 2026, 9:31 a.m.
PD Predicate disambiguation batch_69a25651b9048190a6277b7fec98c1ea completed Feb. 28, 2026, 2:43 a.m.
Created at: Feb. 28, 2026, 2:30 a.m.