Triple

T3185968
Position Surface form Disambiguated ID Type / Status
Subject South Asian ethnic groups E66698 entity
Predicate includesEthnicGroup P1898 FINISHED
Object Gujaratis E59007 NE FINISHED

How this triple was built (2 steps)

Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.

NER Named-entity recognition gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: Gujaratis | Statement: [South Asian ethnic groups, includesEthnicGroup, Gujaratis]
NED1 Entity disambiguation (via context triple) gpt-5-mini-2025-08-07
Target entity: Gujaratis
Context triple: [South Asian ethnic groups, includesEthnicGroup, Gujaratis]
  • A. Gujarati people chosen
    Gujarati people are an Indo-Aryan ethnic group primarily from the Indian state of Gujarat, known for their distinct language, vibrant culture, entrepreneurial tradition, and widespread global diaspora.
  • B. Gujarati
    Gujarati is an Indo-Aryan language primarily spoken in the Indian state of Gujarat and by Gujarati communities worldwide.
  • C. Marwari
    Marwari is an Indo-Aryan language of western India, primarily spoken in Rajasthan and surrounding regions, and closely related to Gujarati and other Rajasthani dialects.
  • D. Hindko
    Hindko is a group of Indo-Aryan dialects spoken primarily in northern Pakistan, especially in parts of Khyber Pakhtunkhwa and Azad Kashmir.
  • E. Sindhi
    Sindhi is an Indo-Aryan language spoken primarily in Pakistan and India, known for its rich literary tradition and distinct script variants.
  • F. None of above.
  • G. Unsure - the case is ambiguous/there is not enough information to decide.

Provenance (3 batches)

The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.

Step Stage Batch ID Status When
creating Elicitation batch_69ad8587c1bc8190a2595f2c22ee1001 completed March 8, 2026, 2:19 p.m.
NER Named-entity recognition batch_69ada6c32ce88190a231be18d38ec5ba completed March 8, 2026, 4:41 p.m.
NED1 Entity disambiguation (via context triple) batch_69b24b8285388190bc47264bd1028ab3 completed March 12, 2026, 5:13 a.m.
Created at: March 8, 2026, 3:06 p.m.