Triple

T11775282
Position Surface form Disambiguated ID Type / Status
Subject Institut_für_deutsche_Sprache E280001 entity
Predicate operates P24 FINISHED
Object COSMAS II corpus search system
COSMAS II corpus search system is a large-scale linguistic search platform for German language text corpora, maintained by the Institut für Deutsche Sprache for research and lexicographic analysis.
E945990 NE FINISHED

How this triple was built (4 steps)

Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.

NER Named-entity recognition gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: COSMAS II corpus search system | Statement: [Institut_für_deutsche_Sprache, operates, COSMAS II corpus search system]
NED1 Entity disambiguation (via context triple) gpt-5-mini-2025-08-07
Target entity: COSMAS II corpus search system
Context triple: [Institut_für_deutsche_Sprache, operates, COSMAS II corpus search system]
  • A. CiteSeerX
    CiteSeerX is a public digital library and search engine that focuses on indexing and providing access to scientific and academic research papers, particularly in computer and information science.
  • B. Darwin Information Typing Architecture
    Darwin Information Typing Architecture (DITA) is an XML-based, topic-oriented architecture and standard for authoring, organizing, and publishing technical content in a modular and reusable way.
  • C. Corpus
    Corpus is a common shortened name for Corpus Christi College, one of the historic constituent colleges of the University of Cambridge.
  • D. BNF bibliographic database
    The BNF bibliographic database is the comprehensive online catalog of the Bibliothèque nationale de France, providing detailed bibliographic records for its collections of books, manuscripts, and other documents.
  • E. Resnik
    Resnik is a surname most notably associated with Judith Resnik, the American astronaut who died in the Space Shuttle Challenger disaster.
  • F. None of above. chosen
  • G. Unsure - the case is ambiguous/there is not enough information to decide.
NEDg Description generation gpt-5.1
Instruction
Generate a one-sentence description of the target entity. 
You are given a context triple in the form (subject, predicate, object), where the object is the target entity. 
# Instructions
Use the triple to infer relevant information about the entity. Describe the entity based on what is most defining, well-known. 
Avoid repeating the information from the triple, unless really essential.
# Response Format
Return only the sentence: "Description: [one-sentence description of the target entity]"
Input
Entity: COSMAS II corpus search system
Triple: [Institut_für_deutsche_Sprache, operates, COSMAS II corpus search system]
Generated description
COSMAS II corpus search system is a large-scale linguistic search platform for German language text corpora, maintained by the Institut für Deutsche Sprache for research and lexicographic analysis.
NED2 Entity disambiguation (via description) gpt-5-mini-2025-08-07
Target entity: COSMAS II corpus search system
Target entity description: COSMAS II corpus search system is a large-scale linguistic search platform for German language text corpora, maintained by the Institut für Deutsche Sprache for research and lexicographic analysis.
  • A. CiteSeerX
    CiteSeerX is a public digital library and search engine that focuses on indexing and providing access to scientific and academic research papers, particularly in computer and information science.
  • B. Darwin Information Typing Architecture
    Darwin Information Typing Architecture (DITA) is an XML-based, topic-oriented architecture and standard for authoring, organizing, and publishing technical content in a modular and reusable way.
  • C. Corpus
    Corpus is a common shortened name for Corpus Christi College, one of the historic constituent colleges of the University of Cambridge.
  • D. BNF bibliographic database
    The BNF bibliographic database is the comprehensive online catalog of the Bibliothèque nationale de France, providing detailed bibliographic records for its collections of books, manuscripts, and other documents.
  • E. Resnik
    Resnik is a surname most notably associated with Judith Resnik, the American astronaut who died in the Space Shuttle Challenger disaster.
  • F. None of above. chosen

Provenance (5 batches)

The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.

Step Stage Batch ID Status When
creating Elicitation batch_69d6ab01d2688190ad8ed6bda487eaa5 completed April 8, 2026, 7:22 p.m.
NER Named-entity recognition batch_69d8a55f415081908eec78cb2c956598 completed April 10, 2026, 7:23 a.m.
NED1 Entity disambiguation (via context triple) batch_69f090a95a908190a99e579e51cbeb4a completed April 28, 2026, 10:49 a.m.
NEDg Description generation batch_69f0bd3e585481908223acfd780a72a2 completed April 28, 2026, 1:59 p.m.
NED2 Entity disambiguation (via description) batch_69f0ef31076c8190b33a6a2778d7ffbb completed April 28, 2026, 5:32 p.m.
Created at: April 8, 2026, 9:41 p.m.