Triple

T18784086
Position Surface form Disambiguated ID Type / Status
Subject modENCODE project E459329 entity
Predicate relatedTo P37 FINISHED
Object ENCODE project NE NERFINISHED

How this triple was built (3 steps)

Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.

NER Named-entity recognition gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: ENCODE project | Statement: [modENCODE project, relatedTo, ENCODE project]
NED1 Entity disambiguation (via context triple) gpt-5-mini-2025-08-07
Target entity: ENCODE project
Context triple: [modENCODE project, relatedTo, ENCODE project]
  • A. modENCODE project
    The modENCODE project is a large-scale genomics initiative aimed at comprehensively identifying and annotating functional elements in model organisms such as Drosophila melanogaster and Caenorhabditis elegans.
  • B. Personal Genome Project
    The Personal Genome Project is a pioneering open-science initiative that publicly shares the genomic and health data of volunteers to advance research and understanding of human genetics.
  • C. Human Genome Project
    The Human Genome Project was an international scientific research initiative that successfully mapped and sequenced the entire human DNA genome, revolutionizing genetics and biomedical research.
  • D. Edico Genome
    Edico Genome was a biotechnology company specializing in high-speed genomic data analysis through its DRAGEN bio-IT platform and FPGA-based acceleration technology.
  • E. All of Us Research Program (genomics components)
    The All of Us Research Program (genomics components) is a large-scale U.S. precision medicine initiative that collects and analyzes genomic data from diverse participants to advance research on health, disease, and individualized care.
  • F. None of above. chosen
  • G. Unsure - the case is ambiguous/there is not enough information to decide.
NED2 Entity disambiguation (via description) gpt-5-mini-2025-08-07
Target entity: ENCODE project
Target entity description: The ENCODE project is a large-scale collaborative research initiative aimed at identifying and characterizing all functional elements in the human genome.
  • A. modENCODE project
    The modENCODE project is a large-scale genomics initiative aimed at comprehensively identifying and annotating functional elements in model organisms such as Drosophila melanogaster and Caenorhabditis elegans.
  • B. Personal Genome Project
    The Personal Genome Project is a pioneering open-science initiative that publicly shares the genomic and health data of volunteers to advance research and understanding of human genetics.
  • C. Human Genome Project
    The Human Genome Project was an international scientific research initiative that successfully mapped and sequenced the entire human DNA genome, revolutionizing genetics and biomedical research.
  • D. Edico Genome
    Edico Genome was a biotechnology company specializing in high-speed genomic data analysis through its DRAGEN bio-IT platform and FPGA-based acceleration technology.
  • E. All of Us Research Program (genomics components)
    The All of Us Research Program (genomics components) is a large-scale U.S. precision medicine initiative that collects and analyzes genomic data from diverse participants to advance research on health, disease, and individualized care.
  • F. None of above. chosen

Provenance (2 batches)

The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.

Step Stage Batch ID Status When
creating Elicitation batch_69d8d396f54c8190ba49db31e8743842 completed April 10, 2026, 10:40 a.m.
NER Named-entity recognition batch_69e5977ffa648190be5f682bba47fb0a completed April 20, 2026, 3:03 a.m.
Created at: April 10, 2026, 11:52 a.m.