Triple
T18784072
| Position | Surface form | Disambiguated ID | Type / Status |
|---|---|---|---|
| Subject | modENCODE project |
E459329
|
entity |
| Predicate | fullName |
P16
|
FINISHED |
| Object | Model Organism Encyclopedia of DNA Elements project |
—
|
NE NERFINISHED |
How this triple was built (2 steps)
Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.
NER
Named-entity recognition
gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: Model Organism Encyclopedia of DNA Elements project | Statement: [modENCODE project, fullName, Model Organism Encyclopedia of DNA Elements project]
NED1
Entity disambiguation (via context triple)
gpt-5-mini-2025-08-07
Target entity: Model Organism Encyclopedia of DNA Elements project Context triple: [modENCODE project, fullName, Model Organism Encyclopedia of DNA Elements project]
-
A.
modENCODE project
chosen
The modENCODE project is a large-scale genomics initiative aimed at comprehensively identifying and annotating functional elements in model organisms such as Drosophila melanogaster and Caenorhabditis elegans.
-
B.
Edico Genome
Edico Genome was a biotechnology company specializing in high-speed genomic data analysis through its DRAGEN bio-IT platform and FPGA-based acceleration technology.
-
C.
Ensembl
Ensembl is a comprehensive genome annotation and browsing platform that provides detailed, regularly updated genomic data for a wide range of species.
-
D.
Personal Genome Project
The Personal Genome Project is a pioneering open-science initiative that publicly shares the genomic and health data of volunteers to advance research and understanding of human genetics.
-
E.
Institute for Integrative Genome Biology
The Institute for Integrative Genome Biology is a research center focused on applying genomic, bioinformatic, and systems biology approaches to understand complex biological processes across diverse organisms.
- F. None of above.
- G. Unsure - the case is ambiguous/there is not enough information to decide.
Provenance (2 batches)
The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.
| Step | Stage | Batch ID | Status | When |
|---|---|---|---|---|
| creating | Elicitation | batch_69d8d396f54c8190ba49db31e8743842 |
completed | April 10, 2026, 10:40 a.m. |
| NER | Named-entity recognition | batch_69e5977ffa648190be5f682bba47fb0a |
completed | April 20, 2026, 3:03 a.m. |
Created at: April 10, 2026, 11:52 a.m.