Triple
T7392995
| Position | Surface form | Disambiguated ID | Type / Status |
|---|---|---|---|
| Subject | George Church |
E170546
|
entity |
| Predicate | knownFor |
P22
|
FINISHED |
| Object | Personal Genome Project |
E661987
|
NE FINISHED |
How this triple was built (2 steps)
Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.
NER
Named-entity recognition
gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: Personal Genome Project | Statement: [George Church, knownFor, Personal Genome Project]
NED1
Entity disambiguation (via context triple)
gpt-5-mini-2025-08-07
Target entity: Personal Genome Project Context triple: [George Church, knownFor, Personal Genome Project]
-
A.
Personal Genome Project
chosen
The Personal Genome Project is a pioneering open-science initiative that publicly shares the genomic and health data of volunteers to advance research and understanding of human genetics.
-
B.
Human Genome Project
The Human Genome Project was an international scientific research initiative that successfully mapped and sequenced the entire human DNA genome, revolutionizing genetics and biomedical research.
-
C.
All of Us Research Program (genomics components)
The All of Us Research Program (genomics components) is a large-scale U.S. precision medicine initiative that collects and analyzes genomic data from diverse participants to advance research on health, disease, and individualized care.
-
D.
Edico Genome
Edico Genome was a biotechnology company specializing in high-speed genomic data analysis through its DRAGEN bio-IT platform and FPGA-based acceleration technology.
-
E.
Genomic Science program
The Genomic Science program is a U.S. Department of Energy research initiative that applies genomics and systems biology to understand and harness biological processes for energy, environmental, and climate-related applications.
- F. None of above.
- G. Unsure - the case is ambiguous/there is not enough information to decide.
Provenance (3 batches)
The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.
| Step | Stage | Batch ID | Status | When |
|---|---|---|---|---|
| creating | Elicitation | batch_69c68a5e2c9081909e713ce866e0060a |
completed | March 27, 2026, 1:47 p.m. |
| NER | Named-entity recognition | batch_69c6f224790c819099ceb7c7ac8d00f6 |
completed | March 27, 2026, 9:09 p.m. |
| NED1 | Entity disambiguation (via context triple) | batch_69c81ecf65388190a149efc77aedcd91 |
completed | March 28, 2026, 6:32 p.m. |
Created at: March 27, 2026, 3:09 p.m.