Triple
T16430695
| Position | Surface form | Disambiguated ID | Type / Status |
|---|---|---|---|
| Subject | Amazon Web Services Open Data |
E399063
|
entity |
| Predicate | primaryServiceUsed |
P11934
|
FINISHED |
| Object | AWS Public Datasets |
E399063
|
NE FINISHED |
How this triple was built (2 steps)
Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.
NER
Named-entity recognition
gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: AWS Public Datasets | Statement: [Amazon Web Services Open Data, primaryServiceUsed, AWS Public Datasets]
NED1
Entity disambiguation (via context triple)
gpt-5-mini-2025-08-07
Target entity: AWS Public Datasets Context triple: [Amazon Web Services Open Data, primaryServiceUsed, AWS Public Datasets]
-
A.
Open Data Index
Open Data Index is a global initiative that evaluates and ranks the openness and accessibility of government data across countries.
-
B.
Amazon Web Services Open Data
chosen
Amazon Web Services Open Data is a program that hosts and freely provides access to large, publicly available datasets—such as satellite imagery, genomics, and climate data—on AWS cloud infrastructure for research, innovation, and application development.
-
C.
Data.gov portal
The Data.gov portal is the U.S. government’s central online repository for accessing and downloading open data sets from federal agencies.
-
D.
Open Data Lab
Open Data Lab is a World Wide Web Foundation initiative that supports the use of open data to drive social impact, innovation, and better governance, particularly in developing countries.
-
E.
Hugging Face Datasets
Hugging Face Datasets is an open-source library that provides a large collection of ready-to-use datasets and efficient data loading tools for machine learning and natural language processing workflows.
- F. None of above.
- G. Unsure - the case is ambiguous/there is not enough information to decide.
Provenance (3 batches)
The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.
| Step | Stage | Batch ID | Status | When |
|---|---|---|---|---|
| creating | Elicitation | batch_69d87f2b9024819085c20e52de95d583 |
completed | April 10, 2026, 4:40 a.m. |
| NER | Named-entity recognition | batch_69e328fe0f488190ac34aa677c980a20 |
completed | April 18, 2026, 6:47 a.m. |
| NED1 | Entity disambiguation (via context triple) | batch_6a004584fa508190a85b1f79ecf9c258 |
completed | May 10, 2026, 8:44 a.m. |
Created at: April 10, 2026, 5:10 a.m.