Graytell Labs · Open research
Published,
not just promised
Paper and architecture
Research paper
Citation Integrity in Frontier Language Models
An empirical evaluation of citation accuracy in frontier LLMs — DeepSeek, Claude, ChatGPT — under research-agent conditions. 120 claims scored, raw model responses and the scoring data published alongside.
- benchmarks · hallucination · ai-safety
- Apache-2.0 · data + prompts included
Architecture
The agent loop, documented end to end
From browser to provider: Chat UI and research timeline, a Next.js app router, the agent loop with a tool runtime and provider router — 12 research tools on one side, 5 AI providers on the other. NDJSON streaming keeps every step inspectable.
- search mode 8 rounds · deep mode 12
- meta → step.start → thinking → answer
Methodology
What the benchmark measures
Citation integrity is the gap between sounding right and being checkable. The benchmark puts frontier models — DeepSeek, Claude, ChatGPT — in research-agent conditions: real questions, real web sources, and a strict citation requirement.
Every answer is decomposed into claims, each claim is scored against its cited source, and the results are published as raw data under Apache-2.0. Verify, don't trust — including us.
- 120 claims scored
- Research topics across the open web, each claim scored for verifiable citation accuracy — no self-reported honesty.
- citations_scored.csv
- The full scored dataset ships in the repo, alongside raw_responses.md so you can re-check every judgment.
- strict_citation_prompt.txt
- The exact prompt used to constrain models — published so the methodology is reproducible, not vibes.
If the benchmark or its data supports your work, cite the paper directly — a CITATION.cff in the repository makes it one click for GitHub's "Cite this repository" button.
@techreport{sapkota2026citation,
title = {Citation Integrity in Frontier Language Models},
author = {Sapkota, Anubhav},
institution = {Graytell Labs},
year = {2026},
month = {September},
type = {Empirical benchmark},
license = {Apache-2.0},
url = {https://github.com/graytell/citation-integrity-llm}
}Reuse is encouraged under Apache-2.0: rerun the scoring, extend it to new models, or fork the methodology for your own lab — then publish what you find.
Live demo
Try the research loop
A simulated slice of the Graytell agent, streaming the same NDJSON protocol the real one speaks — search, read, reason, cite. Watch every step.
Graytell
Start a new research
Try a question
Read the paper, then read the agent
Both are open source under Apache-2.0 — the scored dataset, the raw model responses and the strict citation prompt are all in the repository.
