Historical methods record Version 1.0 / July 2026

Language-model screening.

This record describes the narrow language-model step used to route discrepancies between claimed citation titles and identifier-resolved records in the published CITADEL audit. The model did not determine whether a publication existed, infer author intent, or establish that a reference was fabricated.

§ 01 / Historical implementation

Historical implementation

Model
Claude Haiku (Anthropic), a proprietary large language model
Query period
Feb 16-19, 2026
Configuration
Zero-shot classification through an application programming interface; no fine-tuning or in-context examples
Parameters
Maximum output 300 tokens; temperature and other decoding parameters were not explicitly specified, so provider defaults applied

Records retained after this routing step underwent independent bibliographic searches. Only references remaining unmatched in PubMed, Crossref, OpenAlex, and Google Scholar could meet the study definition of a fabricated reference.

§ 02 / Generic prompt

Generic prompt

This is a generic specification of the title-comparison task as implemented in the published audit. It covers the supplied inputs, the output categories, and the step's nonfinal routing role.

Title-comparison routing task
You are reviewing a discrepancy between a citation title and bibliographic records resolved from identifiers supplied in that citation.

You will receive:
- the claimed citation title;
- the title resolved from the supplied PMID, when available;
- the title resolved from the supplied DOI, when available; and
- limited bibliographic metadata.

Using only the supplied information, assign one routing category:
- "title_variant": the claimed title plausibly represents the same work as an identifier-resolved record;
- "identifier_discrepancy": a supplied identifier appears to resolve to a different work;
- "unresolved_discrepancy": the resolved records do not account for the claimed title and external verification is required; or
- "insufficient_information": the evidence is incomplete or supports more than one interpretation.

This is not a determination that a publication exists or is fabricated. Do not infer author intent or use writing style as evidence.

Return only:
{
  "category": "title_variant | identifier_discrepancy | unresolved_discrepancy | insufficient_information",
  "rationale": "One sentence based only on the supplied bibliographic evidence"
}

*CITADEL is a versioned, continually updated system. Model behavior can change across versions, and model-prompt configurations are reevaluated when updated. This record describes the historical implementation relevant to the published audit and does not report the current production configuration.

§ 03 / Suggested citation

Suggested citation

Topaz M. CITADEL language-model screening record. Version 1.0. 2026. https://www.maxtopaz.com/citadel/methods