What a Harness Can Promise That a Prompt Can't
Dan PetersonSynopsis — AI-drafted from Dan's notes
Luis Villa, a former Wikimedia Foundation lawyer and community lead, answers a question put to him on Bluesky: when an AI agent’s log says it read a source and the source supports the claim, is that true, or just more generated text? His answer is that it depends on the harness, not the model. Ask a chat tool to read a citation and check it, and all you get back is the model’s own account of what it did. You can’t tell whether it fetched the page, reached the real page, or answered from its training data.
A harness built for the job can move the checkable steps out of the model and into ordinary code. First, the program fetches the URL before the model sees anything. A dead link, a paywall or an empty page stops the run, and the log of that step is written by the program, not narrated by the model. Second, instead of asking whether a source supports a claim, the harness asks for the exact passage that does, then searches the fetched text for it. An invented quotation can’t pass, and a human reviewer gets real text to judge.
Villa is open about the costs. Models quietly “correct” the quotes they are asked to copy (fixing typos, straightening curly quotes, dropping words), so the matching has to forgive formatting noise without forgiving rewording. Fetching is messy too, since a page can render differently for a script than for a browser. What survives is observability: logs you can trust for the steps the code runs, and an honest line between what is verified and what is still the model’s judgement. A model can still misjudge whether a real passage supports a claim. Villa closes by reframing the question for Wikipedia: not whether a chatbot can check citations, but which parts of that job can be taken away from the chatbot altogether.