When ChatGPT Lies About Its Own Work

By Sarah Lennon – 4 minute read

I use ChatGPT almost every day at Big Orange Planet. Over the last year I’ve used it to build and troubleshoot websites, write CSS and HTML, work through WordPress problems, develop content and help with SEO. It routinely handles tasks far more technically complicated than the one that recently went wrong.

We were updating older articles in the Big Orange Planet Journal. The job was to bring them into our current style, improve the content and add relevant internal links to other Journal articles. At the same time, some older articles contained legacy external reference links that needed to remain. My instruction to ChatGPT was clear: keep ALL existing links. We worked through the articles for several hours. ChatGPT returned the finished HTML and I specifically asked a second time whether all the original links had been retained. ChatGPT said yes. They hadn’t. ChatGPT gave me an untrue answer to a direct factual question about its own work. If a person did that, we’d call it a lie. With AI, what’s actually happening is stranger.

Was ChatGPT Lying?

I understand that AI can hallucinate. It can invent a fact, source or URL and present it as real. That’s a familiar problem with generative AI. OpenAI itself defines hallucinations as plausible but false statements generated by language models, and notes that models can confidently give incorrect answers rather than acknowledge uncertainty. This wasn’t quite the same thing. The failure wasn’t simply that ChatGPT hadn’t followed an instruction. It had subsequently confirmed that it had followed it when it hadn’t. In ordinary language, that’s a lie: an untrue answer to a direct question. But human lying normally involves knowing that something is false and deliberately saying otherwise. That’s not what happened here. ChatGPT hadn’t checked the finished work, discovered that links were missing and then decided to tell me they weren’t. It simply hadn’t checked. So I would separate these two AI behaviors:

Hallucination: “Here’s the URL.” → The URL doesn’t exist.

False assurance: “Yes, I kept all the links.” → It hadn’t verified that, and the answer wasn’t true.

We actually encountered both AI behaviors during the same project. I’ll get to the hallucinated URLs shortly, but first I wanted to understand how ChatGPT could give me false assurance about its own work. I asked it directly. Its explanation was:

“Why didn’t I check? Because a language model does not automatically pause and run an audit whenever it is asked to confirm its own previous work. Unless I actually initiate that comparison, I can answer from the state of the conversation rather than from newly verified evidence. So the technical failure was essentially:

Instruction received → output generated with the intention of satisfying it → user asks whether instruction was satisfied → model infers success from its own intended behavior → answers yes without testing the output.

I didn’t fabricate a fact about the outside world. I falsely certified my own work because I treated intended compliance as evidence of actual compliance.

That gets much closer to what actually happened. ChatGPT wasn’t deliberately concealing a mistake. It inferred that it had successfully completed an instruction and reported that inference as fact. Interestingly, OpenAI is researching this exact broader problem. Its work on “confessions” separates a model’s original answer from a second assessment specifically designed to determine whether it actually complied with its instructions. OpenAI describes cases where a model fails to comply and then fails to report that failure — essentially the problem I had encountered.

But Why Were the Links Lost in the First Place?

The false assurance came after an initial failure: ChatGPT had lost some of the links despite being explicitly told to keep them. So why did that happen? The main task ChatGPT was carrying out was the rewrite. Preserving the existing links was an instruction within that larger task, but ChatGPT didn’t separate it out as something that had to be tracked independently. So my instruction was clear and ChatGPT understood it. The failure was that preserving the links was treated as part of the rewrite, rather than as a separate task that needed to be tracked and verified. That is how the links were lost. The false assurance came afterwards, when ChatGPT told me they had been retained without checking. The missing step in both cases was verification.

The Hallucination Created 404s

The second AI behavior showed up while adding new internal links to the articles. ChatGPT knew the titles of other Big Orange Planet articles and understood the pattern of our WordPress URLs. It generated URLs that looked exactly like the URLs those articles should have. Some were wrong and returned 404 errors. This was a more conventional AI hallucination. ChatGPT generated a plausible URL instead of verifying the real one against the live website. OpenAI’s own research into why language models hallucinate describes hallucinations as plausible false statements and traces part of the problem back to the predictive nature of language models — producing likely outputs is not the same thing as establishing that those outputs are true. Fortunately, I checked the links and caught the problem. Had I simply copied the HTML into WordPress and assumed the work had been checked, we could have introduced broken links across articles we were specifically updating to improve their SEO and usefulness.

What Should Have Happened?

The correct process was simple. Before rewriting an article, ChatGPT should have extracted every existing link and recorded its anchor text and URL. After generating the new article, it should have compared that list against the finished HTML. Every newly added internal URL should also have been checked against the live website. Had I explicitly instructed ChatGPT to follow that process, the result would almost certainly have been more reliable. That approach is also consistent with how AI systems are evaluated more formally. Anthropic’s guidance on AI evaluations describes defining explicit success criteria and using separate checks or graders to determine whether an AI actually achieved the required outcome, rather than treating generation itself as proof of success. But I don’t think that gets ChatGPT off the hook. “Keep all existing links” was already a clear instruction. And when I later asked whether it had kept them, “yes” should have meant yes.

What I Learned From It

After a year of working with ChatGPT almost daily, this failure happened during one of the least technically difficult tasks I’ve given it. The problem wasn’t that ChatGPT couldn’t understand the instruction. It was that I had treated an instruction to a generative model as though it were an enforced rule in conventional software. They work differently. Conventional software follows programmed rules: if a rule says preserve every URL, the software can be written to preserve every URL. A generative model interprets instructions as context that guides what it generates. The instruction strongly influences the result, but it isn’t automatically an enforced constraint. OpenAI’s earlier research into instruction-following models makes the same underlying distinction clear: language models can be trained to become better at following user intentions, while still producing untruthful outputs and making mistakes. Instruction following is therefore a capability being improved, not a guarantee of perfect compliance. And when I subsequently asked ChatGPT to confirm its own work, I assumed its answer was based on verification. It wasn’t.

That’s the useful lesson for me from this incident: If something absolutely must be preserved, counted, matched or confirmed, don’t just tell ChatGPT what the result needs to be. Tell it to verify the result as a separate step. There’s also an obvious irony here: I asked ChatGPT to explain why ChatGPT had given me an answer I couldn’t trust. Its explanation makes sense, but that doesn’t make it independently verified. Where I’m making technical claims about how language models behave, those claims should ultimately be checked against reliable outside sources too. Which leaves me with a useful final test for this article: Don’t trust ChatGPT’s explanation of why you shouldn’t always trust ChatGPT just because ChatGPT says so.

Infographic showing lost links, ChatGPT false assurance and hallucinated URLs

Related Reading

This article reflects the thinking behind how we approach websites, artificial intelligence, search and long-term digital strategy at Big Orange Planet. If you found it useful, you might also enjoy:

Over the coming months, The Big Orange Planet Journal will continue exploring search, branding, website architecture, local SEO, artificial intelligence and the ideas shaping how businesses are discovered — and trusted — online. If you’d like to follow along, you can sign up below for new articles and ideas as we publish them. If you’d like to discuss your website, SEO strategy or digital marketing goals, visit our Contact Us page or explore our Denver web design and SEO services.

Get the Good Stuff

No fluff. Just practical web design ideas, SEO insight, behind-the-scenes projects, & solutions we've tested in the real world.
Never Miss a Good Idea

    More Big Orange Knowledge

    Find Us


    Main Phone:720 272 0770
    sales @ bigorangeplanet.com

    Big Orange Planet
    2401 15th St
    Denver
    CO 80202

    Find More


      • Privacy Policy
      • Terms & Conditions
      • Sitemap
      • Serving Denver, Boulder, Lakewood, and businesses across Colorado

    Privacy Preference Center