Insights3 min read255 views

I Made a Mistake About TOON (And Here's the Data That Proved Me Wrong)

I correct my TOON recommendation: token savings need to be evaluated alongside accuracy, data structure and the model being used.

Written byIdir Ouhab

I rushed into discussing a topic without conducting proper research.

Cover image for I Made a Mistake About TOON (And Here's the Data That Proved Me Wrong)

Before writing the article of 11 November 2025, I read that TOON is the new way to talk to AI—efficient, fast, and easy to read all in one. Oh boy, I was naive thinking that one parameter is enough to switch completely from my beloved JSON.

Now I feel I cheated, and I hope you'll forgive me. More importantly, I want to show you the data that changed my mind.

Why I Got Excited About TOON

If you haven't already, check my previous blog post where I praised TOON for being an amazing format to save tokens (money). The promise was that using TOON would use 30-60% fewer tokens with the same or better accuracy.

That sounded like a no-brainer. Who wouldn't want that?

The Question That Started My Doubt

I proudly shared my blog post on LinkedIn. Then Till Simon asked a deceptively simple question: "Does the LLM respond with the same quality?"

That question haunted me. So I dug deeper.

What the Data Actually Shows

Here's where things get interesting—and complicated.

The official TOON benchmark snapshot from November 2025:

  • TOON: 68.7% accuracy with 4,389 tokens
  • JSON: 65.7% accuracy with 7,260 tokens
  • Token savings: 39.5%

Looks great, right? TOON wins on both metrics.

But Independent Testing Tells a Different Story:

When improvingagents.com tested TOON on tabular data with GPT-4.1-nano:

  • TOON: 47.5% accuracy (ranked 9th out of 12 formats)
  • JSON: 52.3% accuracy
  • Markdown-KV: 60.7% accuracy (best performer)

In the separate nested-data test with GPT-5-nano:

  • TOON: 43.1% accuracy (dead last)
  • JSON: 50.3% accuracy
  • YAML: 62.1% accuracy

Why Such Different Results?

After analysing both benchmarks, here's what I learned:

1. Dataset and task matter These tests used different datasets and questions. The researchers did not establish which differences caused the gap, so the results do not prove that the official benchmark was designed to favor TOON.

2. Model Performance Varies Wildly

  • GPT-5-nano with TOON: 88.6% on official tests, but only 43.1% on nested data
  • Claude Haiku with TOON: 50.7%
  • Your model choice changes everything

3. Token Savings Are Real, But... In that table test, Markdown-KV had the highest accuracy, but used 52,104 tokens against TOON’s 21,518. This comparison does not establish why the accuracy differed.

What I Should Have Told You

What I would test:

  • TOON, compact JSON and other suitable formats on the same inputs
  • The exact model, prompt and task used in the application
  • Answer quality, tokens and total cost, including retries

Keep JSON where your API contract requires it. These results do not justify a universal rule for GPT-5, Claude, small models or production systems.

What I Learned

Model behavior depends on more than the input format. A single format change affects the entire system in ways we don't fully understand yet. What works in one scenario fails in another.

The real lesson: Test everything yourself. Don't trust the hype—including my original recommendation.

TOON is an interesting specialised tool. My original recommendation was too broad: measure both token use and answer quality on your own task before changing formats.


Sources:

About the author

Idir Ouhab

AI Deployment Engineer at OpenAI, trainer and host of Prompt&Play. I write about what I learn taking AI into production.

Your next step

Is your team working on something similar?

Find out in a minute where your project stands, from prompt to production, and what your next step would be.

Share

Topics

  • toon
  • data analysis
  • correction
  • insights
  • mistake
  • learning
  • evidence
  • research
  • revisit
Always active

Remembers your language and cookie choices. A separate session cookie keeps administrators signed in. These are not used for advertising.

Your choices are valid for 180 days in this browser. Optional purposes start switched off. If browser storage is unavailable, your choice lasts for this page only.

You can withdraw permission here at any time. If optional content has already loaded, the page reloads to stop it; unsent form changes may be lost.

How cookies and storage are used