I rushed into discussing a topic without conducting proper research.

Before writing the article of 11 November 2025, I read that TOON is the new way to talk to AI—efficient, fast, and easy to read all in one. Oh boy, I was naive thinking that one parameter is enough to switch completely from my beloved JSON.
Now I feel I cheated, and I hope you'll forgive me. More importantly, I want to show you the data that changed my mind.
Why I Got Excited About TOON
If you haven't already, check my previous blog post where I praised TOON for being an amazing format to save tokens (money). The promise was that using TOON would use 30-60% fewer tokens with the same or better accuracy.
That sounded like a no-brainer. Who wouldn't want that?
The Question That Started My Doubt
I proudly shared my blog post on LinkedIn. Then Till Simon asked a deceptively simple question: "Does the LLM respond with the same quality?"
That question haunted me. So I dug deeper.
What the Data Actually Shows
Here's where things get interesting—and complicated.
The official TOON benchmark snapshot from November 2025:
- TOON: 68.7% accuracy with 4,389 tokens
- JSON: 65.7% accuracy with 7,260 tokens
- Token savings: 39.5%
Looks great, right? TOON wins on both metrics.
But Independent Testing Tells a Different Story:
When improvingagents.com tested TOON on tabular data with GPT-4.1-nano:
- TOON: 47.5% accuracy (ranked 9th out of 12 formats)
- JSON: 52.3% accuracy
- Markdown-KV: 60.7% accuracy (best performer)
In the separate nested-data test with GPT-5-nano:
- TOON: 43.1% accuracy (dead last)
- JSON: 50.3% accuracy
- YAML: 62.1% accuracy
Why Such Different Results?
After analysing both benchmarks, here's what I learned:
1. Dataset and task matter These tests used different datasets and questions. The researchers did not establish which differences caused the gap, so the results do not prove that the official benchmark was designed to favor TOON.
2. Model Performance Varies Wildly
- GPT-5-nano with TOON: 88.6% on official tests, but only 43.1% on nested data
- Claude Haiku with TOON: 50.7%
- Your model choice changes everything
3. Token Savings Are Real, But... In that table test, Markdown-KV had the highest accuracy, but used 52,104 tokens against TOON’s 21,518. This comparison does not establish why the accuracy differed.
What I Should Have Told You
What I would test:
- TOON, compact JSON and other suitable formats on the same inputs
- The exact model, prompt and task used in the application
- Answer quality, tokens and total cost, including retries
Keep JSON where your API contract requires it. These results do not justify a universal rule for GPT-5, Claude, small models or production systems.
What I Learned
Model behavior depends on more than the input format. A single format change affects the entire system in ways we don't fully understand yet. What works in one scenario fails in another.
The real lesson: Test everything yourself. Don't trust the hype—including my original recommendation.
TOON is an interesting specialised tool. My original recommendation was too broad: measure both token use and answer quality on your own task before changing formats.
Sources:

Discussion
Comments
No published comments