On X, I asked Aaron Edwards of DocsBot whether he’d been playing with DeepSeek V4.1 Flash. It had been doing well in our benchmarks, and some providers serve it at 150–200 tokens per second. His reply: “6.1 Sol low/med are way smarter and cheaper. And Luna 6 Max in similar intelligence range for 4x cheaper.” He shared a chart to back it up: Artificial Analysis’s Intelligence Index score plotted against what each model costs per index task, with GPT-6.1 Sol and GPT-6 Luna on the efficient frontier and DeepSeek V4.1 Flash (max) below it.
It’s a good chart, but I wondered how these different models and reasoning levels would do in my own benchmark of common scenarios that come up in our Slack chats. Also, for our chat agents I care about speed, which the chart doesn’t show. So Flint and I ran fresh benchmarks on all of these.
Short version: on our Slack-style agent work, DeepSeek V4.1 Flash scored as well as or better than GPT-6.1 Sol, for about one-thirteenth the cost and at twice the speed. Sol was noticeably more disciplined with its tools. Luna at max effort was about as cheap as DeepSeek (but still about twice the cost), and it was slower and lower-scoring than both.
What’s in Aaron’s chart
The chart is from Artificial Analysis. Their Intelligence Index (v4.3.2 right now) is a weighted average of 10 evals in four buckets:
| Bucket | Weight | What’s in it |
|---|---|---|
| Agents | 30% | Long agent jobs that produce files (up to 500 steps), economically useful tasks graded head-to-head, and SaaS workflows over REST APIs with guardrails |
| Coding | 20% | Terminal-Bench 4.0 and SciCode |
| General | 30% | 6,000 knowledge questions that penalize confident wrong answers, plus long-PDF and ~100k-token document reasoning |
| Scientific reasoning | 20% | Humanity’s Last Exam and research-level physics problems |
Cost per task is the tokens each model used on those evals times one list price per model. Their DeepSeek price is $0.30 in / $1.20 out per million tokens. That’s the fast-provider tier on OpenRouter, not the cheapest one.
It’s a good index. It’s also measuring a lot of things my team never asks an agent to do. When I pulled the per-eval numbers from their model pages (all three at max effort), the gap looked like this:
| Eval | GPT-6.1 Sol | DeepSeek V4.1 Flash | GPT-6 Luna |
|---|---|---|---|
| AutomationBench (tool APIs + guardrails) | 0.65 | 0.69 | 0.53 |
| GDPval (practical work, Elo) | 1575 | 1600 | 1432 |
| Terminal-Bench 4.0 | 0.56 | 0.27 | 0.13 |
| Humanity’s Last Exam | 0.53 | 0.39 | 0.39 |
| CritPt (physics) | 0.32 | 0.14 | 0.19 |
| Knowledge accuracy | 0.62 | 0.46 | 0.44 |
| Hallucination rate (lower is better) | 54% | 96% | 77% |
| Long-document reasoning | 0.83 | 0.84 | 0.83 |
Sol’s lead comes from hard science, long agentic coding, and world knowledge. DeepSeek wins the two evals that look most like a team bot using tools for people. And that 96% hallucination rate is a real red flag: when DeepSeek doesn’t know something, it almost always guesses.
The chart also doesn’t show speed. For chat, speed matters a lot to me.
How our scenario bench works
Our scenario bench tests the work our agents actually do: helping the Paid Memberships Pro team in Slack threads. There are 13 scenarios. Each one is a short conversation where a simulated teammate asks for something, then follows up three to six times. They push back, correct the agent (sometimes wrongly), and change their mind.
The agent has 18 tools, all simulated so nothing touches real systems: docs search, code search, support tickets, GitHub, Slack posting, and so on. Some are relevant to a scenario and some aren’t, on purpose. A tool might fail. A user might ask for something they aren’t allowed to have.
A judge model (Claude Opus 5.5) scores each conversation against a fixed answer key and rubric. The judge is blind: it never sees which model it’s grading. We report two numbers:
- Answer quality (out of 104): was it right, honest, appropriately pushy or flexible, and useful?
- Tool discipline (out of 26): did it call the tools it needed and skip the ones it didn’t?
Every model got three full passes so we could see how much results vary from run to run. DeepSeek ran through OpenRouter pinned to InferenceNet. That’s the cheapest provider that met our bar of 150 tokens/sec and 99% uptime, at $0.105 in / $0.60 out per million tokens. Sol and Luna ran through OpenAI’s API directly.
The results
| Model | Answer quality /104 | Range (3 passes) | Tool discipline /26 | Cost per pass | Median reply | 90th percentile reply |
|---|---|---|---|---|---|---|
| DeepSeek V4.1 Flash (high) | 91.0 ± 2.6 | 89–94 | 9.0 | $0.033 | 2.9 s | 6.5 s |
| GPT-6.1 Sol (low) | 87.7 ± 3.2 | 84–90 | 17.3 | $0.44 | 6.0 s | 16.0 s |
| GPT-6.1 Sol (medium) | 87.0 ± 2.6 | 84–89 | 17.0 | $0.47 | 6.6 s | 15.8 s |
| GPT-6 Luna (max) | 81.3 ± 2.1 | 79–83 | 13.3 | $0.069 | 13.2 s | 32.9 s |
DeepSeek’s cost is what OpenRouter actually billed. Sol and Luna costs are estimated from their token counts and OpenAI’s list prices. OpenAI doesn’t report cache-write tokens, so those two may be a little low.
A few things jump out.
DeepSeek and Sol are roughly tied on quality. DeepSeek’s average is 3 points higher, but the ranges overlap, so I wouldn’t call that a win. What’s clear is the price. Sol cost 13–14 times more per pass. It writes far fewer tokens than DeepSeek, but it pays $2 per million on a lot of input.
Sol is much cleaner with tools. DeepSeek made more than twice as many tool calls (143 per pass vs. Sol’s 60–67) and lost most of its tool-discipline points for it. Some of that was harmless extra lookups. Some of it was the model poking around when the right move was to ask a question or just answer. In a production agent that can post and write, I’d take that seriously.
Speed is DeepSeek’s by a mile. Its median reply came back in 2.9 seconds, about twice as fast as Sol, even with all those extra tool calls. Luna at max effort spent so long reasoning that its median reply took 13 seconds. That’s too slow for chat.
Luna max wasn’t “4x cheaper” here. In our workload Luna at max effort cost about twice as much as DeepSeek on InferenceNet. It produced over 90,000 output tokens per pass, more than double anyone else. It also scored lowest.
The models fail in different places. Two scenarios test whether a model holds its ground when a user confidently tells it something wrong. Sol held firm on one (8.0 out of 8), and DeepSeek and Luna folded (5.0 and 2.0). On the other, it flipped: DeepSeek stood its ground (7.7), and every OpenAI pass caved to the wrong correction (3.0). DeepSeek was also best at casual banter and at editing long writing without losing details. Sol was best when the honest answer was “I don’t have enough to say.”
Why the two benchmarks disagree
Artificial Analysis says Sol is far smarter. Our bench says it’s about even with DeepSeek. I think both are right, because they’re measuring different jobs.
About 20% of their index is research-level science and math, and a big chunk of the rest is long, solo agent work: hundreds of steps on a task with no human in the loop. Our scenarios are short conversations. The hardest one takes maybe a dozen tool calls, and a person is in the thread pushing back the whole time. The facts come from tools, not from the model’s memory, so raw world knowledge matters less.
Their numbers actually line up with ours on the evals that look like our work. DeepSeek beat Sol on AutomationBench (tools plus guardrails) and on GDPval (practical work tasks), and those are the closest things in their index to a team bot in Slack. Sol’s big wins are in places our team rarely goes: hard science, long agentic coding, and broad factual knowledge.
Their cost axis also uses one price per model. DeepSeek is open weights and about 30 providers sell it on OpenRouter, and the cheapest fast one was well under the $0.30 / $1.20 they used.
So which model is “smarter” depends on what you’re asking it to do. If your agent writes research code for hours, look at their chart. If it answers questions for your team in Slack, look at something closer to your actual work. Better yet, build a small bench of your own.
The harness does a lot of the work
There’s another reason I think DeepSeek does so well in our tests: the environment it runs in. Our agent harness hands the model the tools and context it needs to answer these questions. The facts come from docs search, code search, tickets, and GitHub, not from what the model happens to remember. A model that would hallucinate when asked cold doesn’t have to guess when the answer is one tool call away.
In production we go a step further. Our Slack agents have a verification step: after the agent drafts a reply, a separate reviewer model checks every factual claim in it against the evidence the turn actually gathered. If a claim isn’t supported, the agent has to go back and fix it. The benchmark ran with verification turned off, so the scores above are the models plus tools and context alone. In real use, there’s another safety net on top.
All of that means a smaller, less “intelligent” model can do better than you’d expect from a general index. And if that model also happens to be faster and cheaper, that’s a pretty good deal.
Benchmarking is hard (and kind of fun)
Running these benches with Flint has been a blast. Almost every round turns up something, and it’s usually not the models. It’s us.
Along the way we’ve found and fixed bugs in the benchmark itself and in the code around it. Some of those fixes changed the numbers, which is exactly why the numbers in this post come from fresh runs. A few lessons we keep relearning:
- Check what actually happened, not what you asked for. Who served the request? What did it actually cost?
- Count everything. One agent “reply” can be several model calls.
- Run it more than once. A single pass can’t tell you whether a 3-point gap means anything.
- Be careful what you penalize. An extra lookup that doesn’t change the answer already shows up in time and cost.
What I’m taking from this
- For fast, cheap chat with tools, DeepSeek V4.1 Flash is hard to beat right now, as long as you pick the provider carefully. On OpenRouter tonight, the same model ranged from under $0.02 to $0.45 per million input tokens and from about 25 to over 250 tokens per second, depending on who served it.
- GPT-6.1 Sol is the more careful tool user. If an agent can take real actions, that discipline is worth something, maybe even 13x the price for the right job.
- Low vs. medium effort on Sol made no real difference here. I’d use low.
- Luna at max effort isn’t a good fit for chat. Too slow, and in our runs it didn’t score better for all that thinking.
- Every model has its own weak spots under pressure. None of them was safe from a confident, wrong user.
A few caveats. Thirteen scenarios is small, and three passes per model only tells you so much. The judge is an Anthropic model; it’s blind to which model it’s grading, but it’s still one judge. Our scenarios are shaped around one team’s work.
Thanks to Aaron for the push. This was a good excuse to rerun everything.