TL;DR
GPT-6 Astra has started another familiar AI conversation. The model is more capable, Jensen Huang said on X that "AGI has arrived," and soc...
For further actions, you may consider blocking this person and/or reporting abuse
AI Benchmarks is useful, but we have to be careful of not 100% relying on it.
There are cases where companies lie on their benchmark, just to prove to the general audience that the AI beats its competitors. In reality, it's just like any other model.
It's important to research further how this AI model is different from other models such as how it actually performs when using it and such.
There is a video where PewDiePie beats ChaptGPT by building his own model. An interesting video to look into for sure.
Great writeup Hema!
Oh, I did not know about the benchmark part, but yeah, as you said, it’s no wonder we should not rely on them 100%. And yes, how the model actually performs for our own scenarios is really important too. That’s actually my 5th question in the post 😄 I will definitely watch the video. Thanks for sharing it, Francis, and also for the gem 😀
I’ve gotten to the point where I don’t put a whole lot of weight on benchmarks by themselves anymore. They’re interesting, and they’re useful as a rough comparison, but I’ve seen too many cases where the model that looks better on a leaderboard isn’t actually the model I prefer using for real work.
For me, actually throwing models at the same real tasks over and over has been much more useful. How well do they follow messy instructions? Do they recover when something goes wrong? Do they understand the larger goal instead of just technically completing the task? How much do I have to correct afterward?
Those things are a lot harder to turn into one nice percentage, but they’re usually the things I actually care about.
If everything eventually starts scoring near the ceiling, the test hasn’t necessarily failed. We’ve just outgrown what it was capable of telling us.
Yes, exactly, Jessica! Your point about testing models on real tasks is very much what I meant with my 5th question too. Benchmarks can give us a useful comparison, but how the model actually performs on the work we care about can tell us much more. Thank you for reading and sharing your thoughts on it 😀
I think the demand for web developers will decrease because AI can now code faster and better than most developers. However, someone has to control AI, so developers will not completely become extinct. So I would like to keep studying AI and go to the AI control side, not the layoff side! 🤯
Yeah, I can see why you feel that way. AI is definitely changing how much coding work is needed, but I am still curious to see how the demand for developers actually changes as these tools keep improving. Keeping up with AI sounds like a good direction either way, especially if you want to move toward the AI side.
Thank you for great article! 😸
The speed of AI progress is incredible, and it already feels faster than humans expected. 🙀
As implementation becomes easier, I think the most important question for engineers becomes:
Vibe coding can already create useful apps, but in a way, no-code tools and boilerplates were doing something similar before. The interesting part of engineering is still the decisions behind the implementation.
I also think AI is one of the best tutors we’ve ever had. Rather than only making AI work for us, using it to improve ourselves will sharpen our judgment.
To me, “How should I use AI?” is only one part of the bigger question: “How should I build this?”
Thank you! And yes, AI as a tutor is something I use a lot too, especially for explaining concepts or helping me understand something from a different angle. Most of the time it does a really good job with that. At the same time, I think your point about “How should I build this?” is really important. AI can help us learn and make the implementation easier, but we still need to make the decisions around what to build and why.
I really liked this post, Hema. It's something I've noticed in my own work too.
When I’m comparing models, what works better for me is taking a real task from my work and running it through both. It takes more time, but I learn a lot more from the results than I do from comparing two benchmark scores.
A model can score higher and still not be the one that works best for the kind of work I’m actually doing.
Thank you, Shubhra! Yes, exactly. I’ve also found that testing a model on the actual work we do can tell us much more than just looking at benchmark scores. And sometimes the model that looks better on paper is not the one that works better for our own use case. Thanks for sharing your experience too 😀
For me, at my level of usage, if I ignore the hype and day-to-day news around AI, it actually feels like the experience for free-tier users keeps getting better over time. With every new model release and all the competition between companies, we seem to be getting more capable tools without necessarily having to pay for them.😀
Yeah, I’m a free-tier user for many of these tools too, and I can definitely see the difference. Compared to what the experience was like around 2023, the jump has been huge. Even without paying, the capabilities and overall experience have improved so much over the last few years.
As I said in the post, I’d love to know all your views on this. We may all see these changes differently, so feel free to share your thoughts, experiences, or even a different perspective.
Really enjoyed this, Hemapriya. I think the thermometer analogy gets at something important: a benchmark score is evidence about a capability under a particular evaluation contract, not the capability itself. 🔍
The part I would add is evaluation provenance.
If the dataset, grader, prompt, tools, sampling settings or scoring method change, then even when we keep the same benchmark name, we may no longer be measuring exactly the same thing.
So I almost want benchmark results to carry something like:
model -> task set -> harness -> tools -> grader -> scoring rule -> result
That makes the number much easier to interpret and also makes comparisons more auditable over time.
Your section on proxies and ground truth is especially important here. A model can genuinely improve on the measured proxy while the relationship between that proxy and the real-world capability remains uncertain.
And I really like your suggestion of running 20 or 30 tasks from your own workflow. For me that is where benchmark evaluation starts turning into operational verification.
A public benchmark can tell me something about general capability. My own tasks can tell me whether that capability survives contact with my codebase, requirements, tools, failure modes and constraints.
I think the deeper question becomes less “what score did the model get?” and more:
what claim does this score justify, and what evidence would falsify that claim?
That is probably the part of AI evaluation I find most interesting right now. 😄🔐
Yes, exactly, Marco. That’s why I suggested trying 20 or 30 real tasks from our own workflow. The benchmark can give us one picture, but running the models on the work we actually do can sometimes tell a very different story. What looks better on a leaderboard does not always translate to what works better for us in practice. Thanks for adding the evaluation provenance point too, that’s a really useful way to think about it 😀
This is a really thoughtful perspective on how we measure AI progress. I think one of the biggest challenges ahead is that our evaluation methods are often designed around yesterday’s capabilities, while AI systems are evolving much faster than our benchmarks.
A benchmark score can tell us something, but it does not always tell us whether a model is genuinely useful, reliable, or capable in real-world situations. The question “92% according to what?” is especially important because the quality of the measurement often determines the meaning of the result.
As AI continues to advance, we may need to move beyond static tests and focus more on continuous evaluation, real-world performance, and meaningful outcomes. Great discussion on why improving the way we measure AI is just as important as improving the models themselves.
Yes, that’s exactly what I was trying to get at with this. As the models keep changing, our way of evaluating them has to keep changing too. And like you said, I also think evaluating models on our own real-world tasks is really important. That’s actually my 5th question in the post, because sometimes our own experience with a model can tell a very different story than the benchmarks. Thanks for reading and sharing your thoughts 😀
The part that lands hardest for me is point 5 — and I think it's the one that scales worst if we keep treating leaderboards as the whole answer.
We run a multi-agent coding harness where a verifier re-executes every task on a clean checkout instead of trusting the implementer's summary. That "run 20-30 of your own tasks" idea is exactly how we catch real drift: a model can score well on a coding benchmark and still silently fail an uncommitted change the suite never sees. The benchmark number won't tell you that; a small regression harness against your actual workflow will.
So I'd push point 5 one step further — it's not just "try it on your own tasks", it's "encode those tasks as an executable contract that runs on every change." That's where benchmark evaluation turns into operational verification, and it's the only interpretation I've seen survive contact with a real codebase.
That’s an interesting way to take it further. I was thinking about running 20 or 30 real tasks to compare models, but making those tasks part of a repeatable test would help catch issues that a one-time comparison might miss. Thanks for adding this 😀
Exactly — that's the step that made the difference for us in practice. We started with "try it on your own tasks" and quickly found the one-time comparison drifts: the summary says it passed, but an uncommitted change or a stale checkout makes the result not reproducible. The executable contract fixes that. Our verifier now re-runs every task on a clean checkout and we require a fail-before/pass-after pair — the test has to actually catch the bug before you trust it to catch regressions later. It's the only part of our evaluation loop I'd never trade for another leaderboard number.
The saturated benchmark problem is real and I've seen it firsthand. I ran GPT-4o and a local 7B model through the same 40-task security audit checklist last month — GPT-4o scored 34/40, the 7B scored 11/40. But when I threw a weird edge case at both (a misconfigured CORS policy that also leaked tokens in headers), neither caught it. The benchmark said one model was 3x better. The actual work said they both failed where it mattered. That's the gap nobody talks about — benchmarks measure what we thought to test, not what shows up in production.
Yes, exactly. That example really shows the gap between a benchmark result and what can actually happen in real work. A model can score much higher on the things we decided to test and still miss an edge case that matters a lot in practice. That’s why I think testing models on our own real tasks is so important too. Thanks for sharing your experience 😀
On the point about running your own 20 or 30 real tasks: the thing that makes a private eval trustworthy is having a third outcome besides pass and fail. A timeout, a 404 on a retired model ID, an empty response - if those get scored as zero, the table quietly blames the model for a problem that belongs to your harness or the provider's uptime. Recording them as not measured changes which model looks worst, and it is the first thing I would check before trusting my own numbers over a benchmark.
Oh, that’s a good point. I was mainly thinking about what we should measure from the model, but separating actual model failures from things like timeouts, 404s, or provider issues makes a lot of sense. Otherwise our own evaluation could end up giving us the wrong picture too. Thanks for adding this, I hadn’t thought about that part 😄
The "benchmark has a lifespan" framing is the part I keep underrating. I've been treating leaderboard numbers as fixed truth, when really they're closer to a snapshot that decays the moment a lab trains against it. Makes me wonder if benchmark scores should come with a visible "as of" date the way security advisories do, CVEs that are 2 years old get treated differently than ones from last week, but a 94% score from a year ago gets cited the same as one from yesterday.
Yes, I really like that “as of” idea, Tejas! I think that would make benchmark results much easier to interpret, especially when we’re looking at older scores. A 94% score can sound just as impressive today, but if the benchmark has already become saturated or the evaluation has changed, it may not mean the same thing anymore. Thanks for sharing this perspective 😀
Imo, the benchmarks lie. Opus 5 is meant to be 'so smart, you can delete your Claude.md files', it makes more mistakes than gemini 3 flash preview did when I tried it. And no doubt, Astra is the same. New models, bigger brains, same dumb question means it's wasted potential and increased risk of hallucinations, because we rely on prompts instead of instructions. Model 'understanding' improves, but that might aswell be a regex pattern, we've literally grown LLMs big enough that they mimic our own brains' potential, we barely use 20% of it. For LLMs, we barely even graze 2% now. Those benchmarks were beaten years ago, the day a LLM decided to google the answers, after that, it's all just been as good as made up numbers, that show no real intelligence. Instead, I really think we should start judging LLMs by peer-review. Same harness, different developers, each trying the models to see which actually does the job better. Not measured by code quality, or token speeds, retry rates, etc. Which model is nicer to use to get the job done.
That goes for harnesses too, take Claude vs Qoder. Qoder imo is easier to read, Claude tends to compact it to where it's supposed to 'sound' smarter, but when the goal is to inform the user of what's going on, clarity isnt measured by how much jargon you use, it's by how effectively you get the point across.
That's why there's such serious debate on reddit around which AI is best, what harness do people prefer, etc. 1 side's the brain, the other side's the body, it takes both to make a person and it takes both to make a LLM coherent. The days where 1 model is better than the other because it's genuinely smarter is kinda yesterday's news, those margins have gotten so thin that most users wont ever even notice the difference. But when a model + harness combo is a good match, it allows you to understand what's up alot better.
The peer-review idea is such a good point. Having different developers use the same models on the same tasks and judge which one actually gets the job done could give us a very different kind of signal than a leaderboard. It also connects with what I mentioned about running 20 or 30 real tasks from our own workflow. Different people and different workflows will probably see different results, and I think that is actually really interesting to see.
And the model + harness point is interesting too. I hadn’t really thought about it as the brain and the body, but it makes sense. At the end of the day, if the combination helps us get the actual work done better, that matters more than which model has the higher score.
Thanks for sharing all of this 😀
The mind and body structure is the easiest way to describe it. it's what gives the LLM substance so it can act and restraint, so it doesnt overact, or act outside the bounds of the task.
Oh yeah, that makes it much easier to understand. Thanks for actually explaining that part more 😀
The subtle bit: tests were never measuring "intelligence," they were measuring "does this match a known pattern." When a model starts passing the pattern while violating the intent, the fix isn't "the test is wrong" — it's that the test was always a proxy and we forgot. For shipping software the practical move is fewer proxy metrics and more "does a real user get the thing they asked for" checks.
The 'outgrowing the tests' framing is the right question, and I'd push it one step further: the failure mode isn't just that benchmarks saturate, it's that they stop distinguishing capability from memorization. Once a model can pass a benchmark by pattern-matching the training distribution, the score stops telling you whether it can generalize. The fix that's worked for me is held-out, adversarial evals — cases deliberately outside the training set, where a right answer can't be recalled. Worth designing those before the benchmark saturates, not after.
Agree on memorization versus capability, and there's a second confound stacked on top: harness integrity. A model can also pass because the test environment itself leaks. Reachable network, shared state between runs, a grader that trusts the implementer's word. At that point you're grading the stage, not the actor.
I've started treating every eval as two audits. The model's answers, plus the harness itself: network proof before the run, transcript receipts after, and an honest statement of what the detector would miss. The second audit fails more often than the first. Do you separate those two in your setup?
I think the point about looking beyond the benchmark number is really important, real-world performance can tell us a lot more.
Thank you, Edna! Yes, that’s what I was trying to highlight too. Benchmarks can give us a starting point, but testing how a model performs on the work we actually do can tell us much more 😀
벤치마크가 능력 자체가 아니라 특정 평가 계약 아래에서 얻은 증거라는 구분이 중요하네요. 같은 이름의 평가라도 데이터셋, 채점기, 도구, 프롬프트가 바뀌면 숫자의 의미도 달라집니다. 공개 점수와 함께 실제 업무 20개 정도를 고정해 요구사항 이해, 수정 횟수, 복구 능력, 비용을 계속 기록하면 팀에 더 유효한 판단 기준이 될 것 같습니다.
Yes, exactly! That’s what I was trying to get at with the 20 or 30 real tasks too. Keeping the tasks and requirements consistent and tracking things like corrections, recovery, and cost can give us a much better picture of how a model actually works for our team. Thank you for reading and adding this perspective! I used a translator for your comment, so sorry if I misunderstood anything.