Hi Guys!!! As you know I wasn't good for well, a week and Now.....
Let's Dive In!!!
I never consciously decided which AI gets which job. Somehow...
For further actions, you may consider blocking this person and/or reporting abuse
Okay, now Iβm genuinely curious π
I clearly have a whole AI hierarchy that I never consciously created.
Do you have one too?
Which AI gets the serious work, which one gets the dumb little questions, and which one do you trust for one oddly specific thing for absolutely no logical reason? π
I just bounce around different AI models if anything. I don't really have a specific one that is task on a specific thing. It's more of using it as a good search tool more than writing code since it's very useful of finding information on what I need.
I tend not to use it as much of relying on the AI to do coding because it removes my ability to think critically whenever I am solving a problem. A good methodology is not having to do too much and too little of anything.
Yeah yeah cool. Neither do we since you know start of campus placement season, we don't use AI for coding but you know we do it for hacks as students, cause there the goal is to be generally good.
But thanks for the detailed write-up. But Sir, I'd like to know where do you get this great gifs and images from? Just Curiousπ
No problem! For the images/Gifs, I just search them on the internet (Google specifically). Nothing really fancy though :)
I have noticed a pattern for my workflow. It is roughly:
Claud CLI/code until I run out of credits, because it's the only one im paying for. It has historically been able to crunch some tough problems for me as well as turn data into nice graphs and charts right in the chat for me, including full static diagrams with an interactive/click able overlays with fairly good precision.
Github copilot CLI is nearly neck and neck with claude. When claude becomes frustrating, I have found copilot CLI picks on the slack in an unquestioning, no nonsense way.
Gemini excellent tremendously at video and image generation for supporting documents.
Ollama: Child has taken over my personal computer. Protect fancy requests at all costs. π€£β¨οΈπ¦. Ollama is also great for locally structure workflows that are repetitive like testing. I have used it extensively and am extremely thankful for the open source models for learning.
Iβm a much simpler guy than that; I donβt use an IDE by defaultβI do my programming using CLI-based tools. Since I have company subscriptions to ChatGPT and Copilot, I use one of those. If I really need an editor, I use Vim. For the corporate Confluence, thereβs a tool called Roboto that I use to generate key company documents. I also subscribe to Gemini, but mostly just for the Google storage space; I do use it to make 10-second videos occasionallyβmaybe one or two a week. Of course, lately I havenβt felt the urge to program as a hobbyβwhy burn energy on AI? Maybe in the winter, Iβll pick up a long-cherished project again.
I think you really missed one big LLM i.e the real, natural human brainπ
Unfortunately, the folks who are heavily using or leveraging AI tools are outsourcing their thinking and some even forget they are alive π€£
TBH I'm no longer surprised from this
Recently saw someone online explaining that they created an ai in their brain so when they aren't in situation to open an ai platform, they ask to ai in brain, what to do? And I'm like bro? Did you just learn to think?π€£
I relate to this so much. I definitely have a βserious workβ AI and a βjust vibingβ AI without ever consciously deciding to set it up that way.
Claude Opus is usually my serious work model, especially when Iβm planning something complicated, debugging a weird issue, or working on something where I really care about the reasoning and details. Then I use Antigravity with Gemini 3.8 Flash for smaller changes, experiments, quick ideas, or when Iβm basically just poking at something to see what happens.
I also use Codex with GPT-5.6 Sol for some oddly specific things, especially documentation and little repetitive fixes that come up often. Codex has the best memory for that kind of stuff in my experience, so if I run into the same issue again, I can usually just ask it to fix it and it already knows exactly what I mean. That makes it ridiculously useful for those recurring tasks.
And I think you nailed the reason why this happens. Itβs not always about which model is technically βbest.β You start learning which one you trust for certain kinds of work, which one is fast enough, which one has enough usage available, and honestly which one annoys you the least for a particular task.
At some point you realize youβve accidentally become your own AI router.
Great! Its just that I'm a student and I have no premium subscription so I don't get to try a lot of stuff but I use antigravity cause its free for studentsπ
I have Antigravity for students too! I have to pay for Claude and Codex though if I want to use those.
Love this post. My hierarchy is pretty simple: Claude for serious work and personal projects, Gemini for playing around, and a rotating cast of other AIs for experiments that never quite graduate to real work.
The trust thing you described is exactly it. It's not about which one is best overall, it's which one I'd let near something that actually matters.
Currently trying antigravity cli. Not for serious work but for boring repetitive tasks! π
Gemini into antigravity IDE or the windows app... There is a difference where the IDE of course feel more like a complete final product over a web app transfert to a Windows app for Exemple...The IDE is very good... And of course the pro plan of Gemini is one of the best for tokens... I use also Ollama deepseek 4.1 flash on a 100$ plan give you for 300$ credits so far I didn't even using half... Yertarday... Ollama or DS was getting a higher price per tokens usage.. They have reset my monthly usage back from 0 to 300 for no additional cost... For this use it's a win win.. Using it into Hermes desktop... Is a good duoo for casual and heavy work!!
Yeah, but can't you use the NVIDIA NIM API and run that free claude code server locally? Or is it just good enough for prototyping and not in production? Cause my friends use these models like this way only.
Why do you want to get stuck in a tool that an agent does not understand or even know where it lives?
Remember that a lot of these tools came from CLI or sandboxes; using native tools for native agents is much more controllable.
I have been using, for example, a lot of new desktop apps where the agents finally told me, "I am in Hermes CLI," or "Codex," or "Claude Code."
Remember that, and this is always something I build my tools on that even Claude Code desktop did not really know about and cannot control the application you are using. Ask Claude Code in a desktop app,
"OK, create me a new project, call it Project 2." It will never be able to do it. If the native agents do not have an MCP of their own application, it is impossible for them to control it, so they are stuck inside their CLI. It is the same as if you use NIM on Claude Code CLI and you create a command; ask Nemotron to use this command, and it will have to figure out where it lives, what tools it has access to, and how to use them.
Sir, 1st I don't understand what english you've written here.
2nd: using NIM API, you get almost all models including the deepseek and Ollamas and Qwen too. And all claude skills!
The line that hit me: "I just kept using whichever tool annoyed me the least for a particular job, until those choices turned into a system."
I'm a beginner β I started learning Python about two weeks ago and I've been writing tutorials about it. I don't have an Antigravity setup or a model pecking order. But I do have exactly one rule, and I didn't realize it was a rule until I read this.
I use ChatGPT for syntax and quick "what does this error mean" checks. Fast, throwaway, no attachment.
And I use it for absolutely nothing else. Not for learning concepts. Not for understanding why code works. Because the one time I tried to learn a Python topic entirely through ChatGPT, I realized a week later that I couldn't explain any of it back. I'd read good answers and retained nothing.
So now I only use AI after I've tried to figure it out myself. And if I'm learning a new concept, I read documentation first and ask AI second.
It's not a hierarchy of models. It's a hierarchy of moments. AI is the second opinion, never the first.
That's not a smart system. It's just the thing I landed on because everything else annoyed me more. Which is exactly what you're describing.
Great post. Made me realize I do have a system β I just never named it.
Funny that the one you trust less is Gemini 3.1 Pro, because it came out clean on the one thing I measured this week: when a tool quietly did nothing, it never once reported the task as done, across 13 scenarios. That makes me think 'hallucinates more' and 'lies about finishing' are separate axes. In an IDE agent the second one is the expensive one. A model that invents an API but tells you the build failed costs you a minute; one that says the refactor landed when it didn't costs you a deploy.
This is so real π
I definitely have a "don't waste Claude credits on this" brain now. Half the time I'm not even picking the smartest model, I'm picking the one that's the least effort for that task.
Somehow they all ended up with different jobs without me ever deciding it.
I'm a creature of habit, I Started using Antigravity And Now With Gemini 3.8 it's Just Great. Sometimes I Use Opus, almost Never. On The Chat I use gemini web And ChatGPT web, I tried To Change All That Flow a lot of Times But I Always Come Back π
What I find interesting is that this eventually becomes less about model preference and more about risk-based routing. A quick formatting task and a production-facing architectural change shouldn't need the same level of reasoning, verification, or tool access. The useful question may be βwhat is the cost of being wrong here?β rather than βwhich model is smartest?β That could lead to a more deliberate routing system where low-risk tasks get fast, inexpensive models, while high-impact tasks automatically trigger stronger reasoning, additional verification, or human review. The hierarchy people develop informally today could eventually become an explicit policy layer around their AI workflows.
The unwritten routing rules are the part that stands out, because implicit routing is also unauditable. When the bulk-work model hallucinates its way into a shipped bug, the postmortem has no record of which model touched what, so the cost of the headroom trade stays invisible and the hierarchy never gets challenged. A boring fix would be logging the model per task the way you log a deploy. Do you ever do a review pass over just the vibing AI's changes, or does it all land inline like human work?
No, By Vibing AI I mean the one where you do normal convos like about F1 or cricket like that.
The part about Gemini hallucinating more but using it anyway because of usage headroom β that's the real admission here. I caught myself doing the same math: 40+ API calls a day means I can't burn all of them on the expensive model, so the "good enough" one handles 80% of the work and I only switch when something breaks in a way I can't explain. The uncomfortable truth is most of us aren't optimizing for quality, we're optimizing for not hitting a limit mid-afternoon.
YEP.
But I'll admit sometimes it works when claude opus model doesn't
This is such an interesting observation about how our relationship with AI tools has evolved. We often donβt realize when a tool moves from being just a helper to becoming part of our personal workflow and decision-making process. The real skill is not only knowing which AI to use, but understanding when to trust it and when to question it. AI has become less about replacing our thinking and more about shaping how we think and create. Great perspective!
Haha so true! At The Printing World, we use a lighter AI model for drafting quick packaging descriptions and a stricter one to double-check exact die-line specs. Picking the right tool for the job keeps bad print errors from happening!
Cool! I never thought of this way. Thanks for the read! Have a great day!
To me there is no rigid hierarchy. It is entirely organic and constantly evolving based on what actually works right now. Take image generation for my presentations. A few months ago, Grok created absolute masterpieces. It was my go-to. Today, it suddenly feels kind of dumb for my current needs. The pivot? Switched to Gemini, and it is hitting the mark. Same story with coding assistance. A few months ago, a specific tool was spitting out pure gold. Today, it feels slow and outdated. The pivot? Back to Sonnet. In business, just like in our personal workflows, we often cling to tools, processes, or frameworks simply because that is what we used before or because we invested time in learning it. Agility is not about picking the right tool forever. It is about having the courage to abandon what is stupid today, even if it worked brilliantly yesterday.
The interesting part is that model selection probably shouldn't be the only routing decision. The task itself can determine how much autonomy and verification it needs. A quick formatting check might need a fast model with almost no tool access, while a production change may justify a stronger model plus tests or an approval step. That makes βwhich AI should I use?β a broader question of capability, risk, and verification cost. Over time, I could see personal AI workflows becoming less like a fixed model hierarchy and more like a policy: choose the cheapest capability that can complete the task while still providing enough evidence that the result is safe to use.
I agree with this. Different AI models are suited to different kinds of work.
For me, itβs similar to programming tools: I donβt expect one language, framework, database, or library to be the right choice for every problem.
The same applies to AI. I might use Midjourney for images, ChatGPT for reasoning or development, Gemini for certain research tasks, DeepSeek for other coding tasks, and so on. But even choosing the right category isnβt enough. Different models have different strengths and weaknesses within the same category. So I donβt really look for the βbest AI.β I look at the task first, then the capabilities I need, and then choose the tool and model that fit that particular job. The interesting part is learning where each one actually performs well.
The screenshot loop is the exact thing, an agent that can act in the world turned loose on a two-second glance job. More capability made the mistake more expensive, thats the whole trap.
My split is close to yours but I named Claude and ChatGPT first. Claude for the deep build, the planning, the stuff that has to hold together. ChatGPT for the spark, quick throwaway ideas, is this even possible questions. And I keep a local model for file-level edits where I read every diff, because small and scoped beats smart and loose when I already know what I want.
The one that took me a while to figure out was Gemini. It hallucinates and always agrees, so instead of fighting that I gave it the one job where that actually helps, raw brainstorm material, a vision board, stuff I treat like art pieces and mine later for what to do and what not to do. The flaw becomes the feature once you stop asking it to be reliable.
The moment it clicked was watching a small local model hand me confident broken code that looked totally legit until I checked it. Thats when I stopped asking which one is smartest and started matching the tool to the shape of the task, same as you.
This is actually how I think AI usage is evolving too. Instead of asking, βWhich AI is the best?β, the more useful question is, βWhich AI do I trust for this specific job?β
That distinction becomes even more important with AI agents like Husnex. An agent doesn't necessarily need to be the best at everything. It needs to be reliable within a defined workflow, understand the right context, and produce an action that is actually useful.
The interesting part is that trust becomes task-specific. A model can be great for calculations but unreliable for coding, or strong at planning but unnecessary for routine tasks. Over time, we naturally build these little AI hierarchies based on real usage rather than benchmark scores.
The real shift may be from choosing one βbest AIβ to building a personal stack of specialized AI tools and agents.
This is actually a really interesting way to think about AI tools. Iβve found that the βbestβ model changes completely depending on the task, and trust tends to become very specific rather than applying to the model as a whole.
Man, this hits so close to home.I recently realized Iβve done the exact same thing, but on absolute overdrive. At home, I run a self-hosted AI orchestrator I call Promethius, which manages 22 different LLMs. Yes, 22! When you have that many models under one roof, you either develop a subconscious routing hierarchy or you go completely insane trying to choose. For me, it became an accidental game of algorithmic matchmaking.My "screenshot loop" moment happened when I caught myself spinning up a massive, resource-heavy reasoning engine just to check if a single JSON string was valid. Like your Antigravity agent burning tokens to look at a draft, Promethius was practically starting a small house fire using maximum compute for a task a tiny, fast model could do in milliseconds.The biggest shift in my system happened as I climbed the Claude evolutionary ladder from 4.5 into the Claude 5 generation. My routing completely rewrote itself once I got my hands on Claude Fable 5.1 and the new Opus 5.5, and it forced me to completely re-map how I distribute tasks.Now, my modern Promethius routing protocol is all about tiered intensity. Claude Opus 5.5 gets the center-of-the-universe role because it is fast, adaptive, and significantly cheaper than older models, so it handles the bulk of my actual production coding and refactoring. If it gets stuck on a truly bizarre architectural edge case, I wake up Claude Fable 5.1, crank its effort parameters to max, and let it brute-force the logic. Everything elseβthe quick regex formatting and the basic rubber-duck questionsβgets tossed to the lightweight local 8B models so I don't burn premium tokens.You absolutely nailed it: more capability isn't always better. Sometimes, the "best" AI is just the one that matches the exact shape of the task without over-engineering a simple glance.
This really hits home. π Iβve noticed the same thing with AI tools, sometimes the βbetterβ answer isnβt what keeps me coming back. Itβs the tool that makes the whole process feel less frustrating.
Once you start choosing tools based on how they fit different tasks, you almost accidentally build your own AI workflow.
Fantastic article! I have made up my own LLM choices thru trial and error. You would be amazed how many people have this mental model and donβt even realize this. I use GPT-5.for deeper research or debug problems that others cannot. Usually GPT is very chatty but it does point out the issue and gives me options. I use GLM-5. (GPT is more expensive hence GLM) to research new projects and use it to break down the sub work with complexity indicators. Then I run GPT-5.* to confirm. I would use GLM to do complex modules and use deepseek flash for easier ones and deepseek pro for medium complex ones. I use local LLM Qwen 3.5 9B for menial tasks like updating some docks, fix links or updating GitHub issues, create branches. Itβs basically a mix of whatβs the best suited LLM and does not break the bank. For various field tests like with docker or real world e2e, when deepseek flash will get confused, I will try GLM and if that does not help I will occasionally bring in GPT to find out whatβs the actual issue and then get Deepseek flash to fix it. Itβs a bit of song and dance switching LLMs but it keeps token costs low, optimizes time. I use a mix of LLMs on generating dev.to article ideas with hooks, and will use a mix of LLMs with seed texts to write and usually have GPT do final read.i did Muse 1.3 and 1.2, itβs fast but a lot of the code written thru it, GPT or GLM would later rewrite due to fundamental design issues which pop up very late in software when I do field tests. I used Muse free on opencode and now have stopped. The rework is not worth it
I have built a set of dev.to rules processing last 6 months of popular articles, patterns that I have saved as Md and I will run through it. I do occasionally deviate but it helps land the articles right. I have a separate set of rules for hashnode, medium etc. :)
I've been reading more and more cases about agentic AI loop which burns unintended tokens. It's either they read files over and over (or in your case, taking screenshots repeatedly). From what I can infer, AI is just not what they advertise. It's good yes, but not that good that it can grant you wishes from a single command "make no mistakes". It's unreliable so it's much safer to know and do the job yourself than designate it to something that can't be held accountable.
the interesting bit is nobody wrote these routing rules down, so when the system breaks it just feels wrong instead of being debuggable. an explicit router you can inspect would at least tell you which rule fired and why, instead of you reverse engineering your own habits after the fact.
The split between "serious work" and "vibing" AI use is something I see in my own workflow. I use a coding agent for structured tasks and a conversational agent for brainstorming, and the two rarely overlap. The interesting part is that the "vibing" use case often leads to the most creative insights, even though it feels less productive. How do you think about evaluating the quality of unstructured AI interactions?
The "trust enough" part got me. Opus does pretty much all my routine stuff, and I only switch to Fable when Opus gets stuck on the same thing twice.
My weird one is proxies. I work on proxy stuff and every agent I tried was super confident and totally wrong about proxy errors. Give it a 407 and it would just stick the proxy password in the site's auth header π Switching models didn't help, none of them really knew that area.
Ended up building a small MCP server that gives it actual docs for that, and now it's fine. So I guess my hierarchy is more about what context the model has than which model it is.
Gemini for throwaway (the web interface is great for picking up old questions), Qwen 3.8 flash (qoder efficient) for bulk work, Qwen 3.8 max for serious thinking. Gemini pro for 'double check qwen's work', Claude Opus for 'I seriously need to make sure they didnt screw up'. Though most of the time, I actually trust Qwen more than I trust Opus, idk, for me even 3.8 Flash thinks 'smarter' for me, it pre-optimizes further than Opus and just sticks to better design principles imo.
This is so relatable! I definitely have a hierarchy too - one model for quick stuff, another for the hard problems.
For anyone testing different AI models, I've been using jzstoken for my API keys. Makes it easy to switch between models without signing up for each one separately. They give you $5 free to start, so you can test everything without spending a dime.
Honestly, I relate to this. π I donβt really have a βone AI for everythingβ either. I just use whichever one makes me fight with it the least for that particular task. At some point, you stop comparing models and start assigning them jobs like coworkers.
What do you use as the check when Gemini gives a calculation answer you no longer double-check?
Actually you always have answer keys to PYQs, so ....
Glad you're back bro! good post and good GIF!
Thanks! That's from friends btw but you wouldn't know thatπ.
HA HA. talk about post, not unnecessary stuff! But that's good cause we all default back to chatgpt just cause its fast.
Well for exam notes. Not for hackathon stuff, who uses antigravity there huh?π
The screenshot loop is the part I relate to most. Iβve hit the same thing building small tools: giving an agent more power can make a simple task slower and more expensive.
GIF xD
Nice write-up Dhruv :D
Thanks di!π
So relatable. At The Printing World, we use strict, precise tools for checking exact box dielines and specs, but stick to quick, low-stakes tools for rapid marketing drafts. Matching the tool to the actual risk level is everything.
Great topic and text