Essay 5 min read

The right model, not the best model

Test the model on your work before paying more.

Table of contents

I pay $100 a month for ChatGPT and another $100 for Claude. I can still hit the limits when I send too much everyday work to the strongest models.

A few months ago, I tried GPT-5.6 Luna on several real tasks. I expected less. The results were good enough for me, fast, and very cheap.

My expectations were out of date. I still use powerful models. I just want them on the work that needs them.

What my setup looks like now

I recently started using oh-my-pi. It lets me assign models to roles. Here is a simplified view of my setup on September 25, 2026:

What needs doingModelEffort
Everyday workClaude Opus 5.5Low
Planning, advice, and deeper reasoningGPT-6 AstraLow
Subtasks and web searchGPT-6 LunaExtra high
Understanding imagesGPT-6 SolLow
Design workClaude Fable 5.1Low

I also set fallback models for some roles. I choose a model and an effort level for the job. This is my current setup, not a ranking to copy.

Choose again for each part of the job

I would not hand Luna the whole job of refining business requirements inside a legacy application. That needs judgment about the existing system and what a change will affect.

I often use a stronger model to lead the work and delegate smaller tasks to other agents. I choose a model for the whole problem, then choose again for each piece.

A hard project can contain easy work. A short task can hide a difficult decision.

A clear spec helps. It does not make every task easy. If the lead model has to redo a subagent’s work, the cheaper model saved nothing.

I test models on work I understand

I test a promising model on several real tasks I know well enough to do myself. I check the result, the fixes it needs, and how it feels to work with.

Personal preference matters. Two correct answers can differ in code style, structure, or how much work they leave me.

Still, a polished answer can be wrong. That is why I start with tasks I can judge myself.

I am considering a personal benchmark for coding, landing pages, marketing, and writing. I have not built it yet. Code can sometimes be checked directly. A campaign that sounds good still has to work on real people.

Earlier this year I wrote Chasing models won’t make you better at using them. I still think learning one model well matters more than testing every release. But that should not become an excuse to keep an outdated default.

Products need more than my opinion

My taste can guide my own work. It should not decide which model serves thousands of users.

For an application, define an acceptable result for each task. Which mistakes can be fixed? Which make the answer useless? Build evals from the work users actually send through the system.

When a cheaper model arrives, test it before switching. Check quality, time, retries, and corrections. The price of one call tells you too little.

Evals can miss important failures. Some work also needs human judgment. But a product decision needs more than one person’s reaction to a few answers.

This is the approach I recommend. I have not measured a production model switch of my own.

Test the default before paying more for it

For me, better choices mean more work within the subscriptions I already pay for. The monthly bill is the same. Others might be able to move to a cheaper plan, depending on its limits and features.

In a product billed by usage, a saving can repeat across thousands of users. So can a quality problem.

You can test your own default without rebuilding your setup:

  1. Pick a few tasks you do often and can judge yourself.
  2. Give them to your usual model and a smaller one, with the same context.
  3. Compare the result, your corrections, and the total time.
  4. Move a task to the smaller model only when it earns that place.

Luna did not replace every model in my workflow. It showed me that one of my defaults was out of date.

Get the next one by email

I’ll send you new articles, plus the interesting links, sources, working notes, and behind-the-scenes details that shaped them.

Free. Unsubscribe anytime. Powered by Substack.