🔍 Read the full analysis: My AI Stack In September 2026: Who Builds, Who Digs, Who Decides on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Thorsten Meyer says he uses Claude Opus 5.5 for building and GPT-6.1 Sol for detailed reviews, following Sol’s release on Sept. 29. His comparison of Artificial Analysis Intelligence Index v4.3.x scores finds a wide gap in estimated task costs, but the index does not establish which model performs best on a particular team’s work.
Thorsten Meyer says he has made Claude Opus 5.5 his main model for building software and GPT-6.1 Sol his model for detailed investigation and review, following Sol’s release on Sept. 29. His updated stack reflects a comparison in which several models score within about 20 points on the Artificial Analysis Intelligence Index, while estimated costs per task vary widely.
Meyer’s comparison uses the Artificial Analysis Intelligence Index v4.3.x, which he describes as a measure of general capability rather than a verdict on any particular workload. In the listed top settings, Opus 5.5 scores 58 at an estimated $5.98 per task; GPT-6.1 Sol at xhigh scores 51 at $0.39. The figures are estimates from the index, not reported costs for Meyer’s own completed tasks.
The other models in his comparison fill narrower roles. Sonnet 5.5 scores 56 at max effort and costs an estimated $7.60 per task; Fable 5.1 scores 53 at $7.63; and GPT-6 Astra scores 53 at $3.26. GPT-6 Luna scores 37 at $0.07, which Meyer assigns to classification, extraction and routing. He says Astra and Fable are options for cases where his own tests favor them.
Effort settings also change the estimates. For Opus 5.5, the index lists a score of 54 for $1.82 per task at high and 56 for $3.46 at xhigh. Meyer uses those settings for development and harder problems, respectively. At max, the score rises to 58 while estimated cost reaches $5.98. For Sonnet 5.5, the listed max setting costs $7.60 for a score of 56; Meyer favors its high setting, listed at 47 for $1.08.
Opus builds. Sol reviews. Jev decides.
One price tape, six models
Score against cost, at every effort setting
The effort dial moves the bill more than the model
Claude Opus 5.5
Claude Sonnet 5.5
GPT-6.1 Sol: near-Astra scores at a fraction of the price
Three published settings
| Setting | Index | Cost per task | Output tokens | First token |
|---|---|---|---|---|
| medium | 48 | $0.21 | 15M | 5.3 s |
| high | 50 | $0.32 | 25M | 57 s |
| xhigh | 51 | $0.39 | 36M | 69 s |
Same score band, very different bill
My stack: who builds, who reviews
Cheaper tokens are not cheaper work
Read the numbers with four warnings
Part 2: Jev, the model that decides instead of writing
One call in, typed answers out
Three question types
Confidence is the superpower
Three uses running in my publishing operation
The fit test, then the shadow test
- Replay 300 to 500 past decisions
- Compare overall and per confidence band
- Read 20 disagreements, decide who was right
- High band at 95% or better?
- Own flag, off by default
- Canary on 5 to 10 units
- Roll out in the confident band only
24 use cases, sorted by how well they fit
Proven in production
- 1Relevance gate
- 2Language check
- 3Classifier fallback
Publishing and content
- 4Thin-source detector
- 5Same-event dedupe
- 6Product fits roundup
- 7Disclosure present
- 8Headline quality
- 9Comment moderation
Commerce and support
- 10Support-ticket routing
- 11Return-reason coding
- 12Review to feature complaints
- 13Catalogue taxonomy
- 14Order-fraud pre-triage
Software and AI systems
- 15LLM guardrail
- 16RAG passage filter
- 17Citation check
- 18Tool and intent routing
- 19Log-line triage
- 20PR risk triage
Business ops and home
- 21Inbox triage
- 22Expense categorisation
- 23Lead qualification
- 24Smart-home intent
Limits, cost and one hard rule
Why the Review Seat Matters
Meyer’s central decision is to use a model from another family to review work produced by his main builder. He says GPT-6.1 Sol’s estimated cost of $0.32 to $0.39 per task makes routine review practical for him. That is his operating choice, not evidence that every Sol review will catch errors or that the same economics apply to other users.
The split also shows why a benchmark score alone may not settle model selection. The index’s general scores sit relatively close for several listed models, but their estimated task costs and response characteristics differ. Meyer says teams should shadow-test models on their own work before switching. He also cautions that more effort cannot supply missing requirements, and that a second model can share a flawed specification with the first.
His review rules distinguish evidence from approval: a passing test does not by itself mean work should ship. When a review finds a failure, Meyer says he sends the failing case and evidence back for correction rather than simply asking the builder to try harder.
How the Model Roles Took Shape
The stack follows releases and index entries listed across September. Meyer says Opus 5.5 was released on Sept. 22, Luna on Sept. 22, Sonnet 5.5 on Sept. 28 and GPT-6.1 Sol on Sept. 29. Fable 5.1 is listed with a Sept. 1 release, while Astra is listed with a Sept. 3 release.
Sol’s launch price, according to Meyer, is $2 per million input tokens and $10 per million output tokens, the same as the prior GPT-6 Sol. He reports three available effort settings in the index: medium, high and xhigh. At medium, Sol scores 48 at an estimated $0.21 per task; at high, 50 at $0.32; and at xhigh, 51 at $0.39. The index lists output-token totals of 15 million, 25 million and 36 million for those settings.
Meyer reports that high and xhigh take 57 and 69 seconds, respectively, to produce a first token in the index. He says this makes them a poor fit for interactive use. He also notes that the index had not yet published low or max settings for GPT-6.1 Sol when he wrote his update.
““The practical reading: Sol is not the model I ask to build. It is the model I can afford to run on everything.””
— Thorsten Meyer
What the Index Cannot Settle
The index measures general capability, and Meyer says its results are not a verdict on an individual workload. His article does not report a controlled head-to-head test across his own projects, so the listed scores and estimated costs do not establish which model will be most effective for another team.
Meyer also cautions that a one-point score difference may fall within measurement noise. The index had not published GPT-6.1 Sol’s low or max results at the time of his update. The figures do not settle how model performance may change as new measurements appear, or whether the reported task costs match usage in a particular workflow.
One cost discussion in the source text is incomplete: it begins an illustrative example about model prices and human review time but ends mid-sentence. No conclusion from that example can be reported from the supplied material.
Test Before Changing Defaults
Meyer’s next step for teams considering a similar setup is to shadow-test models against their own tasks before changing defaults. The supplied update does not announce a later test date or a planned change to his chosen roles. Its immediate recommendation is to compare model quality and cost on the work that needs to be done, then use the results to set defaults.
Further comparisons may be possible as the index adds settings that were not yet listed for GPT-6.1 Sol. Until then, Meyer’s reported setup is Opus 5.5 at high or xhigh for building, Sol at high or xhigh for details and review, and other models for selected tasks. Whether he changes that arrangement will depend on his tests and later index data.
Key Questions
Which models does Meyer use for building and review?
Meyer says he uses Opus 5.5 as his main builder and GPT-6.1 Sol for detailed investigation and review.
What does GPT-6.1 Sol cost in the index comparison?
The index lists Sol at an estimated $0.32 per task at high and $0.39 at xhigh. These are index estimates, not guaranteed costs for an individual workload.
Does the comparison prove Sol is better value for every team?
No. Meyer says the index measures general capability and recommends shadow-testing models on a team’s own tasks before switching.
Why does Meyer avoid Opus 5.5 at max effort for routine work?
In the index figures he cites, max scores 58 at an estimated $5.98 per task, compared with high at 54 for $1.82. Meyer uses high for development and xhigh for harder work, reserving max as rarely worthwhile.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
