The Model Is Only 10%: The Real Lesson of the New SDLC
AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: The Model Is Only 10%: The Real Lesson of the New SDLC on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

TL;DR

A recent Google whitepaper emphasizes that in AI-assisted software engineering, the model itself is only a small part of the system. The majority of behavior depends on how the AI is configured, guided, and integrated through harnesses and context engineering.

A new Google whitepaper titled The New SDLC With Vibe Coding emphasizes that the model accounts for only 10% of the behavior in AI-assisted software development. The paper argues that the real driver of system performance and reliability is how the AI is configured, guided, and integrated, through what it calls the harness and context engineering. This shift in understanding has significant implications for how organizations approach AI development and deployment.

The whitepaper, authored by Addy Osmani, Shubham Saboo, and Sokratis Kartakis, states that 85% of professional developers now use AI coding agents regularly, with over half using them daily. It reports that approximately 41% of new code is generated by AI, marking a rapid integration of AI tools into software workflows.

The core insight is that model size and capabilities are less important than the harness—the prompts, tools, rules, and observability layers surrounding the model. Evidence from public benchmarks shows that tweaking harness components can significantly improve AI performance without changing the model itself, sometimes by over 13 points in test scores.

The paper stresses that failures in AI agents are often due to configuration errors—missing tools, vague rules, or noisy context—rather than the model’s raw ability. This means organizations should focus on building, owning, and improving their harnesses rather than constantly chasing the latest model release.

At a glance
reportWhen: published March 2026
The developmentThe new Google whitepaper highlights that in AI-driven SDLC, the model accounts for only 10% of system behavior, shifting focus to harness design and context management.
The Model Is Only 10% — The New SDLC With Vibe Coding
AI Dispatch · Field Notes
Google · Osmani, Saboo & Kartakis · May 2026

The model is only 10%

A Google whitepaper argues software’s biggest shift is from writing code to expressing intent. Its sharpest claim: the model you obsess over is the smallest part of the system — the scaffolding around it does the real work.

A spectrum, not a binary — the differentiator is how outputs get verified
Vibe Coding
Casual prompts · “does it seem to work?” · disposable code · high risk
Structured AI-Assisted
Detailed prompts + constraints · manual testing · features in real codebases
Agentic Engineering
Formal specs · automated tests + evals + CI gates · production scale · low risk
Tests verify the deterministic; evals verify the rest. Without both, it’s vibe coding — however clever the prompt.
The idea worth building your strategy around
Agent = Model + Harness
~10%
HARNESS — prompts · tools · context · hooks · sandboxes · observability
MODEL~90% IS YOUR SURFACE AREA, NOT THE PROVIDER’S
Outside Top 30 → Top 5 on Terminal Bench 2.0 by changing only the harness — same model.
“Most agent failures, examined honestly, are configuration failures” — a missing tool, a vague rule, a noisy context.
The economics: it’s a token-cost problem (CapEx vs OpEx)
Vibe Coding
Low CapEx · High OpEx
Looks free, hides debt: token burn (fix-it loops), maintenance tax (AI spaghetti), security remediation. Crosses over to 3–10× more per feature.
Agentic Engineering
High CapEx · Low OpEx
Pay upfront (specs, evals, context), then ship cheaply. Levers: context engineering for first-pass success + intelligent model routing — cheap models for the easy work.
85%
of devs use AI coding agents (51% daily)
41%
of all new code is AI-generated
~90%
of agent behavior is the harness, not the model
+19%
longer on some tasks (METR) — verification is the cost
The read

The clearest map yet of how serious AI development works — and mostly tool-agnostic. But it’s a Google funnel: the concepts are neutral, the on-ramps point to Gemini, Jules & the ADK. If the harness is 90% and it’s yours, your moat and your costs both live there — so own your scaffolding, route across models, and remember: AI amplifies whatever engineering culture it lands in.

Source: Osmani, Saboo & Kartakis, “The New SDLC With Vibe Coding,” Google (May 2026). Figures are the paper’s own, incl. METR & LangChain. Analysis is the author’s.
thorstenmeyerai.com

Implications for AI Development Strategies

This shift in understanding changes how organizations should invest in AI. Instead of prioritizing access to the largest or most advanced models, companies should focus on designing effective harnesses and managing context. This approach can lead to cost savings, improved reliability, and greater control over AI behavior. It also suggests that long-term competitive advantage lies in mastering system configuration rather than model size, which has significant implications for AI governance and operational costs.

lweiyupeixx Press Model Separator Press Type Automatic Model Parts Detacher Part Separation Tool Hobby Assembling Model Ergonomic

lweiyupeixx Press Model Separator Press Type Automatic Model Parts Detacher Part Separation Tool Hobby Assembling Model Ergonomic

  • Press Type Model Separator: Effortlessly separates model components
  • High-Strength ABS Construction: Stable and durable material
  • Ergonomic Design: Comfortable operation for users

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on AI in Software Engineering

By early 2026, AI tools have become integral to software development, with a majority of developers using AI coding agents regularly. The industry has largely focused on acquiring and deploying advanced models, such as GPT variants and Claude, but recent research indicates that system behavior depends more on how these models are integrated and controlled.

The whitepaper builds on prior discussions about ‘vibe coding’—quick, unstructured prompts—and introduces a spectrum that includes disciplined, structured approaches like agentic engineering. This evolution reflects a broader understanding that verification, judgment, and system design are more critical than raw model capabilities.

“The model is only 10% of what determines behavior; the harness is 90%.”

— Addy Osmani

Google Antigravity Pro: Harness Google AI to accelerate software delivery, improve code quality and scale modern AI-assisted development. (AI Coding)

Google Antigravity Pro: Harness Google AI to accelerate software delivery, improve code quality and scale modern AI-assisted development. (AI Coding)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unclear Aspects of Harness Optimization

While evidence shows harness design significantly impacts AI performance, the optimal methods for creating, maintaining, and scaling these configurations are still being developed. It remains unclear how quickly organizations can adopt these practices at scale and what specific tools will emerge to support this shift.

AI Context Engineering: Architecting Intelligence Through Prompt Structures, Tools, and Memory

AI Context Engineering: Architecting Intelligence Through Prompt Structures, Tools, and Memory

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI System Design and Adoption

Organizations are expected to reevaluate their AI strategies, investing more in system architecture, context management, and configuration tools. Future research and industry practice will likely focus on developing standardized harness components, best practices for context engineering, and metrics to measure system reliability beyond model capabilities. Monitoring how these approaches influence cost, security, and performance will be key in the coming months.

AI Observability Systems: AI monitoring frameworks | AI observability tools | performance analytics in AI | AI performance metrics | AI system monitoring | real-world AI applications

AI Observability Systems: AI monitoring frameworks | AI observability tools | performance analytics in AI | AI performance metrics | AI system monitoring | real-world AI applications

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why is the model size less important than the harness?

Because the behavior of AI systems depends more on how the model is integrated, guided, and constrained through prompts, tools, and rules, rather than the model’s raw capabilities alone.

What does system harness include?

It includes prompts, rules, tools, context policies, observability layers, and other configuration components that shape how the AI operates within a system.

How can organizations improve AI reliability?

By focusing on designing effective harnesses, managing context carefully, and continuously refining configuration rather than solely relying on larger or newer models.

What are the economic implications of this shift?

Investing in system configuration and context engineering can reduce operating costs, improve security, and provide a more sustainable long-term approach than constantly upgrading models.

Will this change how AI tools are developed or used?

Yes, it encourages a focus on system design, configuration, and verification, making AI deployment more predictable, controllable, and cost-effective.

Source: ThorstenMeyerAI.com

COLLEGE MOVE-IN

College move-in / dorm season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

9 Best 4K Monitors For Work And Play In 2026

Discover the best 4K monitors of 2026 for productivity and gaming, including top picks like Dell, LG, and Samsung, based on specs and user needs.

US Government directive to suspend access to Fable 5 and Mythos 5

The US government has issued a directive to halt all access to Anthropic’s Fable 5 and Mythos 5, citing national security concerns over potential jailbreak vulnerabilities.

The Local-First Agentic Operator

A single operator, using agentic AI, has built an 18-product portfolio across diverse domains, challenging traditional organizational needs.

Verizon Surges In Global Coverage

Verizon reports a major surge in its global network coverage, impacting international connectivity and market presence.