Skip to main content
geminihigh Impact

Can Gemini 3.1 Pro Replace Claude—at 87% Lower Cost?

News by OneHuman

Gemini 3.1 Pro launches Feb 19, topping APEX-Agents over Claude Opus 4.6. API pricing 7.5× cheaper than Anthropic. Same $20/month for consumers.

breaking-newsgeminigooglemodel-releaseknowledge-workagentic-aifebruary-2026

Published: February 21, 2026 Reading time: 2.5 minutes


Google's Counterpunch Lands

Two weeks after Anthropic's Claude Opus 4.6 triggered a $285B software selloff, Google shipped its response.

Gemini 3.1 Pro launched February 19, 2026. It tops the APEX-Agents leaderboard—a benchmark measuring performance on real professional tasks—beating both Claude Opus 4.6 and GPT-5.2. And it costs 87% less per token than Anthropic's flagship.

This isn't Google catching up. It's Google repricing the market.

Context: Gemini 3 Pro launched November 2025 with strong benchmarks but no decisive advantage over Anthropic. Gemini 3.1 Pro is the first 0.1-increment Google has ever shipped—a signal that the quarterly release cadence is now the baseline. Google is no longer waiting for a generation jump to ship improvements.

The Numbers That Matter

On ARC-AGI-2—the benchmark designed to test novel reasoning that can't be memorized—Gemini 3.1 Pro scored 77.1%. One generation ago, Gemini 3 Pro scored 31.1%. That's a 46-point jump in a single release cycle. Claude Opus 4.6 sits at 68.8%.

On APEX-Agents—real professional agentic tasks—the leaderboard now reads:

  • Gemini 3.1 Pro: 33.5% ← new #1
  • Claude Opus 4.6: 29.8%
  • GPT-5.2: 23.0%
  • Gemini 3 Pro: 18.4%

"Gemini 3.1 Pro is now at the top of the APEX-Agents leaderboard," said Mercor CEO Brendan Foody. "These results show how quickly agents are improving at real knowledge work."

The Price Gap Is the Story

Google didn't raise prices with the upgrade. Gemini 3.1 Pro uses identical pricing to Gemini 3 Pro:

Gemini 3.1 Pro Claude Opus 4.6
Input (per 1M tokens) $2.00 $15.00
Output (per 1M tokens) $12.00 $75.00
Context window 1M tokens 200K tokens
Consumer plan $20/month $20/month

Same consumer price. 7.5× cheaper API. 5× larger context window. For any company building on top of these models, this changes the architecture conversation immediately.

Where Claude Still Leads

Gemini 3.1 Pro does not top every benchmark. Claude Opus 4.6 retains leads in:

  • Expert task quality (GDPval-AA: Opus 4.6 at 1606 Elo vs. Gemini's 1317)
  • GUI automation (OSWorld: Opus 4.6 at 72.7%; Gemini hasn't published a score)
  • Long-form output (Opus 4.6 supports 128K output tokens vs. Gemini's 64K)
  • Arena text quality (LM Arena: Claude Opus 4.6 still leads on human preferences)

SWE-Bench Verified—real GitHub issue resolution—is essentially tied: Gemini 3.1 Pro at 80.6%, Opus 4.6 at 80.8%.

The pattern: Google wins on price and breadth. Anthropic wins on depth and human preference. For knowledge workers, the threat is now coming from both directions.

The Jobs in the Crosshairs

APEX-Agents doesn't test trivia. It measures performance on long-horizon professional workflows—the kind of tasks that justify $80K–$200K salaries:

  • Multi-step research synthesis (analyst work)
  • Code review and automated bug resolution (junior developer work)
  • Document drafting with source citations (paralegal, associate work)
  • Project coordination and task breakdown (project manager work)

When the #1 model on that benchmark nearly doubles its score in one release cycle, the question for employers isn't "should we pilot this?" It's "how quickly can we restructure headcount around it?"

The pattern emerging: Claude Opus 4.6 accelerated the threat to white-collar work. Gemini 3.1 Pro just made that threat 7.5× cheaper to deploy.

Who Gets Access

Consumers: Gemini 3.1 Pro is rolling out in the Gemini app for Google AI Pro and Ultra subscribers ($20/month). Also available in NotebookLM for Pro/Ultra users.

Developers: Preview access via Gemini API, AI Studio, Vertex AI, Gemini CLI, Antigravity, and Android Studio. Also integrated into GitHub Copilot, Visual Studio, and VS Code—meaning developers already inside the Microsoft ecosystem get access without switching tools.

Two API endpoints are available:

  • gemini-3.1-pro-preview — general use
  • gemini-3.1-pro-preview-customtools — prioritizes custom tool calling in agent workflows

Note: API access experienced outages on launch day due to high demand—a recurring pattern for major model releases. If you hit rate limits, the preview is still rolling out capacity.

Still in preview. GA launch date and final pricing have not been announced.

What You Should Ask Your Employer

Q: Should I be worried about this model? A: If your job involves multi-step research, code review, document synthesis, or project coordination—yes. APEX-Agents measures exactly those tasks.

Q: Is this better than Claude for creative or strategic work? A: Not yet. Human evaluators still prefer Claude Opus 4.6 for nuanced, expert-level output. But "good enough at 87% less cost" is how enterprise procurement works.

Q: Does the 5× larger context window matter? A: For codebase analysis, legal document review, or long research chains—dramatically. An entire codebase fits in a single Gemini prompt. With Claude, you're chunking.

Q: Should I switch from Claude to Gemini? A: Not necessarily. For high-stakes output—board memos, legal briefs, nuanced strategy—Opus 4.6 still has the human preference edge. But for high-volume agentic pipelines where cost compounds, Gemini 3.1 Pro is the rational choice. The answer is: use both, depending on the task.

What Happens Next

30 days: API adoption spikes among cost-conscious users and developers. At $2/1M tokens, any company paying Opus 4.6 rates will run a side-by-side test this quarter. Some won't switch back.

90 days: Watch for Google to remove the "preview" label and announce GA pricing. If they hold $2/$12, Anthropic faces its first real pricing pressure since GPT-4. If Google raises prices at GA—classic bait-and-switch.

6 months: The APEX-Agents score nearly doubled in one generation (18.4% → 33.5%). If that velocity holds, Q4 2026 looks like a model that handles 70–80% of professional cognitive tasks without human supervision.

Watch for: Anthropic's response. They held Opus 4.6 pricing flat after the $285B selloff. Whether they match Google's $2/1M input rate—or double down on quality premium—defines the next chapter of this arms race.

Bottom Line

Gemini 3.1 Pro is the first Google model that competes directly with Anthropic and OpenAI on agentic professional tasks—and undercuts them on price by 7.5×. It doesn't beat Claude in every dimension, but it doesn't need to. It needs to be good enough at a price that forces procurement conversations.

Who this affects most:

  • Developers: Google is now inside your IDE (GitHub Copilot integration). You're going to use it whether you choose to or not.
  • Enterprise buyers: The TCO math changed on February 19. Same-quality agentic output at 87% lower API cost demands a budget review.
  • Knowledge workers: The model that topped the "real professional tasks" leaderboard costs less than your Netflix subscription.

The $20/month price didn't change. The threat level did.


Sources

Share This Article

"Google shipped a model that beats Anthropic and OpenAI on real professional tasks—and costs 87% less. The arms race isn't about who's smartest anymore. It's about who's cheapest."
— News by OneHuman
"Gemini 3.1 Pro scored 33.5 on APEX-Agents vs. Opus 4.6's 29.8—while costing $2/1M input tokens vs. Opus 4.6's $15. Same agentic output, 7.5× less cost. The CFO is already running that math."
— News by OneHuman
"ARC-AGI-2 jumped from 31.1% to 77.1% in one generation—the largest single-gen reasoning gain ever recorded. That's not an incremental upgrade. That's a category shift."
— News by OneHuman
"When AI professional task performance nearly doubles in a single quarterly release, the question isn't 'when will AI replace me?' It's 'why hasn't my employer done the math yet?'"
— News by OneHuman

Get alerted when Gemini changes its pricing or limits.

Independent. No ads. No affiliates. No investors.

Or join Pro for unlimited recommendations and Pro Deep Analysis.