{"id":20723,"date":"2026-08-11T14:35:23","date_gmt":"2026-08-11T14:35:23","guid":{"rendered":"https:\/\/aimade.tech\/?p=20723"},"modified":"2026-08-11T14:35:23","modified_gmt":"2026-08-11T14:35:23","slug":"claude-opus-4-7-vs-gpt-5-4-vs-gemini-3-1-pro-2026-benchmark","status":"publish","type":"post","link":"https:\/\/aimade.tech\/?p=20723","title":{"rendered":"Claude Opus 4.7 vs GPT-5.4 vs Gemini 3.1 Pro: The 2026 Frontier Model Benchmark"},"content":{"rendered":"\n<figure class=\"wp-block-image size-full\"><img data-recalc-dims=\"1\" loading=\"lazy\" decoding=\"async\" width=\"2560\" height=\"1429\" src=\"https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/08\/hero-scaled.jpg?resize=2560%2C1429&#038;ssl=1\" alt=\"Three white cubes labeled with hexagon, circular arrow, and triangle symbols representing Claude Opus 4.7, GPT-5.4, and Gemini 3.1 Pro arranged on a dark wood desk next to a laptop showing a stylized line chart\" class=\"wp-image-20721\" srcset=\"https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/08\/hero-scaled.jpg?w=2560&amp;ssl=1 2560w, https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/08\/hero-scaled.jpg?resize=300%2C167&amp;ssl=1 300w, https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/08\/hero-scaled.jpg?resize=1024%2C572&amp;ssl=1 1024w, https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/08\/hero-scaled.jpg?resize=768%2C429&amp;ssl=1 768w, https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/08\/hero-scaled.jpg?resize=1536%2C857&amp;ssl=1 1536w, https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/08\/hero-scaled.jpg?resize=2048%2C1143&amp;ssl=1 2048w, https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/08\/hero-scaled.jpg?resize=600%2C335&amp;ssl=1 600w\" sizes=\"auto, (max-width: 1000px) 100vw, 1000px\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Three frontier models, four real-world task categories, one honest verdict. Claude Opus 4.7, GPT-5.4, and Gemini 3.1 Pro have been the production-default choices for most AI teams since April 2026 \u2014 and the gap between them is narrower than any of their marketing pages suggest. This benchmark comparison synthesizes published vendor data, third-party leaderboards, and an independent multi-skill agentic coding test to give practitioners the practical verdict: which frontier model to pick, when, and what the real tradeoffs are on cost, context, and capability.<\/p>\n\n\n\n\n\n\n\n<p class=\"wp-block-paragraph\">We&#8217;ve tested where the three diverge on coding, reasoning, pricing, multimodal support, and known weak points. The TL;DR: there is no single winner. Each model has a deployment scenario where it pulls clearly ahead. The rest of this article is the data behind that verdict.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"the-three-flagships\">The three flagships and their release dates<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Reading any 2026 frontier model comparison fairly requires anchoring on the release dates. The staggered shipping means GPT-5.4 had roughly six weeks of solo availability before Gemini 3.1 Pro landed, and Gemini had about ten weeks before Opus 4.7 shipped. Each release triggered the next vendor&#8217;s response. Rather than audit each model&#8217;s claims in isolation, the cleaner read is to compare them on benchmarks that were stable across the release window.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Claude Opus 4.7<\/strong> launched on <a href=\"https:\/\/www.anthropic.com\/news\/claude-opus-4-7\" target=\"_blank\" rel=\"noopener\">April 16, 2026<\/a> as Anthropic&#8217;s flagship, with a single SKU at the top end of the Claude line. Anthropic positioned the launch around software engineering gains and explicitly &#8220;<a href=\"https:\/\/the-decoder.com\/anthropics-claude-opus-4-7-makes-a-big-leap-in-coding-while-deliberately-scaling-back-cyber-capabilities\/\" target=\"_blank\" rel=\"noopener\">deliberately scaling back cyber capabilities<\/a>&#8221; \u2014 a framing worth noting because it tells you what Anthropic decided was the right tradeoff.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>GPT-5.4<\/strong> was released on <a href=\"https:\/\/techcrunch.com\/2026\/03\/05\/openai-launches-gpt-5-4-with-pro-and-thinking-versions\/\" target=\"_blank\" rel=\"noopener\">March 5, 2026<\/a>, with three variants: standard, Pro, and Thinking. The thinking variant is the OpenAI equivalent of Anthropic&#8217;s extended-thinking modes and Google DeepMind&#8217;s &#8220;Deep Think.&#8221; Pro is the higher-precision, slower, more expensive tier. (For broader context on OpenAI&#8217;s positioning, see the earlier <a href=\"\/?p=20507\">Local LLM Setup 2026<\/a> piece on choosing a model for local deployment versus API access.)<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Gemini 3.1 Pro<\/strong> is the February 2026 <a href=\"https:\/\/blog.google\/innovation-and-ai\/models-and-research\/gemini-models\/gemini-3-1-pro\/\" target=\"_blank\" rel=\"noopener\">Google DeepMind flagship<\/a>, with a separate Gemini 3.1 Pro Preview tier that hit general availability in March. The product positioning is multimodal-first: native inputs for text, images, audio, video, PDFs, and code repositories with no preprocessing pipeline. (For a synthesis of every prior frontier-model round from late 2025, see our <a href=\"\/?p=20485\">Gemini 2.5 Pro vs GPT-4.5 vs Claude 3.7 Sonnet rankings<\/a> piece \u2014 the 2026 lineup is the generation that supersedes it.)<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">If you&#8217;re using the API today, all three models can replace each other for most text tasks. The decision logic is about which dimension of the tradeoff matrix matters most to your workload: coding depth, reasoning precision, multimodal breadth, or cost. The next four sections walk through each of those dimensions with primary-source benchmark numbers.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"coding-benchmark\">How the three stack up on coding benchmarks<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Software engineering is where Claude Opus 4.7 has the clearest 2026 lead. Anthropic&#8217;s research page for Opus 4.7 cites <a href=\"https:\/\/www.anthropic.com\/research\/claude-opus-4-7\" target=\"_blank\" rel=\"noopener\">SWE-bench Verified at 87.6%<\/a> (up from Opus 4.6&#8217;s 80.8%) and <a href=\"https:\/\/www.anthropic.com\/research\/claude-opus-4-7\" target=\"_blank\" rel=\"noopener\">SWE-bench Pro at 64.3%<\/a> (up from 53.4%). The first figure \u2014 SWE-bench Verified \u2014 is the canonical coding agent benchmark held by <a href=\"https:\/\/www.swebench.com\/verified\" target=\"_blank\" rel=\"noopener\">Princeton&#8217;s SWE-bench<\/a> team. As of August 2026, Opus 4.7 is among the top-scoring generally available models on that benchmark.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">GPT-5.4&#8217;s coding performance is strong but not at the Opus level. <a href=\"https:\/\/developers.openai.com\/api\/docs\/models\/gpt-5.4\" target=\"_blank\" rel=\"noopener\">OpenAI&#8217;s API documentation<\/a> reports the model&#8217;s SWE-bench Verified numbers in the 74\u201382% range across variants \u2014 better than GPT-5.2, but still behind Opus 4.7. Where GPT-5.4 picks up ground is on production engineering tasks that involve multi-file reasoning across a large codebase: the Pro variant in particular has a reputation for stability on this kind of work.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Gemini 3.1 Pro&#8217;s coding strength is on competitive programming-style benchmarks. The model hits <a href=\"https:\/\/blog.google\/innovation-and-ai\/models-and-research\/gemini-models\/gemini-3-1-pro\/\" target=\"_blank\" rel=\"noopener\">LiveCodeBench Pro at 2887 Codeforces Elo<\/a>, which compares competitively to mid-tier competitive programmers. For &#8220;write me a function&#8221; or &#8220;debug this algorithm&#8221; workflows, Gemini is competitive. For &#8220;refactor this 50-file PR,&#8221; Opus 4.7 still leads.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For context on the coding-agency landscape beyond these three flagships, the <a href=\"\/?p=20635\">Claude vs GPT-5 for Code Review: 2026 Engineering Benchmark<\/a> deep-dive we ran earlier this year shows where GPT-5 (the predecessor) sits on review-specific tasks \u2014 most of those findings still apply, with Opus 4.7 widening the lead on multi-turn review chains. The earlier <a href=\"\/?p=20070\">Introducing Claude Opus 4.7<\/a> coverage captures Anthropic&#8217;s positioning at launch, and the <a href=\"\/?p=1604\">April 2026 roundup<\/a> put this release in the broader release window alongside Meta Llama 4 and Google&#8217;s other Gemini updates.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"reasoning-tests\">Reasoning tests: GPQA, MMLU-Pro, AIME<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Reasoning benchmarks are where the three converge. GPT-5.4 Pro leads <a href=\"https:\/\/openai.com\/index\/introducing-gpt-5-4\/\" target=\"_blank\" rel=\"noopener\">GPQA Diamond at 94.4%<\/a> (the standard variant hits 92.8%) and <a href=\"https:\/\/openai.com\/index\/introducing-gpt-5-4\/\" target=\"_blank\" rel=\"noopener\">AIME 2026 at 99.17%<\/a>. Gemini 3.1 Pro reports <a href=\"https:\/\/deepmind.google\/models\/model-cards\/gemini-3-1-pro\/\" target=\"_blank\" rel=\"noopener\">MMLU-Pro at 93.8%<\/a> on its official model card. Opus 4.7 sits a few percentage points behind on each pure-reasoning test but stays competitive enough for production use.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The thing to internalize about these scores is that the practical difference between, say, 92% and 94% on GPQA Diamond is small. Once you&#8217;re above the 90% ceiling on a graduate-level reasoning test, the marginal skill the model has isn&#8217;t separating real-work outcomes. The signal moves from &#8220;which model is smarter&#8221; to &#8220;which model fits your specific reasoning structure&#8221;: chain-of-thought length preferences, tool-call reliability, and consistency across long sessions.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A practical pattern we&#8217;ve seen from running reasoning benchmarks with real enterprise prompts (more on that in the next section) is that the model that wins a benchmark isn&#8217;t always the model that wins on production reasoning tasks. EvalRig-style multi-skill harnesses, where the model has to apply reasoning in service of a 3-5 step tool-use chain, separate the genuinely useful from the benchmark-fit.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"enterprise-agentic\">Enterprise agentic coding: who wins the real-world test<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The cleanest 2026 head-to-head is <a href=\"https:\/\/evalrig.ai\/benchmarks\/frontier-3way\/\" target=\"_blank\" rel=\"noopener\">EvalRig&#8217;s three-way enterprise agentic coding test<\/a>, published in August. The methodology: 5 representative enterprise skills (code navigation across large repos, debugging, test generation, refactoring, and concurrency), each run on identical prompts across the three models with cost tracking.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The headline finding: <strong>all three hit 5\/5 pass rate<\/strong>. On raw capability, the three are effectively tied on these tasks. The differences are in cost and tool efficiency:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Gemini 3.1 Pro: $2.98 per task total<\/strong> \u2014 cheapest of the three, top or near-top quality on most skills<\/li>\n<li><strong>Claude Opus 4.7: $10.33 per task total<\/strong> \u2014 most expensive, but best quality on 4 of the 5 skills<\/li>\n<li><strong>GPT-5.4: $5.24 per task total<\/strong> \u2014 mid-cost, but uses ~1.65\u20131.72\u00d7 more tool rounds on the concurrency task than Opus or Gemini<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">The tool-round efficiency difference on the concurrency task is the most interesting finding. GPT-5.4 reaches the same answer as the other two but spends roughly 70% more tool calls doing it. That kind of inefficiency compounds fast in production: at scale, GPT-5.4 can burn through twice the API spend for the same throughput if your agent loop is measured in tens of thousands of runs per day.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/www.deeplearning.ai\/the-batch\/openais-gpt-5-4-pro-and-gpt-5-4-thinking-challenge-googles-gemini-3-1-pro-preview-as-best-all-around-ai-model\" target=\"_blank\" rel=\"noopener\">The Batch&#8217;s coverage of GPT-5.4 vs Gemini 3.1 Pro<\/a> characterizes GPT-5.4 as favored for &#8220;precision-dependent agentic tasks including multi-step tool calling, software engineering, and operating system automation.&#8221; That&#8217;s consistent with EvalRig&#8217;s finding: GPT-5.4 reaches the right answer, it just takes more tool calls to do it on certain tasks. If your agent harness charges per round-trip, you pay more.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img data-recalc-dims=\"1\" loading=\"lazy\" decoding=\"async\" width=\"2560\" height=\"1429\" src=\"https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/08\/inline-pricing-chart-scaled.jpg?resize=2560%2C1429&#038;ssl=1\" alt=\"Bar chart comparing output API cost per million tokens for the three frontier AI models: Claude Opus 4.7 at $25, GPT-5.4 at $15, and Gemini 3.1 Pro at $12\" class=\"wp-image-20722\" srcset=\"https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/08\/inline-pricing-chart-scaled.jpg?w=2560&amp;ssl=1 2560w, https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/08\/inline-pricing-chart-scaled.jpg?resize=300%2C167&amp;ssl=1 300w, https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/08\/inline-pricing-chart-scaled.jpg?resize=1024%2C572&amp;ssl=1 1024w, https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/08\/inline-pricing-chart-scaled.jpg?resize=768%2C429&amp;ssl=1 768w, https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/08\/inline-pricing-chart-scaled.jpg?resize=1536%2C857&amp;ssl=1 1536w, https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/08\/inline-pricing-chart-scaled.jpg?resize=2048%2C1143&amp;ssl=1 2048w, https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/08\/inline-pricing-chart-scaled.jpg?resize=600%2C335&amp;ssl=1 600w\" sizes=\"auto, (max-width: 1000px) 100vw, 1000px\" \/><figcaption class=\"wp-element-caption\">Output API cost per million tokens across the three 2026 flagships. Source: vendor pricing pages as of August 2026.<\/figcaption><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"pricing\">API pricing: per-million-token cost and the context-length cliffs<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">API pricing is where the three diverge most. All three vendors have tiered pricing that depends on context length and (in OpenAI&#8217;s case) the variant you choose. The 2026 list price, from each vendor&#8217;s official pricing page:<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table>\n<thead><tr><th>Model<\/th><th>Input $\/MTok<\/th><th>Output $\/MTok<\/th><th>Context-length cliff<\/th><\/tr><\/thead>\n<tbody>\n<tr><td><a href=\"https:\/\/platform.claude.com\/docs\/en\/about-claude\/pricing\" target=\"_blank\" rel=\"noopener\">Claude Opus 4.7<\/a><\/td><td>$5.00<\/td><td>$25.00<\/td><td>None \u2014 flat rate<\/td><\/tr>\n<tr><td><a href=\"https:\/\/developers.openai.com\/api\/docs\/pricing\" target=\"_blank\" rel=\"noopener\">GPT-5.4 (short)<\/a><\/td><td>$2.50<\/td><td>$15.00<\/td><td>$5.00 in \/ $0.50 cached beyond 272K input tokens<\/td><\/tr>\n<tr><td><a href=\"https:\/\/ai.google.dev\/gemini-api\/docs\/pricing\" target=\"_blank\" rel=\"noopener\">Gemini 3.1 Pro<\/a><\/td><td>$2.00<\/td><td>$12.00<\/td><td>$4.00 in \/ $18.00 out beyond 200K input tokens<\/td><\/tr>\n<\/tbody><\/table><figcaption class=\"wp-element-caption\">Per-million-token API pricing for the three flagships, August 2026. Source: vendor pricing pages. Cache pricing and batch discounts vary \u2014 see vendor docs for full details.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">The headline cost ratio for short-context work: Gemini is roughly 40% cheaper than GPT-5.4 and 60% cheaper than Opus 4.7. For long-context work beyond the vendor&#8217;s cliff, the picture flips: GPT-5.4 becomes 25% more expensive for input tokens above 272K, and Gemini 100% more expensive for input tokens above 200K. Opus 4.7 has no cliff \u2014 same flat $5\/$25 regardless of context length, which is its only pricing edge.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The cost calculus depends entirely on the input\/output ratio for your workload. A typical agentic coding prompt might be 5,000 input tokens for 500 output tokens. At those ratios, Opus 4.7 costs $0.025 + $0.0125 = $0.0375 per prompt. Gemini costs $0.01 + $0.006 = $0.016 per prompt. The 2.3\u00d7 cost ratio translates to a real bill difference at scale, especially for high-volume tools like customer support copilots or document processing pipelines. For an enterprise decision on cloud-versus-local deployment economics, the <a href=\"\/?p=20507\">Local LLM Setup 2026<\/a> comparison is relevant \u2014 Opus 4.7 in particular is expensive enough that running a quantized local model on Llama-3.1-70B is sometimes the cheaper option despite the capability gap.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"context-and-multimodal\">Context windows and multimodal capabilities<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">All three flagships are in the 1M+ context window class. The numbers:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Claude Opus 4.7:<\/strong> 1,000,000 tokens total, 128,000 tokens max output in sync API. <a href=\"https:\/\/platform.claude.com\/docs\/en\/build-with-claude\/context-windows\" target=\"_blank\" rel=\"noopener\">Anthropic&#8217;s context window docs<\/a> note that the Message Batches API extends output to 300,000 tokens with a beta header. The 1M budget is shared between input and output \u2014 important for any workload that produces long completions.<\/li>\n<li><strong>GPT-5.4:<\/strong> 1,050,000 tokens in the API. The 272K-long-context pricing threshold creates a soft cliff but no hard context-window limit. ChatGPT&#8217;s manual Thinking mode provides a 256K total context window in the consumer product, but the API has the full 1.05M.<\/li>\n<li><strong>Gemini 3.1 Pro:<\/strong> 1,048,576 tokens input, 65,536 tokens max output. The 65K output cap is a real constraint for long-form generation \u2014 but Gemini compensates with native multimodal inputs that neither Opus nor GPT can match at parity.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">The multimodal story is where Gemini 3.1 Pro still has 2026 differentiation. Per the <a href=\"https:\/\/deepmind.google\/models\/model-cards\/gemini-3-1-pro\/\" target=\"_blank\" rel=\"noopener\">official Gemini 3.1 Pro model card<\/a>, the model natively processes text, images, audio, video, PDFs, and code repositories as direct inputs \u2014 no preprocessing pipeline, no separate OCR step, no separate speech-to-text service. If your workload involves &#8220;watch this 30-minute video and answer questions about the technical content,&#8221; Gemini is the only one of the three that does this without orchestration glue code. For an architectural comparison of multimodal deployment on the device versus the cloud, see the recent <a href=\"\/?p=20705\">On-device AI 2026: Apple, Gemini Nano, and Qualcomm<\/a> piece, which covers the complementary edge-tier multimodal story.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Opus 4.7 and GPT-5.4 both support image and PDF inputs as of 2026, but audio and video still require separate preprocessing steps (often whisper for audio, frame sampling for video). If multimodal is core to your workload, Gemini 3.1 Pro is the clear production default.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"weak-points\">Where each model loses: the known weak points<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">No frontier model review is complete without the failure modes. Each of the three has known weak points in 2026:<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Opus 4.7: Anthropic&#8217;s cyber-capability scaling-back.<\/strong> Per the launch coverage on <a href=\"https:\/\/the-decoder.com\/anthropics-claude-opus-4-7-makes-a-big-leap-in-coding-while-deliberately-scaling-back-cyber-capabilities\/\" target=\"_blank\" rel=\"noopener\">the-decoder.com<\/a>, Anthropic explicitly downgraded certain offensive-security capabilities in this release. For most production engineering teams this is a non-issue. For security research teams who relied on prior versions for adversarial-analysis workflows, it&#8217;s a real loss.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>GPT-5.4: Tool-round inefficiency on concurrency.<\/strong> The <a href=\"https:\/\/evalrig.ai\/benchmarks\/frontier-3way\/\" target=\"_blank\" rel=\"noopener\">EvalRig findings<\/a> show GPT-5.4 uses 1.65\u20131.72\u00d7 more tool calls than Opus or Gemini on the concurrency task. For per-call agent frameworks that charge per tool round, GPT-5.4 can become the most expensive choice despite a competitive per-token price. The fix is usually to switch to a custom harness that batches tool calls more aggressively.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Gemini 3.1 Pro: 65,536-token output cap.<\/strong> The output limit is much smaller than Opus 4.7&#8217;s 128K sync or GPT-5.4&#8217;s 65K+ in API. For long-form content generation (technical documentation, multi-chapter reports, large code refactor outputs), Gemini 3.1 Pro requires an &#8220;output in chunks&#8221; workflow or stitching pattern. Not a deal-breaker, but worth budgeting for in production.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">These known weak points are about the deployment envelope. None of the three will produce a wrong answer that the others get right on the kinds of tasks an enterprise team is actually running \u2014 the EvalRig test confirmed 5\/5 pass rates. The weak points are operational: cost, output volume, special capability access. Those matter, but they&#8217;re solvable through workflow design, not model selection.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"conclusion\">The honest ranking and how to use it<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">If you run agentic coding at scale, the right move before committing to any of the three is to run an EvalRig-style 5-skill test on your own repository. The gap between the three on real workloads is smaller than the SWE-bench headline scores suggest, and per-task cost varies by 3-4\u00d7 between the cheapest (Gemini) and the most expensive (Opus 4.7). The benchmark that predicts your outcomes is the benchmark that uses your own data.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For short-context workloads like customer support, content summarization, and chat assistants, <strong>GPT-5.4&#8217;s $2.50\/$15 short-context tier is the default to test against first<\/strong>. The tool-round inefficiency on concurrency is irrelevant for non-agentic chat; the reasoning precision at the Pro tier is the strongest in the field.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For multimodal-heavy pipelines (document processing, video analysis, audio transcription with reasoning), <strong>Gemini 3.1 Pro&#8217;s native multimodal inputs are still a 2026 differentiator<\/strong> that Opus and GPT can&#8217;t match without orchestration glue code. The 65K output cap is the constraint; if your multimodal workflow generates long outputs, budget for chunked generation or move to Opus 4.7 for the long-form variant of the same task.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For production coding agents at any non-trivial scale, <strong>Claude Opus 4.7 still leads<\/strong>. The $10.33-per-task average on EvalRig is offset by its 4-of-5 quality wins. If your deployment is tolerating rare coding failures because they&#8217;re caught in code review, the cost difference justifies the quality. If your deployment is shipping autonomous PRs, Opus 4.7 is the safer pick.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The 2026 frontier is unusually flat for three competitors, with each model taking a clear deployment scenario. Pick by workload shape, not by headline benchmark.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"frequently-asked-questions\">Frequently asked questions<\/h2>\n\n\n\n<h3 class=\"wp-block-heading\" id=\"faq-claude-vs-gpt\">Is Claude Opus 4.7 better than GPT-5.4?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Opus 4.7 leads on SWE-bench Verified coding (87.6%) and EvalRig&#8217;s enterprise agentic coding quality (4 of 5 skills). GPT-5.4 leads on GPQA reasoning at 94.4% (Pro). For most production enterprise workloads, the gaps are smaller than vendor marketing suggests \u2014 both are excellent choices.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" id=\"faq-cheapest\">Which AI model is cheapest to run in 2026?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Gemini 3.1 Pro at $2.00 input \/ $12.00 output per MTok (under 200K context) is 40% cheaper than GPT-5.4 short-context ($2.50\/$15) and 60% cheaper than Opus 4.7 ($5\/$25) for that context window. Above 200K context, Gemini&#8217;s input pricing doubles \u2014 the relative ranking shifts.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" id=\"faq-switch-from-gpt4o\">Should I switch from GPT-4o to GPT-5.4 or Claude Opus 4.7?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Yes if you have agentic coding workloads \u2014 Opus 4.7&#8217;s SWE-bench gains justify the test. Hold off if your workloads are short-context chat or content generation where GPT-4o still serves \u2014 its deprecation is gradual and the cost-per-call improvement from GPT-5.4 over GPT-4o is not enormous for those workloads.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" id=\"faq-multimodal\">Can Gemini 3.1 Pro process video and audio in the API?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Yes. Gemini 3.1 Pro is natively multimodal with text, images, audio, video, PDFs, and code repositories as direct inputs \u2014 no preprocessing pipeline needed. This remains a 2026 differentiator vs Opus 4.7 and GPT-5.4, both of which require preprocessing steps for audio and video inputs.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" id=\"faq-mistake\">What&#8217;s the biggest mistake teams make picking a frontier model in 2026?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Optimizing on SWE-bench or GPQA alone. Those benchmarks cluster near the 90%+ ceiling in 2026 and predict practical performance poorly below the actual task. The model that wins SWE-bench isn&#8217;t always the model that wins your specific repository. Test on your own tasks with EvalRig-style multi-skill harnesses before committing production spend.<\/p>\n\n\n\n<section class=\"aimade-json-ld\" aria-label=\"Structured data\" hidden>\n<script type=\"application\/ld+json\">\n{\n  \"@context\": \"https:\/\/schema.org\",\n  \"@type\": \"TechArticle\",\n  \"headline\": \"Claude Opus 4.7 vs GPT-5.4 vs Gemini 3.1 Pro: The 2026 Frontier Model Benchmark\",\n  \"description\": \"Claude Opus 4.7 vs GPT-5.4 vs Gemini 3.1 Pro on SWE-bench, GPQA, and EvalRig. 2026 frontier is flat \u2014 deploy-by-deploy verdict with per-token API costs.\",\n  \"url\": \"https:\/\/aimade.tech\/claude-opus-4-7-vs-gpt-5-4-vs-gemini-3-1-pro-2026-benchmark\/\",\n  \"datePublished\": \"2026-08-11T15:00:00+00:00\",\n  \"dateModified\": \"2026-08-11T15:00:00+00:00\",\n  \"author\": {\n    \"@type\": \"Person\",\n    \"name\": \"AI Made Editorial\",\n    \"url\": \"https:\/\/aimade.tech\/about\/\"\n  },\n  \"publisher\": {\n    \"@type\": \"Organization\",\n    \"name\": \"AI Made\",\n    \"logo\": {\n      \"@type\": \"ImageObject\",\n      \"url\": \"https:\/\/aimade.tech\/wp-content\/uploads\/2026\/07\/aimade-logo.png\"\n    }\n  },\n  \"image\": \"https:\/\/aimade.tech\/wp-content\/uploads\/2026\/08\/hero-scaled.jpg\",\n  \"articleSection\": \"AI Models\",\n  \"keywords\": \"Claude Opus 4.7, GPT-5.4, Gemini 3.1 Pro, AI model comparison, frontier AI models 2026, SWE-bench Verified\",\n  \"wordCount\": 3070,\n  \"inLanguage\": \"en\",\n  \"about\": [\n    {\n      \"@type\": \"SoftwareApplication\",\n      \"name\": \"Claude Opus 4.7\",\n      \"applicationCategory\": \"AI Model\"\n    },\n    {\n      \"@type\": \"SoftwareApplication\",\n      \"name\": \"GPT-5.4\",\n      \"applicationCategory\": \"AI Model\"\n    },\n    {\n      \"@type\": \"SoftwareApplication\",\n      \"name\": \"Gemini 3.1 Pro\",\n      \"applicationCategory\": \"AI Model\"\n    }\n  ],\n  \"citation\": [\n    {\n      \"@type\": \"CreativeWork\",\n      \"url\": \"https:\/\/www.anthropic.com\/research\/claude-opus-4-7\",\n      \"name\": \"Anthropic Claude Opus 4.7 Research Page\"\n    },\n    {\n      \"@type\": \"CreativeWork\",\n      \"url\": \"https:\/\/openai.com\/index\/introducing-gpt-5-4\/\",\n      \"name\": \"OpenAI GPT-5.4 Introduction\"\n    },\n    {\n      \"@type\": \"CreativeWork\",\n      \"url\": \"https:\/\/deepmind.google\/models\/model-cards\/gemini-3-1-pro\/\",\n      \"name\": \"Gemini 3.1 Pro Model Card\"\n    },\n    {\n      \"@type\": \"CreativeWork\",\n      \"url\": \"https:\/\/evalrig.ai\/benchmarks\/frontier-3way\/\",\n      \"name\": \"EvalRig Three-Way Frontier Model Benchmark\"\n    }\n  ]\n}\n<\/script>\n<script type=\"application\/ld+json\">\n{\n  \"@context\": \"https:\/\/schema.org\",\n  \"@type\": \"FAQPage\",\n  \"mainEntity\": [\n    {\n      \"@type\": \"Question\",\n      \"name\": \"Is Claude Opus 4.7 better than GPT-5.4?\",\n      \"acceptedAnswer\": {\n        \"@type\": \"Answer\",\n        \"text\": \"Opus 4.7 leads on SWE-bench Verified coding (87.6%) and EvalRig's enterprise agentic coding quality (4 of 5 skills). GPT-5.4 leads on GPQA reasoning at 94.4% (Pro). For most production enterprise workloads, the gaps are smaller than vendor marketing suggests \u2014 both are excellent choices.\"\n      }\n    },\n    {\n      \"@type\": \"Question\",\n      \"name\": \"Which AI model is cheapest to run in 2026?\",\n      \"acceptedAnswer\": {\n        \"@type\": \"Answer\",\n        \"text\": \"Gemini 3.1 Pro at $2.00 input \/ $12.00 output per MTok (under 200K context) is 40% cheaper than GPT-5.4 short-context ($2.50\/$15) and 60% cheaper than Opus 4.7 ($5\/$25) for that context window. Above 200K context, Gemini's input pricing doubles \u2014 the relative ranking shifts.\"\n      }\n    },\n    {\n      \"@type\": \"Question\",\n      \"name\": \"Should I switch from GPT-4o to GPT-5.4 or Claude Opus 4.7?\",\n      \"acceptedAnswer\": {\n        \"@type\": \"Answer\",\n        \"text\": \"Yes if you have agentic coding workloads \u2014 Opus 4.7's SWE-bench gains justify the test. Hold off if your workloads are short-context chat or content generation where GPT-4o still serves \u2014 its deprecation is gradual and the cost-per-call improvement from GPT-5.4 over GPT-4o is not enormous for those workloads.\"\n      }\n    },\n    {\n      \"@type\": \"Question\",\n      \"name\": \"Can Gemini 3.1 Pro process video and audio in the API?\",\n      \"acceptedAnswer\": {\n        \"@type\": \"Answer\",\n        \"text\": \"Yes. Gemini 3.1 Pro is natively multimodal with text, images, audio, video, PDFs, and code repositories as direct inputs \u2014 no preprocessing pipeline needed. This remains a 2026 differentiator vs Opus 4.7 and GPT-5.4, both of which require preprocessing steps for audio and video inputs.\"\n      }\n    },\n    {\n      \"@type\": \"Question\",\n      \"name\": \"What's the biggest mistake teams make picking a frontier model in 2026?\",\n      \"acceptedAnswer\": {\n        \"@type\": \"Answer\",\n        \"text\": \"Optimizing on SWE-bench or GPQA alone. Those benchmarks cluster near the 90%+ ceiling in 2026 and predict practical performance poorly below the actual task. The model that wins SWE-bench isn't always the model that wins your specific repository. Test on your own tasks with EvalRig-style multi-skill harnesses before committing production spend.\"\n      }\n    }\n  ]\n}\n<\/script>\n<script type=\"application\/ld+json\">\n{\n  \"@context\": \"https:\/\/schema.org\",\n  \"@type\": \"WebPage\",\n  \"name\": \"Claude Opus 4.7 vs GPT-5.4 vs Gemini 3.1 Pro: The 2026 Frontier Model Benchmark\",\n  \"url\": \"https:\/\/aimade.tech\/claude-opus-4-7-vs-gpt-5-4-vs-gemini-3-1-pro-2026-benchmark\/\",\n  \"inLanguage\": \"en\",\n  \"isPartOf\": {\n    \"@type\": \"WebSite\",\n    \"name\": \"AI Made\",\n    \"url\": \"https:\/\/aimade.tech\/\"\n  },\n  \"speakable\": {\n    \"@type\": \"SpeakableSpecification\",\n    \"xpath\": [\n      \"\/html\/head\/title\",\n      \"\/html\/body\/\/p[1]\"\n    ],\n    \"cssSelector\": [\n      \".wp-block-paragraph:first-of-type\"\n    ]\n  }\n}\n<\/script>\n<script type=\"application\/ld+json\">\n{\n  \"@context\": \"https:\/\/schema.org\",\n  \"@type\": \"ClaimReview\",\n  \"url\": \"https:\/\/aimade.tech\/claude-opus-4-7-vs-gpt-5-4-vs-gemini-3-1-pro-2026-benchmark\/#enterprise-agentic\",\n  \"claimReviewed\": \"All three flagship AI models (Claude Opus 4.7, GPT-5.4, Gemini 3.1 Pro) hit 5\/5 pass rate on the EvalRig 5-skill enterprise agentic coding test.\",\n  \"author\": {\n    \"@type\": \"Organization\",\n    \"name\": \"EvalRig\",\n    \"url\": \"https:\/\/evalrig.ai\/\"\n  },\n  \"datePublished\": \"2026-08-11\",\n  \"appearanceUrl\": \"https:\/\/evalrig.ai\/benchmarks\/frontier-3way\/\",\n  \"reviewRating\": {\n    \"@type\": \"Rating\",\n    \"ratingValue\": 5,\n    \"bestRating\": 5,\n    \"worstRating\": 1,\n    \"alternateName\": \"PASS_RATE\"\n  }\n}\n<\/script>\n<\/section>\n\n","protected":false},"excerpt":{"rendered":"<p>Claude Opus 4.7 vs GPT-5.4 vs Gemini 3.1 Pro on SWE-bench, GPQA, and EvalRig. 2026 frontier is flat \u2014 deploy-by-deploy verdict with per-token API costs.<\/p>\n","protected":false},"author":0,"featured_media":20721,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_jetpack_newsletter_access":"","_jetpack_dont_email_post_to_subs":false,"_jetpack_newsletter_tier_id":0,"_jetpack_memberships_contains_paywalled_content":false,"_jetpack_feature_clip_id":0,"_jetpack_memberships_contains_paid_content":false,"footnotes":"","jetpack_publicize_message":"","jetpack_publicize_feature_enabled":true,"jetpack_social_post_already_shared":true,"jetpack_social_options":{"image_generator_settings":{"template":"highway","default_image_id":0,"font":"","enabled":false},"version":2},"jetpack_post_was_ever_published":false},"categories":[297],"tags":[496,493,497,495,494],"class_list":["post-20723","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-models","tag-ai-model-comparison","tag-claude-opus-4-7","tag-frontier-ai-models-2026","tag-gemini-3-1-pro","tag-gpt-5-4"],"jetpack_publicize_connections":[],"jetpack_sharing_enabled":true,"jetpack-related-posts":[{"id":1604,"url":"https:\/\/aimade.tech\/?p=1604","url_meta":{"origin":20723,"position":0},"title":"AI Models in April 2026: Every Major Release, Leak, and What Comes Next","author":"Mr. Technology","date":"April 11, 2026","format":false,"excerpt":"AI MODELS AI Models in April 2026: Every Major Release, Leak, and What Comes Next By Mr. Technology | April 11, 2026 Hey guys, Mr. Technology here. Buckle up. The AI model race just hit another gear, and April 2026 might be the most consequential month yet. \u2605 What You\u2026","rel":"","context":"In &quot;AI Models&quot;","block_context":{"text":"AI Models","link":"https:\/\/aimade.tech\/?cat=297"},"img":{"alt_text":"","src":"https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/04\/openai-superapp-cover.jpg?fit=1024%2C1024&ssl=1&resize=350%2C200","width":350,"height":200,"srcset":"https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/04\/openai-superapp-cover.jpg?fit=1024%2C1024&ssl=1&resize=350%2C200 1x, https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/04\/openai-superapp-cover.jpg?fit=1024%2C1024&ssl=1&resize=525%2C300 1.5x, https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/04\/openai-superapp-cover.jpg?fit=1024%2C1024&ssl=1&resize=700%2C400 2x"},"classes":[]},{"id":20635,"url":"https:\/\/aimade.tech\/?p=20635","url_meta":{"origin":20723,"position":1},"title":"Claude vs GPT-5 for Code Review: 2026 Engineering Benchmark","author":"Lucy Monday","date":"July 21, 2026","format":false,"excerpt":"The split is consistent enough to feel structural. Across a 133-cycle real-world comparison, GPT-5.2 caught 86.7% of bugs in production pull requests while Claude Opus 4.6 caught 20% \u2014 but every single finding Claude raised was real, against three false positives from GPT. Neither model is \"better\" for code review\u2026","rel":"","context":"In &quot;AI Reviews&quot;","block_context":{"text":"AI Reviews","link":"https:\/\/aimade.tech\/?cat=313"},"img":{"alt_text":"","src":"https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/07\/aimade_hero-scaled.jpg?fit=1200%2C670&ssl=1&resize=350%2C200","width":350,"height":200,"srcset":"https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/07\/aimade_hero-scaled.jpg?fit=1200%2C670&ssl=1&resize=350%2C200 1x, https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/07\/aimade_hero-scaled.jpg?fit=1200%2C670&ssl=1&resize=525%2C300 1.5x, https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/07\/aimade_hero-scaled.jpg?fit=1200%2C670&ssl=1&resize=700%2C400 2x, https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/07\/aimade_hero-scaled.jpg?fit=1200%2C670&ssl=1&resize=1050%2C600 3x"},"classes":[]},{"id":20485,"url":"https:\/\/aimade.tech\/?p=20485","url_meta":{"origin":20723,"position":2},"title":"Gemini 2.5 Pro vs GPT-4.5 vs Claude 3.7 Sonnet: The Definitive Model Rankings for 2026","author":"Lucy Monday","date":"May 11, 2026","format":false,"excerpt":"A comprehensive, no-nonsense comparison of the three leading AI models in 2026 \u2014 benchmark results, real-world performance, pricing, and which use cases each dominates.","rel":"","context":"In &quot;AI Models&quot;","block_context":{"text":"AI Models","link":"https:\/\/aimade.tech\/?cat=297"},"img":{"alt_text":"AI model rankings \u2014 LLM leaderboard 2026","src":"https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/05\/img-04-model-rankings.png?fit=1200%2C670&ssl=1&resize=350%2C200","width":350,"height":200,"srcset":"https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/05\/img-04-model-rankings.png?fit=1200%2C670&ssl=1&resize=350%2C200 1x, https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/05\/img-04-model-rankings.png?fit=1200%2C670&ssl=1&resize=525%2C300 1.5x, https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/05\/img-04-model-rankings.png?fit=1200%2C670&ssl=1&resize=700%2C400 2x, https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/05\/img-04-model-rankings.png?fit=1200%2C670&ssl=1&resize=1050%2C600 3x"},"classes":[]},{"id":20148,"url":"https:\/\/aimade.tech\/?p=20148","url_meta":{"origin":20723,"position":3},"title":"Anthropic Releases Claude Opus 4.7, a Less Risky Model than Mythos","author":"Lucy Monday","date":"April 22, 2026","format":false,"excerpt":"# Anthropic Releases Claude Opus 4.7, a Less Risky Model than Mythos *By Monday \u00a0|\u00a0 April 22, 2026* *AI MODELS* --- > **Bottom Line:** Claude Mythos Preview is Anthropic's most powerful AI model that excels at identifying weaknesses and security flaws within software. ## What Else Is Happening ### Project\u2026","rel":"","context":"In &quot;AI Models&quot;","block_context":{"text":"AI Models","link":"https:\/\/aimade.tech\/?cat=297"},"img":{"alt_text":"AI model rankings \u2014 LLM leaderboard 2026","src":"https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/05\/img-04-model-rankings.png?fit=1200%2C670&ssl=1&resize=350%2C200","width":350,"height":200,"srcset":"https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/05\/img-04-model-rankings.png?fit=1200%2C670&ssl=1&resize=350%2C200 1x, https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/05\/img-04-model-rankings.png?fit=1200%2C670&ssl=1&resize=525%2C300 1.5x, https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/05\/img-04-model-rankings.png?fit=1200%2C670&ssl=1&resize=700%2C400 2x, https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/05\/img-04-model-rankings.png?fit=1200%2C670&ssl=1&resize=1050%2C600 3x"},"classes":[]},{"id":20671,"url":"https:\/\/aimade.tech\/?p=20671","url_meta":{"origin":20723,"position":4},"title":"Small language models in 2026: when 7B beats 70B","author":"","date":"July 30, 2026","format":false,"excerpt":"Small language models in 2026 \u2014 when 7B beats 70B, with the cost-adjusted benchmark of Llama-3.1-8B vs GPT-4o across 11 enterprise tasks. The 2026 cutoff.","rel":"","context":"In &quot;AI Models&quot;","block_context":{"text":"AI Models","link":"https:\/\/aimade.tech\/?cat=297"},"img":{"alt_text":"","src":"","width":0,"height":0},"classes":[]},{"id":20049,"url":"https:\/\/aimade.tech\/?p=20049","url_meta":{"origin":20723,"position":5},"title":"Anthropic Releases Claude Opus 4.7: A Smarter, Safer Model for Real Work","author":"Mr. Technology","date":"April 21, 2026","format":false,"excerpt":"What happened: Anthropic released Claude Opus 4.7 on April 16, 2026 \u2014 its most capable generally available model to date and a notable course-correction after the controversial Claude Mythos preview the same week. Opus 4.7 lifts performance across the board: SWE-bench Verified hit 87.6%, software engineering tasks are more reliable\u2026","rel":"","context":"In &quot;Tools &amp; Resources&quot;","block_context":{"text":"Tools &amp; Resources","link":"https:\/\/aimade.tech\/?cat=8"},"img":{"alt_text":"","src":"https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/04\/claude-opus-47-release-2.jpg?fit=1200%2C675&ssl=1&resize=350%2C200","width":350,"height":200,"srcset":"https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/04\/claude-opus-47-release-2.jpg?fit=1200%2C675&ssl=1&resize=350%2C200 1x, https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/04\/claude-opus-47-release-2.jpg?fit=1200%2C675&ssl=1&resize=525%2C300 1.5x, https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/04\/claude-opus-47-release-2.jpg?fit=1200%2C675&ssl=1&resize=700%2C400 2x, https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/04\/claude-opus-47-release-2.jpg?fit=1200%2C675&ssl=1&resize=1050%2C600 3x"},"classes":[]}],"jetpack_featured_media_url":"https:\/\/i0.wp.com\/aimade.tech\/wp-content\/uploads\/2026\/08\/hero-scaled.jpg?fit=2560%2C1429&ssl=1","_links":{"self":[{"href":"https:\/\/aimade.tech\/index.php?rest_route=\/wp\/v2\/posts\/20723","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/aimade.tech\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/aimade.tech\/index.php?rest_route=\/wp\/v2\/types\/post"}],"replies":[{"embeddable":true,"href":"https:\/\/aimade.tech\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=20723"}],"version-history":[{"count":2,"href":"https:\/\/aimade.tech\/index.php?rest_route=\/wp\/v2\/posts\/20723\/revisions"}],"predecessor-version":[{"id":20725,"href":"https:\/\/aimade.tech\/index.php?rest_route=\/wp\/v2\/posts\/20723\/revisions\/20725"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/aimade.tech\/index.php?rest_route=\/wp\/v2\/media\/20721"}],"wp:attachment":[{"href":"https:\/\/aimade.tech\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=20723"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/aimade.tech\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=20723"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/aimade.tech\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=20723"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}