back to blog

Hybrid AI: Routing work between frontier APIs and open weight LLMs

Read Time 9 mins | Written by: Cole

Hybrid AI: Routing work between frontier APIs and open weight LLMs

AT&T post-trained its own telecom model on 400 billion tokens, and it still calls Anthropic and OpenAI APIs every day. About 40% of employee AI queries now run on open models, and AT&T is targeting 60% to 70%. That's hybrid AI in practice.

Frontier models from Anthropic and OpenAI take the work that needs the strongest reasoning available. Open-weight models, on AT&T's own GPUs or a cloud endpoint, take the high-volume work that doesn't. A gateway in front of both decides which model gets each task.

"When you are processing hundreds of billions of tokens, infrastructure becomes part of the problem you solve," said Mark Austin, VP of Data Science and AI at AT&T.

The models in that mix will change every quarter. The gateway that decides which one gets each task is the piece to build yourself.

Three tiers of hybrid AI architecture

hybrid-ai-three-tiers-body

Most teams say "hybrid" and picture two options – our GPUs or their API. There are three, and conflating the first two wrecks more business cases than any model choice.

  • Self-hosted. Your GPUs, your inference engineering – vLLM, SGLang, Triton. Full control of where data sits.

  • Hosted open weight models. The same open model on someone else's endpoint: Together, Fireworks, DeepInfra, Groq, Baseten, Workers AI. Identical weights, identical license, none of the operations.

  • Frontier APIs. Anthropic, OpenAI, Google.

AT&T runs all three at once. It built OTel 2.0 across roughly 530 GPUs, most of them AMD Instinct MI300X, and moves about 45 billion tokens a day. Its post-trained telecom model sits in tier one, Phi-4, gpt-oss-120b, and Gemma-4 come through Azure in tier two, and the frontier labs cover what the first two can't.

Cloudflare runs the same three tiers and lands nowhere near the same ratio. Over a trailing 30 days, 91% of its internal AI requests went to OpenAI, Anthropic, and Google, and under 9% went to Workers AI, its own open-weight platform, mostly for documentation review. The ratio is a workload question with no universal answer, and the architecture is the part that transfers.

Teams that announce they're "going local" almost always want tier two and find out late. Cost starts most hybrid conversations, and tier two answers it without buying a single GPU. Tier one earns its keep when data cannot leave your perimeter – regulated records, customer data under residency rules, code you can't send to a vendor's inference path. Choosing tier one for cost alone runs into the next section.

The price spread is what makes tiering worth the engineering. GPT-5.6 Sol runs $5 per million input tokens and $30 output. DeepSeek V4 Flash on Together costs $0.14 and $0.28. How much of that gap you capture depends on model routing for defined workflows.

The GPU utilization math behind self-hosted LLM costs

hybrid-ai-gpu-breakeven

Self-hosting gets pitched as a capex decision. It's a utilization decision, and the threshold is higher than most teams model.

In December 2025, the vLLM team published a throughput figure from multi-node serving of a DeepSeek-class mixture-of-experts model: 2,200 output tokens per second per H200 GPU, sustained, using wide expert parallelism and dual-batch overlap. Their earlier benchmarks ran closer to 1,500 per GPU.

Price that against Together's published H200 on-demand rate of $5.99 per GPU-hour:

  • At 2,200 tokens/sec, one GPU-hour produces 7.92 million output tokens – $0.76 per million output tokens at full utilization
  • At 1,500 tokens/sec, 5.4 million tokens per GPU-hour – $1.11 per million

Now compare renting the same model instead of running it. DeepSeek V4 Pro costs $3.96 per million output tokens on Together.

You need those GPUs roughly 19% busy, around the clock, to break even against simply buying the same model from a hosted endpoint. At pre-optimization throughput, 28%.

That threshold assumes inference engineering on par with vLLM's reference deployment, and counts nothing for ops engineers, the InfiniBand fabric, or the hours your fleet sits idle. Reserved capacity and a steady 24/7 workload pull it back down – the math turns on how consistently you keep the fleet fed.

Together, Fireworks, Groq, and Baseten already built that engineering and spread it across thousands of tenants. Tier one bets your utilization beats theirs.

How to decide which AI tasks go to open models

hybrid-ai-task-routing

Every task sorts on three questions: where its data can go, how much volume it carries, and how much quality loss it tolerates.

Data can't leave your perimeter → self-hosted. Regulated records and residency-bound data decide the tier before cost does.

High volume, tolerant of a small quality drop → hosted open-weight. Classification, extraction, summarization, documentation, first-pass drafts. Most of the savings live here.

Complex reasoning, agentic, or high-stakes → frontier API. Multi-step agent work, security review, and anything customer-facing where a wrong answer costs more than the tokens.

An eval gate before any task moves down a tier. Cheap routing looks like a win on your billing dashboard within hours. The quality loss shows up nowhere unless you measure it on your own workloads first.

What belongs in a custom AI harness

hybrid-ai-harness-build-vs-buy

The harness is the code around the model – the gateway, the tool wiring, the context assembly, the guardrails. In a hybrid setup it's also whatever survives a model swap, so build it with all three tiers behind one interface.

One gateway, OpenAI-compatible. Cloudflare's reason for routing everything through a single proxy Worker from day one: "centralizing through a Worker meant we could add per-user attribution, model catalog management, and permission enforcement later without touching any client configs." That's the whole argument for owning the boundary.

Task-to-tier policy as configuration. AT&T's gateway matches each task to a model, and moving a task between tiers should be a config change rather than a deploy. That's how a 40% open-model share becomes 70% without a rewrite.

Context and tool assembly. Cloudflare collapsed 34 GitLab tools into 2 through its MCP Server Portal, cutting roughly 15,000 tokens of schema overhead to near zero. Tool surface area is a cost line.

Redaction and audit at the boundary, so no feature team has to remember it – see how Uber's AI and MCP gateways are structured.

Buy the rest. Inference serving, gateway plumbing (LiteLLM, Kong AI Gateway, Envoy AI Gateway), sandboxing, and observability all have mature options, and your version will be worse. Stripe makes that case better than any argument – it built its own internal coding-agent stack, then acquired OpenRouter in August 2026 rather than build the routing layer too.

Why open models stall before production at enterprise scale

 

hybrid-ai-open-model-production-gap_1

Mozilla's first State of Open Source AI report, published with SlashData in July 2026, found that half of developers adding AI features already use both open and closed models. The arrangement needs no selling. Shipping it is the problem.

53% of open-model projects reach production, against 63% of closed-model projects – and that gap widens with company size. Closed-model production rates climb from 54% at small companies to 73% at enterprises. Open models barely move, 53% to 57%.

Large organizations get worse at shipping open models as they grow.

The pressure runs the other way at the same time. McKinsey's 2026 State of AI survey found one in five organizations limiting AI use because of operating costs. Cost pushes teams toward the cheaper tier while operational maturity keeps them from shipping it.

That gap is a harness problem, and the strongest argument for building the layer on purpose instead of inheriting whatever ships with your models.

Custom harnesses depreciate faster than you can amortize them

hybrid-ai-harness-depreciation

Most vendors skip something important: the harness you build this quarter has a short half-life, and planning for that changes what you build.

Manifest shipped an LLM router in March 2026 that sorted requests into four complexity tiers. They deprecated it in June and shut it off on September 1, 2026 – six months, start to finish. Bruno Perez's postmortem landed on a structural problem rather than an implementation one: "The prompt alone does not contain the whole task; it is just the trigger."

So build the thin, durable parts – the boundary, the policy, the evals, the audit trail. Skip the thick parts the model vendors are actively commoditizing. A harness you can throw away in a quarter without losing your governance is the design target.

Where engineering leaders should start with hybrid AI

hybrid-ai-where-to-start

Every team building internal AI tooling ends up with a hybrid split, planned or not. AT&T's 40% and Cloudflare's 9% are both what it looks like decided on purpose – the cheap tier on work that tolerates it, the expensive tier on work that doesn't.

  1. Build the boundary first. One gateway, OpenAI-compatible, every internal AI call through it. Do this before you pick models, because it's what makes every later decision reversible.
  2. Find the boring, high-volume task. Classification, extraction, documentation, summarization, first-pass drafting. Cloudflare's 9% is documentation review. AT&T started with development-stage work before touching anything customer-facing.
  3. Move that one task to a hosted open-weight endpoint. Same weights, no GPU commitment, reversible in an afternoon.
  4. Price tier one last, against the 19% utilization threshold rather than frontier API rates, and only after you know which data actually has to stay inside.

AMD's account of what it took AT&T to get there is worth the read if you want to learn more.

Codingscape builds the harness layer – the gateway, the tool wiring, and the internal tools and production AI around it – so the model choice stays a config change instead of a rebuild. Let’s talk if you’re ready to build hybrid AI at your company.

Don't Miss
Another Update

Subscribe to be notified when
new content is published
Cole

Cole is Codingscape's Content Marketing Strategist & Copywriter.