Cisco Routed 50-60% of 90,000 Agent Requests to Open Models
When Cisco handed a personal AI agent to all 90,000 of its employees on August 27, 2026, the headline number was the headcount. The number that should interest anyone writing an AI budget is 50 to 60.
That is the share of requests to MyAgent — Cisco's new company-wide agent — that the company says are served by open-weight models running on GPUs Cisco owns. Another 20% to 30% of requests never reach a generative model at all; they are handled by ordinary software automation. Only a residual share is sent to frontier APIs such as Azure OpenAI, Anthropic's Claude, and Google's Gemini.
It is a deliberate inversion of how most enterprises have deployed AI over the past two years. The default pattern has been to pick a frontier provider, wire the whole company into it, and watch the invoice grow. Cisco built the routing layer first and treated the frontier model as an escalation path rather than a default.
The Architecture: A Router, Not a Model Contract
MyAgent is not a new model. It is an execution layer sitting on Circuit, the internal AI platform Cisco has been building since 2023. Circuit had already reached more than 100,000 users and roughly 90% employee adoption before MyAgent shipped, which is why the agent rollout could go company-wide on day one rather than as a pilot.
The orchestration layer decides where each request goes based on the nature of the task, acceptable latency, and the reliability the job demands. Three destinations:
- Self-hosted open-weight models on Cisco's own GPU servers — the majority path - Deterministic software automation for anything that does not need a language model: ticket creation, threshold alerts, scheduled checks - External frontier APIs — Azure OpenAI, Claude, Gemini, plus Cisco's own Deep Network Model — reserved for the hard remainder
Each agent takes instructions only from the employee it belongs to; there is no agent-to-agent chatter between colleagues. Context is pulled through permissioned connectors that respect existing access rights across Outlook, Webex, Jira, and SharePoint. Any action with an effect outside the system requires explicit human sign-off, and a policy server can block destructive operations and prevent Cisco data from being used to train third-party models.
A single personal agent can call more than 800 backend subagents. Agentic interactions on Circuit grew roughly 350% quarter over quarter in the period before the MyAgent launch.
Why the Split Makes Financial Sense
The economics behind Cisco's mix are straightforward. Agentic workloads consume orders of magnitude more tokens than chat, because an agent does not generate one response — it generates many, across tool calls, reasoning steps, and retries. Routing that volume to a $10-per-million-token frontier API by default produces a bill that scales with agent activity, not with headcount.
Cisco's CFO, Mark Patterson, described the architecture in Fortune as a lever for budget discipline. "This won't burn a ton of tokens with frontier models," he said. Unconfirmed estimates cited in coverage point to an annual token bill of roughly $900 million, or about $200 per employee per week — though Cisco has not published official cost figures, error rates, or a productivity audit for MyAgent.
The pattern matches independent production data. Vercel's AI Gateway index put open-weight models at roughly 62% of token volume but under 9% of spend in late August 2026. High-volume routine work is where cheap models pay off, and Cisco's routing ensures that routine work does not subsidize frontier pricing.
What Cisco Actually Uses It For
The use cases Cisco's leadership highlights are concentrated in high-value, document-heavy functions. Patterson said that between 80% and 90% of the first draft of the Management's Discussion and Analysis (MD&A) section of SEC filings is now AI-generated. His personal agent runs a "CFO cockpit" dashboard that aggregates performance by product, geography, and customer segment and suggests corrective actions. A separate agent automatically compares Cisco's performance to peers on revenue growth, earnings per share, and R&D expenditure.
For the broader workforce, the primary uses are email summarization, spam filtering, reply drafting, and cross-system task coordination across Outlook, Webex, Jira, and SharePoint.
Token consumption is tracked by a Splunk supervising tool, which Cisco acquired back in 2023.
Governance and the Supervision Question
Cisco describes MyAgent as moving "beyond chat-based interaction toward supervised autonomous workflows." The word "supervised" is doing a lot of work in that sentence, and the public announcement leaves the most operational question unanswered: where exactly do the review or approval checkpoints occur as a task moves between applications and data sources?
The stated controls are specific on paper but unmeasured in practice. Each agent instance is strictly personal and receives instructions only from the employee it belongs to. Any action with an effect outside the system requires explicit human validation. A policy server can block destructive operations and prevent Cisco data from being used to train third-party models. Cisco's AI Defense banner monitors interactions between users, models, and agents for prompt manipulation, malicious behavior, and unauthorized data transfers.
What Cisco has not published is how often those controls are actually invoked. If approval prompts fire on every cross-system action, the agent's value proposition collapses back to a chatbot with extra steps. If they fire rarely, the "supervised" label rests on trust rather than evidence. The company has not published workforce-wide outcome measures — not error rates, not time saved, not incidents caught or missed.
This is not unique to Cisco. It is the central gap in enterprise agent deployments across the industry: vendors describe governance architectures in detail but rarely publish the data that would show whether they work. Until that data exists, every "supervised autonomous" claim is an architecture diagram, not an outcome.
The Takeaway for Buyers
Cisco's stated routing percentages leave 10% to 30% of traffic unaccounted for, and the company has not disclosed error rates or outcome measures. So the architecture is the takeaway, not the ROI.
The core insight is unglamorous but durable: the routing layer and the model layer are separate decisions. Settle routing first, and every model choice underneath stays reversible. That matters in a market where the price and capability leaderboard reshuffles quarterly — and where the difference between a $900 million bill and a $90 million one is not which model you picked, but which fraction of your traffic you sent to it.
For enterprises planning an agent rollout, the Cisco case suggests a specific sequence: build the governed routing layer before committing to a model contract, reserve frontier APIs for the tasks that actually need them, and treat the open-weight and automation paths as the default rather than the fallback.
The sequencing also matters for vendor negotiations. An enterprise that has already built its own routing layer can benchmark each model provider on price and latency for the specific slice of traffic it routes to them, rather than accepting a single bundled rate for everything. That changes the buyer's leverage from the start.
The headline for most companies will not be 90,000 agents. It will be how much of their traffic never needed a frontier model in the first place.
Editorial sources
Every claim in this briefing traces back to the references below.
- Cisco MyAgent rollout coverage (Thimaya Subaiya, Cisco EVP Operations, Aug 27 2026) — Reports Cisco's company-wide MyAgent deployment, Circuit platform, and routing architecture https://superpowerdaily.com/posts/cisco-rolls-out-myagent-to-90-000-employees-for-supervised-cross-app-work
- Cisco Sends Most 90,000-Employee AI Agent Traffic to Open Models — Provides the 50-60% open-weight, 20-30% automation routing figures and Vercel AI Gateway benchmark https://ai2.work/blog/cisco-sends-most-90-000-employee-ai-agent-traffic-to-open-models
- How Cisco Empowered 90,000 Employees with an AI Assistant — Reports CFO Mark Patterson's comments on cost discipline, MD&A AI generation, and Splunk token tracking https://www.dawnliphardt.com/how-cisco-empowered-90000-employees-with-an-ai-assistant/