Tôi Là Tùng
Back to Blog

What Is System 1 in an AI Agent? The Cheap Reflex Layer

System 1 is the fast, low-cost reflex layer of an AI Agent: intent classification, routing and gating before a large model is called. With listed token prices.

What Is System 1 in an AI Agent? The Cheap Reflex Layer | Tôi là Tùng, toilatung, Nguyễn Thanh Tùng, Tùng Sóc Sơn

TL;DR: System 1 in an AI Agent is the fast, low-cost reflex layer. It uses a small model for intent classification, routing, raw data extraction and syntax gating, and passes work to System 2 (a large reasoning model) only when it meets complex logic. The cost figures in this article come from a hypothetical calculation based on listed prices, not from measurements of a running system.

When an agent receives every request, from classifying a message to planning many steps, and sends all of it to the same large model, the token bill grows with the number of requests rather than with their difficulty. Most requests in real operation are short and repetitive. This article explains what System 1 is, where it sits in the architecture, and how to estimate cost honestly, that is, only from listed prices and clearly stated assumptions.

What is System 1 in an AI Agent?

System 1 is the agent's reflex layer: a small, fast, low-cost model that handles short decisions with a narrow scope before any large model is called. System 2 is the slower and more expensive reasoning layer, activated only when System 1 decides a request needs planning or complex synthesis.

Typical System 1 jobs:

  • Intent classification: is this request an information question, a booking, a complaint or content creation.
  • Prompt routing: pick the right agent or tool for the request.
  • Raw data extraction: pull a name, phone number or order code from free text.
  • Syntax gating: check whether the input has the right format or signs of prompt injection before it moves on.

Why borrow Kahneman's psychology model?

Because Daniel Kahneman's split between fast and slow describes fairly well how to divide responsibility between model tiers. In Thinking, Fast and Slow (2011), Kahneman calls System 1 fast, automatic, low-effort thinking, and System 2 slow, deliberate, effortful thinking.

This is a design analogy, not a claim that language models work like the human brain. The value of the analogy is that it forces the designer to answer a practical question: which decisions are simple enough for a cheap reflex, and which truly need reasoning. The article Dual-Engine AI Agent: lessons from System 1 and System 2 checks this split against the source code of a specific system and notes its limits.

How do System 1 and System 2 differ?

System 1 optimizes for speed and price per narrow decision. System 2 optimizes for reasoning quality on multi-step tasks. The table takes unit prices from the providers' official pricing pages, checked on 10 October 2026. Prices can change, so verify again before budgeting.

TierExample modelInput (USD per million tokens)Output (USD per million tokens)Typical job
System 1Claude Haiku 5.5 (prompts up to 100,000 tokens)0.100.50Classify, route, extract
System 1Gemini 2.5 Flash-Lite0.100.40Classify, route, extract
System 1DeepSeek-V4.1-Flash (cache miss, off-peak)0.150.60Classify, short summaries
System 2Claude Sonnet 5.5210Planning, long writing, synthesis
System 2Claude Opus 5.5420Multi-step reasoning, complex context
System 2Gemini 3.1 Pro Preview (prompts up to 200,000 tokens)212Reasoning, long analysis

On latency, the providers above do not commit to a fixed figure on their pricing pages, so there is no latency column. Measure end-to-end p50 and p95 latency on your own workload.

Price sources: Anthropic, Google Gemini API, DeepSeek API. DeepSeek charges peak-hour rates at twice the off-peak rates, and older model names have been replaced by new ones on its pricing page.

How does a System 1 router work?

The router receives a request, returns a label with a confidence level, and escalates to System 2 only when the label is in the complex group or confidence is low. When unsure, the default must be to escalate, not to handle it alone.

The code below is a sketch to illustrate the logic, not code running in a specific system:

type Route = "simple" | "complex" | "unsafe";

async function route(input: string) {
  const { label, confidence } = await system1.classify(input); // small model

  if (label === "unsafe") return reject(input);       // gating
  if (label === "simple" && confidence >= 0.8) {
    return system1.handle(input);                      // handled by reflex
  }
  return system2.handle(input);                        // escalate when complex or unsure
}

The 0.8 threshold in the example is an illustrative value. A real threshold must be calibrated on a labeled set of your own requests. Deterministic rules such as access checks or size limits belong in ordinary code, with no model call.

How does cost change with System 1?

In the hypothetical scenario below, a routed architecture cuts token cost by roughly 75–80% compared with running everything through Claude Sonnet 5.5. This is the result of a calculation from stated assumptions, not a measurement from real traffic.

Scenario assumptions:

  • One month has 10,000 requests. Each request is 1,000 input tokens and 300 output tokens.
  • Baseline: 100% of requests go through Claude Sonnet 5.5 (2 USD input, 10 USD output per million tokens).
  • Routed option: every request goes through Claude Haiku 5.5 first. System 1 handles 80% or 85% of requests itself. The rest go to Claude Sonnet 5.5. For escalated requests, System 1 generates only about 50 output tokens for the label.
  • Not counted: cache storage fees, retry cost, monitoring cost and design effort.
OptionTotal token cost per month (USD)Versus all through Sonnet 5.5
100% through Claude Sonnet 5.550.00Baseline
System 1 handles 80%, 20% go to Sonnet 5.512.25About 75.5% lower
System 1 handles 85%, 15% go to Sonnet 5.59.81About 80.4% lower

How to read this correctly: the 80–85% rate is an assumption, and the quality of System 1 on real tasks decides whether it holds. If System 1 mishandles part of the requests, the cost of fixing errors and business damage can exceed the money saved. This is the point made in cheap AI models are not always cheap to run: a token price is not yet the cost of an accepted task.

Once real measurements exist, the table should add end-to-end p50 and p95 latency, fallback rate, misclassification rate and cost per accepted task.

What risks does System 1 carry and how are they controlled?

The main risk is misclassification, letting a complex request slip into the cheap path. Reduce it by running in observation mode first, logging every escalation, and placing an approval gate at consequential steps.

  • Run shadow mode first: let System 1 classify in parallel without authority to act, then compare its labels with known correct results.
  • Escalate by default when unsure: low confidence goes up to System 2.
  • Record fallbacks: each escalation is logged to measure the rate and to notice when System 1 starts to degrade.
  • Keep a human approval gate: steps with hard-to-reverse consequences still need a person, however confident System 1 is. See Zero Trust for AI Agents.

The System 1 layer is one component of the wider architecture described in What is an AI Business OS, where it acts as a gate in front of specialized agents.

When should you not use System 1?

When request volume is too small, when every request is complex, or when there is no labeled dataset to check against. An extra tier pays off only if most requests are truly simple and repetitive.

With a few hundred requests per month, the token price gap is usually smaller than the effort of building and testing a router. When every request needs long context and multi-step reasoning, System 1 only adds latency and failure points. Without labeled data you have no way to check whether System 1 is right or wrong, and building this tier is guesswork.

Frequently asked questions (FAQ)

Does System 1 in an AI Agent fully replace System 2?

No. System 1 handles short, narrow decisions at low cost. Tasks that need multi-step planning or long-context synthesis still need System 2. The two tiers complement each other and the router decides which path each task takes.

Which model should I use as System 1?

There is no general answer. Try a few small models, such as Claude Haiku, Gemini Flash-Lite or DeepSeek Flash, on the same labeled set of your own requests, then choose by classification accuracy, end-to-end latency and cost per accepted task. A price table is only one input to the decision.

Does System 1 really cut cost by 80%?

The 75–80% figure in this article is the result of a calculation scenario that assumes System 1 handles 80–85% of requests itself. The real reduction depends on that rate, on classification quality and on your workload. Only measurements on real traffic can confirm it.

Should System 1 be tested before it is allowed to act?

Yes. Run it in observation mode (shadow mode), letting System 1 classify in parallel and comparing with known correct results, before granting authority to act. Every escalation should be logged to measure the fallback rate.

Conclusion

System 1 applies a simple idea: do not use an expensive tool for a cheap job. Its value lies in answering which decisions are simple enough for a reflex, with an escalation mechanism for when it is unsure. This article uses listed prices and a hypothetical scenario to show the possible scale of the difference, without measurements from a real system. I will publish performance numbers only once the measurement set is clean enough to stand behind.

Lead Magnet Special Edition

Nhận Bộ Thư Viện Prompt & SOP AI Workflow Vận Hành Doanh Nghiệp 2026

Tặng miễn phí Ebook PDF + Notion Template quản lý AI System thực chiến từ Tôi Là Tùng. Gửi trực tiếp vào hòm thư công việc của bạn.

Bảo mật 100%• Nhận file PDF & Notion• Hủy đăng ký 1-Click
🎁 Miễn Phí 100%

Tặng File Cấu Hình AI Stack Tự Host (n8n + Flowise + pgvector)

Docker-compose sẵn dùng để tự dựng hạ tầng AI của riêng bạn trong vài phút — không cần trả phí SaaS hàng tháng.

Nguyễn Thanh Tùng — AI System Designer
Written by Tùng
Nguyễn Thanh Tùng · AI Director