← Andy's AI Field Notes

Field guide · ecosystem map

The AI stack, without the fog

How harnesses, agents, tools, data, models, and inference engines fit together

When I first tried to understand modern AI, the hardest part was not any single concept. It was the pile of names: ChatGPT, Codex, Claude Code, Cursor, MCP, LangGraph, Ollama, vLLM, frontier models, open-weight models, RAG, embeddings, memory, agents, subagents. They sounded like one thing. They are not one thing.

This is a map of the stack. It is not a leaderboard, a vendor comparison, or a taxonomy for winning arguments online. The goal is simpler: when a new tool appears, you should be able to say, calmly, I know where that fits.

One quick translation: people often say “AI” for the whole experience. In this article, model means the trained engine underneath: GPT-family models, Claude-family models, Gemini, or open-weight models. ChatGPT is the product you touch; a model is one of the engines it can run on.

The durable shape. Products change names. Protocols evolve. Frameworks get popular and then less popular. But the layers underneath are steady: a user works through a harness; the harness assembles context, tools, data, and state; a model reasons over what it receives; an inference system runs the model; observability and governance keep the whole thing honest.

Examples note. Names like ChatGPT, Codex, Claude Code, Cursor, MCP, LangGraph, Ollama, vLLM, OpenClaw, and Hermes are examples of the current ecosystem, not an exhaustive list. Treat them as landmarks on the map.

Part 1

The map

Most confusion comes from flattening the ecosystem into one word: AI. Start by separating the layers, and the arguments get quieter.

The whole picture

AI is a stack, not a product

The most useful first move is to stop asking, “What AI tool is this?” and start asking, “Which layer is this?” A model, a chat product, an agent runtime, an MCP server, an embedding index, and an inference engine all participate in AI systems, but they do very different jobs.

01 USER SURFACE Chat, coding, research, workplace, support, internal copilots 02 HARNESS / APPLICATION Prompts, state, UX, permissions, tool loading, context assembly TOOLS MCP, OpenAPI, functions DATA Docs, code, search, memory ORCHESTRATION Loops, graphs, plans, jobs 03 INFERENCE / SERVING Provider infrastructure, serving engines, local runtimes, GPU scheduling, KV cache 04 MODEL / AI ENGINE GPT-family, Claude-family, Gemini, or open-weight models producing tokens GOVERNANCE EVALS TRACING
Harness
The thing you touch. It owns the product experience and decides what the model sees.
Context
The working set: instructions, messages, tool schemas, retrieved passages, files, and memory.
Serving
The runtime that makes the model fast enough, cheap enough, and available enough to use.
Model
The trained AI engine that reasons over the assembled request and produces output.

The thing you actually use

The harness is what makes the model feel useful

A raw model API is powerful, but it is not yet a product. It accepts a request and returns tokens. The harness is the software around it: the chat window, coding environment, research assistant, customer-support console, workflow runner, or enterprise assistant.

Coding harness
Reads files, searches repositories, edits code, runs tests, manages plans, and asks for approval before risky actions.
Codex, Claude Code, Cursor
Work harness
Connects to documents, mail, calendars, tickets, CRM records, dashboards, and company permissions.
Claude Cowork, ChatGPT-style work assistants, internal copilots
Chat harness
Optimizes for conversation, file upload, search, memory, multimodal input, and a low-friction user experience.
ChatGPT, Claude, Gemini
Custom harness
Your application: domain UI, auth, business rules, retrieval, tools, logging, costs, and deployment.
SDK plus app code

The model does not decide the product. The harness does. It decides what documents are read, which tools are exposed, whether a tool call needs approval, how much history is carried forward, when to compact, where memory lives, what gets measured, and what happens when the model gets something wrong.

Same model, different harness

Imagine the same underlying model, running on the same serving setup, behind two products.

Plain chat

The harness sends your message, prior conversation, maybe an uploaded file, and product instructions. The model answers. The product feels conversational because the harness is built for conversation.

Coding agent

The harness sends repo instructions, selected files, tool schemas, current task state, and permission rules. It can search, edit, run tests, observe failures, and loop. The product feels agentic because the harness is built for work.

Part 2

Inside the harness

Once you see the harness as software, the ecosystem starts looking familiar again: SDKs, state machines, queues, permissions, retries, logs, and product tradeoffs.

Building one yourself

An AI app is still an app

If you build your own harness, you do not start with “agent magic.” You start with normal application decisions. What is the user trying to do? What data can the app access? What actions are allowed? What should be deterministic code, and what should be delegated to the model?

SDK

Provider SDKs and APIs handle the request shape: messages, model choice, tool definitions, streaming, structured output, files, and errors. They make the model callable from your software.

State

Your app decides what persists: chat history, task state, retrieved passages, tool results, user preferences, approvals, and durable artifacts like files or tickets.

Policy

Your app decides what the model may do. Reading a file, sending an email, changing a database row, or spending money should not all have the same permission shape.

Product

Your app decides how work feels: chat, form, wizard, editor, dashboard, queue, background job, or mixed mode. The UI is part of the intelligence because it constrains the work.

A useful rule of thumb: use the model where interpretation, language, ambiguity, or synthesis matter. Use ordinary code where correctness, permissions, transactions, and repeatability matter.

Workflow

Natural language can describe a flow. It does not replace control.

A tempting thought is: if the model understands instructions, why use a workflow library at all? Sometimes that is right. For a small task, a paragraph of instructions is enough. But production systems usually need more than an instruction. They need durable state, retries, branches, approvals, cancellation, and observability.

Use natural language when the stakes are low

Exploration, drafting, local coding help, one-off analysis, and tasks where a human watches the result closely.

Use workflow code when the path matters

Approvals, compliance, scheduled jobs, customer-facing automation, transactional changes, and anything that must resume after failure.

Use graph runtimes when loops matter

Agent workflows are often cyclic: plan, act, observe, revise. Graph runtimes exist because that loop needs control, not just enthusiasm.

Do not make the model your only state machine. If the only record of the workflow is “the model remembers what we said,” the system is fragile. The harness should know the current step, the allowed transitions, the prior tool results, and the stop condition.

Orchestration

The control plane is where the ecosystem is getting crowded

This is the layer behind many new projects: agent runtimes, local assistants, workflow orchestrators, coding agents, and personal automation systems. OpenClaw and Hermes-style projects live around here. They are not new kinds of models. They are harnesses and orchestrators with opinions about memory, tools, schedules, skills, and where the work should run.

Intent
A person asks for work, or a scheduled job wakes up.
Plan
The orchestrator decides whether this is one model call, a tool call, a loop, a graph, or a handoff.
State
The harness records progress outside the model: task state, memory, files, approvals, trace events, and results.
Act
The model asks for tools, writes drafts, edits files, queries data, or delegates to a smaller loop.
Stop
The system decides the task is done, blocked, waiting for approval, or ready for a human to review.

The important distinction is not whether the orchestrator feels autonomous. It is whether the control plane is explicit enough that you can debug it when the model confidently takes the scenic route.

Part 3

Tools, agents, and subagents

Tools are how the system reaches outside the current message: reading from services, calling APIs, writing data, or changing state. Agents are loops around model calls. Subagents are isolated loops with their own context and job.

Tool use

The model does not touch the world directly

When a model “uses a tool,” it is usually not executing the action itself. The harness gives the model a menu of tool descriptions and schemas. The model emits a structured request. The harness validates it, asks for approval if needed, runs the tool, and sends the result back as more context.

1

Expose capability

The harness tells the model that a tool exists: name, description, input schema, and sometimes examples or constraints.

2

Request action

The model chooses a tool and proposes arguments. This is still text-shaped prediction, just constrained into a structured form.

3

Execute outside the model

The harness calls the database, browser, shell, calendar, CRM, search index, or HTTP API. Permissions live here.

4

Return observation

The tool result goes back into the conversation. The model now reasons over the observation and decides what to do next.

MCP and OpenAPI fit here. MCP is a protocol for exposing capabilities to AI applications. OpenAPI is a mature way to describe HTTP APIs. Both can become tool surfaces, but neither removes the need for product judgment about permissions, trust, cost, and what the model should see.

Agents

An agent is a loop, not a spell

The word agent gets stretched until it covers everything. For engineering purposes, keep the definition small: an agent is a system that can take multiple steps toward a goal by alternating between model calls, actions, observations, and state updates.

Useful distinctions once the word “agent” gets noisy.
Pattern What happens Best fit
Chat User asks, model answers, maybe with a tool call. Conversation, explanation, drafting, simple lookup.
Workflow The app owns a mostly fixed path. The model fills in judgment at selected steps. Approvals, forms, ticket triage, repeatable operations.
Agent The system loops: plan, act, observe, revise, stop. Ambiguous tasks where the next step depends on what the last step found.
Subagent A smaller isolated agent handles a piece of work and reports back. Research, code search, document review, parallel tasks, context isolation.

Autonomy is not the goal by itself. The useful question is not “Can the agent keep going?” It is “Should it keep going without a new constraint, approval, or piece of evidence?” A good agent has stopping rules, not just momentum.

Part 4

Data, models, and inference

Once tools and orchestration are visible, the lower layers become easier to place: where knowledge lives, what model runs, and what infrastructure serves it.

Knowledge placement

How knowledge reaches the request

Your company has documents, tickets, code, wikis, emails, CRM records, logs, PDFs, dashboards, and database rows. The model does not know any of that just because it exists. The harness has to choose how knowledge reaches the request.

Common routes for getting knowledge into a request, and the mistake each route prevents.
Route Use when Failure mode
Request context Small, current, task-critical. Put the exact message, file excerpt, image, tool result, or instruction in the request right now. Context bloat. Carrying too much turns every later request into a tax.
Structured query / API The answer lives in fields, rows, metrics, or records. Query databases, CRM, tickets, dashboards, or APIs with filters, joins, and permissions. RAG where SQL was right. Let the source system calculate exact answers; do not ask retrieval to approximate a metric.
Targeted corpus search You can locate the source by title, keyword, metadata, owner, date, or code symbol. Search for candidates, then inspect the useful parts. Whole-corpus dumping. Search is a locator; the model still needs the specific source text or extracted result.
Retrieval / RAG You need recall across lots of unstructured text. Use semantic, keyword, or hybrid retrieval to bring relevant passages into the request. Retrieval treated as truth. Chunking, freshness, ranking, and citations still need design.
Context memory Future requests need stable background without rediscovering it. Store compact user, project, or workflow context the harness can reintroduce when relevant. Stale summaries. Memory needs to be inspectable, editable, and easy to override.

The practical sequence. If the source is structured, query it. If an unstructured source is findable by name, keyword, metadata, or code symbol, use targeted search, then read. If the source is large, language varies, or recall needs to happen repeatedly, add retrieval or RAG. Add context memory only for stable background that should reappear across future work.

Models

Frontier and open-weight models trade control for convenience

A model is the trained system that processes the request and generates the next tokens, pixels, or structured outputs. In a managed API, you usually choose a model and let the provider run it. In a self-hosted setup, you choose the weights and the serving engine that will load them. Neither choice is inherently better; the right one depends on quality, control, privacy, latency, cost, and how much infrastructure you want to own.

Frontier API model

  • Usually best quality and fastest access to new capabilities.
  • Provider owns serving, scaling, safety layers, model updates, and many product integrations.
  • You trade away some control over internals, deployment location, and exact runtime behavior.

Open-weight model

  • You can run it locally, privately, or inside your own infrastructure.
  • You control serving choices, quantization, fine-tunes, routing, and data boundaries.
  • You also own latency, hardware, upgrades, evals, security, and operational surprises.

Serving

The inference engine runs the model

A model is the trained artifact: the weights, often billions of learned numbers. The inference engine is the software that loads those weights, manages memory and batching, and turns requests into output. This is where abstract ideas from the earlier articles become physical: GPUs, memory, KV cache, context length, throughput, latency, and failure modes.

Local runtime

A friendly local runtime for downloading and running models on your machine. Great for experimentation, privacy-sensitive local work, demos, and learning what serving feels like without building a cluster.

Serving engine

A high-throughput serving engine for open models. It focuses on efficient GPU utilization, batching, OpenAI-compatible serving, and the less glamorous work that makes many requests run economically.

Provider infra

Managed APIs hide most serving mechanics. That is a feature. You still feel the consequences through price, latency, rate limits, context windows, caching behavior, and model availability.

Why builders should care. Even if you never run your own model, serving mechanics explain why long context costs more, why cached prefixes matter, why output is expensive, why small models are sometimes enough, and why latency is shaped by more than network time. For the inside view, see Where Have All the Tokens Gone? and Prompt Caching, From the Inside.

Part 5

How to choose

The point of the map is not vocabulary. It is better decisions: what to buy, what to build, what to expose, what to measure, and what to keep boring.

Decision questions

Three questions that actually get used

Long checklists look responsible and then vanish. These three questions are easier to keep in your head, and each points to a different kind of engineering work.

Harness
What should the harness own? Context assembly, permissions, tool loading, state, retries, stop rules, cost budgets, and what the user sees. If the answer matters for correctness or safety, it belongs in the harness, not only in a prompt.
Knowledge
How should knowledge reach the request? Query structured sources, search targeted corpora, retrieve passages, use request context, or reintroduce memory. Choose the narrowest route that preserves truth.
Loop
Does this need durable state or repeated action? If yes, you are designing workflow or agent orchestration. Define progress, approvals, evaluation, observability, cost limits, latency expectations, and stop conditions.

The missing layers

Evals, observability, and governance belong in the design

They are easy to leave out of a first diagram because they are not as exciting as models and agents. In production, they become the difference between a demo and a system.

Evals

Small test sets, scenario checks, grading rubrics, regression tests, and human review loops. Without evals, every model or harness change is a bet you cannot measure.

Observability

Traces, tool-call logs, token counts, cache hit rates, latency, failures, refusals, and user corrections. You cannot improve what you cannot see.

Governance

Auth, permissions, data boundaries, audit logs, tool confirmations, retention rules, and policy. The model is not where all trust decisions belong.

If you only take one thing. AI is not one thing. It is a normal software system wrapped around a strange new compute primitive. The harness is where much of the product behavior lives: what the model sees, what it can do, what state survives, what gets measured, and when the system should stop.