Field guide · ecosystem map
How harnesses, agents, tools, data, models, and inference engines fit together
When I first tried to understand modern AI, the hardest part was not any single concept. It was the pile of names: ChatGPT, Codex, Claude Code, Cursor, MCP, LangGraph, Ollama, vLLM, frontier models, open-weight models, RAG, embeddings, memory, agents, subagents. They sounded like one thing. They are not one thing.
This is a map of the stack. It is not a leaderboard, a vendor comparison, or a taxonomy for winning arguments online. The goal is simpler: when a new tool appears, you should be able to say, calmly, I know where that fits.
One quick translation: people often say “AI” for the whole experience. In this article, model means the trained engine underneath: GPT-family models, Claude-family models, Gemini, or open-weight models. ChatGPT is the product you touch; a model is one of the engines it can run on.
The durable shape. Products change names. Protocols evolve. Frameworks get popular and then less popular. But the layers underneath are steady: a user works through a harness; the harness assembles context, tools, data, and state; a model reasons over what it receives; an inference system runs the model; observability and governance keep the whole thing honest.
Examples note. Names like ChatGPT, Codex, Claude Code, Cursor, MCP, LangGraph, Ollama, vLLM, OpenClaw, and Hermes are examples of the current ecosystem, not an exhaustive list. Treat them as landmarks on the map.
Most confusion comes from flattening the ecosystem into one word: AI. Start by separating the layers, and the arguments get quieter.
The whole picture
The most useful first move is to stop asking, “What AI tool is this?” and start asking, “Which layer is this?” A model, a chat product, an agent runtime, an MCP server, an embedding index, and an inference engine all participate in AI systems, but they do very different jobs.
The thing you actually use
A raw model API is powerful, but it is not yet a product. It accepts a request and returns tokens. The harness is the software around it: the chat window, coding environment, research assistant, customer-support console, workflow runner, or enterprise assistant.
The model does not decide the product. The harness does. It decides what documents are read, which tools are exposed, whether a tool call needs approval, how much history is carried forward, when to compact, where memory lives, what gets measured, and what happens when the model gets something wrong.
Imagine the same underlying model, running on the same serving setup, behind two products.
The harness sends your message, prior conversation, maybe an uploaded file, and product instructions. The model answers. The product feels conversational because the harness is built for conversation.
The harness sends repo instructions, selected files, tool schemas, current task state, and permission rules. It can search, edit, run tests, observe failures, and loop. The product feels agentic because the harness is built for work.
Once you see the harness as software, the ecosystem starts looking familiar again: SDKs, state machines, queues, permissions, retries, logs, and product tradeoffs.
Building one yourself
If you build your own harness, you do not start with “agent magic.” You start with normal application decisions. What is the user trying to do? What data can the app access? What actions are allowed? What should be deterministic code, and what should be delegated to the model?
Provider SDKs and APIs handle the request shape: messages, model choice, tool definitions, streaming, structured output, files, and errors. They make the model callable from your software.
Your app decides what persists: chat history, task state, retrieved passages, tool results, user preferences, approvals, and durable artifacts like files or tickets.
Your app decides what the model may do. Reading a file, sending an email, changing a database row, or spending money should not all have the same permission shape.
Your app decides how work feels: chat, form, wizard, editor, dashboard, queue, background job, or mixed mode. The UI is part of the intelligence because it constrains the work.
A useful rule of thumb: use the model where interpretation, language, ambiguity, or synthesis matter. Use ordinary code where correctness, permissions, transactions, and repeatability matter.
Workflow
A tempting thought is: if the model understands instructions, why use a workflow library at all? Sometimes that is right. For a small task, a paragraph of instructions is enough. But production systems usually need more than an instruction. They need durable state, retries, branches, approvals, cancellation, and observability.
Exploration, drafting, local coding help, one-off analysis, and tasks where a human watches the result closely.
Approvals, compliance, scheduled jobs, customer-facing automation, transactional changes, and anything that must resume after failure.
Agent workflows are often cyclic: plan, act, observe, revise. Graph runtimes exist because that loop needs control, not just enthusiasm.
Do not make the model your only state machine. If the only record of the workflow is “the model remembers what we said,” the system is fragile. The harness should know the current step, the allowed transitions, the prior tool results, and the stop condition.
Orchestration
This is the layer behind many new projects: agent runtimes, local assistants, workflow orchestrators, coding agents, and personal automation systems. OpenClaw and Hermes-style projects live around here. They are not new kinds of models. They are harnesses and orchestrators with opinions about memory, tools, schedules, skills, and where the work should run.
The important distinction is not whether the orchestrator feels autonomous. It is whether the control plane is explicit enough that you can debug it when the model confidently takes the scenic route.
Tools are how the system reaches outside the current message: reading from services, calling APIs, writing data, or changing state. Agents are loops around model calls. Subagents are isolated loops with their own context and job.
Tool use
When a model “uses a tool,” it is usually not executing the action itself. The harness gives the model a menu of tool descriptions and schemas. The model emits a structured request. The harness validates it, asks for approval if needed, runs the tool, and sends the result back as more context.
The harness tells the model that a tool exists: name, description, input schema, and sometimes examples or constraints.
The model chooses a tool and proposes arguments. This is still text-shaped prediction, just constrained into a structured form.
The harness calls the database, browser, shell, calendar, CRM, search index, or HTTP API. Permissions live here.
The tool result goes back into the conversation. The model now reasons over the observation and decides what to do next.
MCP and OpenAPI fit here. MCP is a protocol for exposing capabilities to AI applications. OpenAPI is a mature way to describe HTTP APIs. Both can become tool surfaces, but neither removes the need for product judgment about permissions, trust, cost, and what the model should see.
Agents
The word agent gets stretched until it covers everything. For engineering purposes, keep the definition small: an agent is a system that can take multiple steps toward a goal by alternating between model calls, actions, observations, and state updates.
| Pattern | What happens | Best fit |
|---|---|---|
| Chat | User asks, model answers, maybe with a tool call. | Conversation, explanation, drafting, simple lookup. |
| Workflow | The app owns a mostly fixed path. The model fills in judgment at selected steps. | Approvals, forms, ticket triage, repeatable operations. |
| Agent | The system loops: plan, act, observe, revise, stop. | Ambiguous tasks where the next step depends on what the last step found. |
| Subagent | A smaller isolated agent handles a piece of work and reports back. | Research, code search, document review, parallel tasks, context isolation. |
Autonomy is not the goal by itself. The useful question is not “Can the agent keep going?” It is “Should it keep going without a new constraint, approval, or piece of evidence?” A good agent has stopping rules, not just momentum.
Once tools and orchestration are visible, the lower layers become easier to place: where knowledge lives, what model runs, and what infrastructure serves it.
Knowledge placement
Your company has documents, tickets, code, wikis, emails, CRM records, logs, PDFs, dashboards, and database rows. The model does not know any of that just because it exists. The harness has to choose how knowledge reaches the request.
| Route | Use when | Failure mode |
|---|---|---|
| Request context | Small, current, task-critical. Put the exact message, file excerpt, image, tool result, or instruction in the request right now. | Context bloat. Carrying too much turns every later request into a tax. |
| Structured query / API | The answer lives in fields, rows, metrics, or records. Query databases, CRM, tickets, dashboards, or APIs with filters, joins, and permissions. | RAG where SQL was right. Let the source system calculate exact answers; do not ask retrieval to approximate a metric. |
| Targeted corpus search | You can locate the source by title, keyword, metadata, owner, date, or code symbol. Search for candidates, then inspect the useful parts. | Whole-corpus dumping. Search is a locator; the model still needs the specific source text or extracted result. |
| Retrieval / RAG | You need recall across lots of unstructured text. Use semantic, keyword, or hybrid retrieval to bring relevant passages into the request. | Retrieval treated as truth. Chunking, freshness, ranking, and citations still need design. |
| Context memory | Future requests need stable background without rediscovering it. Store compact user, project, or workflow context the harness can reintroduce when relevant. | Stale summaries. Memory needs to be inspectable, editable, and easy to override. |
The practical sequence. If the source is structured, query it. If an unstructured source is findable by name, keyword, metadata, or code symbol, use targeted search, then read. If the source is large, language varies, or recall needs to happen repeatedly, add retrieval or RAG. Add context memory only for stable background that should reappear across future work.
Models
A model is the trained system that processes the request and generates the next tokens, pixels, or structured outputs. In a managed API, you usually choose a model and let the provider run it. In a self-hosted setup, you choose the weights and the serving engine that will load them. Neither choice is inherently better; the right one depends on quality, control, privacy, latency, cost, and how much infrastructure you want to own.
Serving
A model is the trained artifact: the weights, often billions of learned numbers. The inference engine is the software that loads those weights, manages memory and batching, and turns requests into output. This is where abstract ideas from the earlier articles become physical: GPUs, memory, KV cache, context length, throughput, latency, and failure modes.
A friendly local runtime for downloading and running models on your machine. Great for experimentation, privacy-sensitive local work, demos, and learning what serving feels like without building a cluster.
A high-throughput serving engine for open models. It focuses on efficient GPU utilization, batching, OpenAI-compatible serving, and the less glamorous work that makes many requests run economically.
Managed APIs hide most serving mechanics. That is a feature. You still feel the consequences through price, latency, rate limits, context windows, caching behavior, and model availability.
Why builders should care. Even if you never run your own model, serving mechanics explain why long context costs more, why cached prefixes matter, why output is expensive, why small models are sometimes enough, and why latency is shaped by more than network time. For the inside view, see Where Have All the Tokens Gone? and Prompt Caching, From the Inside.
The point of the map is not vocabulary. It is better decisions: what to buy, what to build, what to expose, what to measure, and what to keep boring.
Decision questions
Long checklists look responsible and then vanish. These three questions are easier to keep in your head, and each points to a different kind of engineering work.
The missing layers
They are easy to leave out of a first diagram because they are not as exciting as models and agents. In production, they become the difference between a demo and a system.
Small test sets, scenario checks, grading rubrics, regression tests, and human review loops. Without evals, every model or harness change is a bet you cannot measure.
Traces, tool-call logs, token counts, cache hit rates, latency, failures, refusals, and user corrections. You cannot improve what you cannot see.
Auth, permissions, data boundaries, audit logs, tool confirmations, retention rules, and policy. The model is not where all trust decisions belong.
If you only take one thing. AI is not one thing. It is a normal software system wrapped around a strange new compute primitive. The harness is where much of the product behavior lives: what the model sees, what it can do, what state survives, what gets measured, and when the system should stop.