A harness is the engineering layer that turns a raw foundation model into a reliable, production-grade system. This guide walks through structured output, tool calling, agent loops, guardrails, and evaluation — using DeepSeek's models as the running example.
A foundation model is, at its core, a function: text in → text out. A harness is everything you wrap around that function to make it useful and safe in the real world.
The word is used in two related ways:
| Kind | Question it answers | Examples |
|---|---|---|
| Runtime harness | "How do I make the model do something useful?" | Tool calling, agent loops, structured output, retrieval, guardrails, observability |
| Evaluation harness | "How do I measure whether it's doing a good job?" | lm-evaluation-harness, Inspect AI, custom eval suites |
Picture the model as the engine. The harness is the chassis, steering, brakes, gauges, and safety systems around it. The engine matters, but the car doesn't work without the harness.
Here's what you get for free with a raw model, and what you have to build yourself.
| Capability | Raw model | With a harness |
|---|---|---|
| Act on the world | ✗ can only talk | ✓ tool calling: query APIs, run code, update state |
| Typed output | ✗ free text | ✓ JSON schema / function calling → guaranteed shape |
| Context & memory | ✗ only what you paste | ✓ retrieval (RAG), session history, vector store |
| Safety | ✗ no guardrails | ✓ input/output validation, PII redaction, refusal handling |
| Measurability | ✗ unknown quality | ✓ evaluation harness with scores and regressions |
Without a harness you hit the classic failure modes: the model hallucinates a JSON field, "answers" a question it should have routed to a tool, leaks a prompt injection, or quietly degrades across releases with no way to notice.
DeepSeek is a good case study because it gives you both things you want for harness engineering: open-weight models you can self-host, and a cheap, OpenAI-compatible API.
The API is a drop-in for the OpenAI SDK — you only change the base_url:
from openai import OpenAI
client = OpenAI(
api_key="sk-...", # your DeepSeek API key
base_url="https://api.deepseek.com", # OpenAI-compatible
)
resp = client.chat.completions.create(
model="deepseek-chat",
messages=[{"role": "user", "content": "Explain DNS in one sentence."}],
)
print(resp.choices[0].message.content)
DeepSeek exposes two main modes through the API:
| Model id | Family | Best for |
|---|---|---|
deepseek-chat | V3 (general) | Fast chat, tool calling, structured output, agent loops |
deepseek-reasoner | R1 (reasoning) | Math, logic, planning — emits chain-of-thought before answering |
The reasoner model returns its thinking separately from its answer:
resp = client.chat.completions.create(
model="deepseek-reasoner",
messages=[{"role": "user", "content": "9.11 or 9.9, which is bigger?"}],
)
msg = resp.choices[0].message
print("reasoning:", msg.reasoning_content) # chain-of-thought
print("answer: ", msg.content) # final answer
Because the weights are open, you can also self-host: vLLM or SGLang for production serving, Ollama or llama.cpp for local development. A self-hosted vLLM server exposes the same OpenAI-compatible endpoint, so your harness code doesn't change — only the base_url.
A solid runtime harness has four pillars. Each is small on its own; together they're the difference between a demo and a system.
Don't parse prose. Ask the model to emit a typed function call and you get a guaranteed shape every time.
resp = client.chat.completions.create(
model="deepseek-chat",
messages=[{"role": "user",
"content": "Book: Clean Code, Robert C. Martin, 2008"}],
tools=[{
"type": "function",
"function": {
"name": "save_book",
"description": "Save a book record",
"parameters": {
"type": "object",
"properties": {
"title": {"type": "string"},
"author": {"type": "string"},
"year": {"type": "integer"},
},
"required": ["title", "author", "year"],
},
},
}],
)
tool_call = resp.choices[0].message.tool_calls[0]
print(tool_call.function.arguments) # valid JSON: {"title": "Clean Code", ...}
Tools let the model act instead of just talk. You declare a schema; the model returns a call; you run the code.
def get_weather(city):
# your real implementation here
return {"city": city, "temp_c": 22, "condition": "sunny"}
TOOLS = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}]
The loop is the harness's heartbeat: prompt → model → tool → model → … → answer.
import json
def run_agent(prompt, tools, call_tool, max_turns=8):
messages = [{"role": "user", "content": prompt}]
for _ in range(max_turns):
resp = client.chat.completions.create(
model="deepseek-chat", messages=messages, tools=tools,
)
msg = resp.choices[0].message
if not msg.tool_calls:
return msg.content # final answer
messages.append(msg)
for call in msg.tool_calls:
args = json.loads(call.function.arguments)
result = call_tool(call.function.name, args)
messages.append({
"role": "tool",
"tool_call_id": call.id,
"content": json.dumps(result),
})
raise RuntimeError("agent hit the turn limit")
Walk the loop step by step:
Validate what leaves the system. A guardrail is just a deterministic check on the model's output.
ALLOWED = {"accepted", "rejected"}
def guard_label(raw: str):
label = raw.strip().lower()
if label not in ALLOWED:
return None, "unsafe output: " + raw[:40]
return label, None
You cannot improve what you cannot measure. An evaluation harness runs your model (or your whole agent) against a fixed set of tasks and scores it, so every prompt change, model swap, or tool addition is a measurable decision.
A minimal eval loop is surprisingly little code:
def run_evals(cases, solve, score):
results = [score(solve(q), expected) for q, expected in cases]
passed = sum(1 for r in results if r)
return {
"pass": passed,
"total": len(cases),
"mean": sum(results) / len(results),
}
# cases = [("9.11 or 9.9?", "9.9"), ("capital of Nepal?", "Kathmandu"), ...]
# run_evals(cases, solve=ask_model, score=exact_match)
For real work, use a battle-tested framework instead of reinventing the wheel. They all accept DeepSeek because it's OpenAI-compatible:
| Tool | Focus |
|---|---|
lm-evaluation-harness (EleutherAI) | Standard benchmarks (MMLU, GSM8K, HellaSwag, …) |
Inspect AI (UK AI Safety Institute) | Safety + capability evaluations with scorers |
OpenAI Evals | Custom function-based evals |
| LangSmith / Braintrust / Phoenix | Evals + tracing + observability in one |
Here's where DeepSeek's design pays off, and how to capture it.
| API | Self-hosted (vLLM) | |
|---|---|---|
| Setup | Minutes | GPU + serving infra |
| Cost at scale | Per-token | Fixed GPU cost |
| Privacy | Data leaves your network | Data stays yours |
| Latency control | Limited | Full (batching, quantization, KV cache) |
Route reasoning work (math, logic, planning, debugging) to deepseek-reasoner, and everything else (chat, tool calling, structured extraction) to deepseek-chat. The reasoner thinks longer and costs more tokens — spend them where they earn their keep.
A well-known pattern: use a large reasoning model (like R1) to generate high-quality traces, then fine-tune a smaller model on those traces. You keep most of the quality at a fraction of the cost — and because DeepSeek's models are open-weight, this is fully legal and practical.
Instrument the harness, not just the model: tokens in/out, latency per hop, tool-call success rate, cache hit rate, and cost per request. When quality regresses, the trace tells you whether it was the prompt, the retrieval, or the tool.
Putting it together: a small but real harness that has tools, a loop, structured output, and a guardrail.
import json
from openai import OpenAI
client = OpenAI(api_key="sk-...", base_url="https://api.deepseek.com")
TOOLS = [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}
]
def call_tool(name, args):
if name == "get_weather":
return {"city": args["city"], "temp_c": 22, "condition": "sunny"}
raise ValueError("unknown tool: " + name)
def run_agent(prompt, max_turns=8):
messages = [{"role": "user", "content": prompt}]
for _ in range(max_turns):
resp = client.chat.completions.create(
model="deepseek-chat", messages=messages, tools=TOOLS,
)
msg = resp.choices[0].message
if not msg.tool_calls:
return msg.content
messages.append(msg)
for call in msg.tool_calls:
result = call_tool(call.function.name,
json.loads(call.function.arguments))
messages.append({
"role": "tool",
"tool_call_id": call.id,
"content": json.dumps(result),
})
raise RuntimeError("agent hit the turn limit")
if __name__ == "__main__":
answer = run_agent("What is the weather in Kathmandu?")
print(answer)
That's the whole pattern: a loop that alternates between the model and real tools, with the harness — not the model — in control of every side effect.
max_turns (or token budget) prevents infinite tool-call loops.deepseek-reasoner emits hidden reasoning tokens you pay for — use it deliberately.