← Back to Home

🤖 AI Harness Engineering with DeepSeek

A harness is the engineering layer that turns a raw foundation model into a reliable, production-grade system. This guide walks through structured output, tool calling, agent loops, guardrails, and evaluation — using DeepSeek's models as the running example.

1. What "AI Harness" Actually Means

A foundation model is, at its core, a function: text in → text out. A harness is everything you wrap around that function to make it useful and safe in the real world.

The word is used in two related ways:

KindQuestion it answersExamples
Runtime harness"How do I make the model do something useful?"Tool calling, agent loops, structured output, retrieval, guardrails, observability
Evaluation harness"How do I measure whether it's doing a good job?"lm-evaluation-harness, Inspect AI, custom eval suites

Picture the model as the engine. The harness is the chassis, steering, brakes, gauges, and safety systems around it. The engine matters, but the car doesn't work without the harness.

Your app a user's request
↓
Harness: context & tools system prompt, history, retrieved docs, tool schemas
↓
Model (DeepSeek) returns text or a tool call
↓
Harness: validate & act execute tools, guard the output, loop back if needed
↓
Answer typed, validated, logged
💡 The core idea: the model is only one component. Most of the reliability, safety, and utility of an AI system comes from the harness you build around it — not from the model weights.

2. Why a Raw Model Isn't Enough

Here's what you get for free with a raw model, and what you have to build yourself.

CapabilityRaw modelWith a harness
Act on the world✗ can only talk✓ tool calling: query APIs, run code, update state
Typed output✗ free text✓ JSON schema / function calling → guaranteed shape
Context & memory✗ only what you paste✓ retrieval (RAG), session history, vector store
Safety✗ no guardrails✓ input/output validation, PII redaction, refusal handling
Measurability✗ unknown quality✓ evaluation harness with scores and regressions

Without a harness you hit the classic failure modes: the model hallucinates a JSON field, "answers" a question it should have routed to a tool, leaks a prompt injection, or quietly degrades across releases with no way to notice.

⚠️ Mindset shift: don't think "which prompt gives the right answer once?" Think "which system gives the right answer reliably, measurably, and safely, over thousands of calls?"

3. Meet DeepSeek

DeepSeek is a good case study because it gives you both things you want for harness engineering: open-weight models you can self-host, and a cheap, OpenAI-compatible API.

The API is a drop-in for the OpenAI SDK — you only change the base_url:

from openai import OpenAI

client = OpenAI(
    api_key="sk-...",                     # your DeepSeek API key
    base_url="https://api.deepseek.com",  # OpenAI-compatible
)

resp = client.chat.completions.create(
    model="deepseek-chat",
    messages=[{"role": "user", "content": "Explain DNS in one sentence."}],
)

print(resp.choices[0].message.content)

DeepSeek exposes two main modes through the API:

Model idFamilyBest for
deepseek-chatV3 (general)Fast chat, tool calling, structured output, agent loops
deepseek-reasonerR1 (reasoning)Math, logic, planning — emits chain-of-thought before answering

The reasoner model returns its thinking separately from its answer:

resp = client.chat.completions.create(
    model="deepseek-reasoner",
    messages=[{"role": "user", "content": "9.11 or 9.9, which is bigger?"}],
)

msg = resp.choices[0].message
print("reasoning:", msg.reasoning_content)  # chain-of-thought
print("answer:   ", msg.content)            # final answer

Because the weights are open, you can also self-host: vLLM or SGLang for production serving, Ollama or llama.cpp for local development. A self-hosted vLLM server exposes the same OpenAI-compatible endpoint, so your harness code doesn't change — only the base_url.

💡 Compatibility is the superpower: because DeepSeek speaks the OpenAI protocol, every tool, SDK, and eval framework that works with OpenAI works with DeepSeek by changing one URL.

4. Building a Runtime Harness

A solid runtime harness has four pillars. Each is small on its own; together they're the difference between a demo and a system.

a) Structured output

Don't parse prose. Ask the model to emit a typed function call and you get a guaranteed shape every time.

resp = client.chat.completions.create(
    model="deepseek-chat",
    messages=[{"role": "user",
               "content": "Book: Clean Code, Robert C. Martin, 2008"}],
    tools=[{
        "type": "function",
        "function": {
            "name": "save_book",
            "description": "Save a book record",
            "parameters": {
                "type": "object",
                "properties": {
                    "title": {"type": "string"},
                    "author": {"type": "string"},
                    "year": {"type": "integer"},
                },
                "required": ["title", "author", "year"],
            },
        },
    }],
)

tool_call = resp.choices[0].message.tool_calls[0]
print(tool_call.function.arguments)  # valid JSON: {"title": "Clean Code", ...}

b) Tool calling

Tools let the model act instead of just talk. You declare a schema; the model returns a call; you run the code.

def get_weather(city):
    # your real implementation here
    return {"city": city, "temp_c": 22, "condition": "sunny"}

TOOLS = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Get current weather for a city",
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
        },
    },
}]
⚠️ You run the tool: the model never executes code itself. It only requests a call. Your harness performs it and feeds the result back — this is what keeps the system safe and auditable.

c) The agent loop

The loop is the harness's heartbeat: prompt → model → tool → model → … → answer.

import json

def run_agent(prompt, tools, call_tool, max_turns=8):
    messages = [{"role": "user", "content": prompt}]
    for _ in range(max_turns):
        resp = client.chat.completions.create(
            model="deepseek-chat", messages=messages, tools=tools,
        )
        msg = resp.choices[0].message
        if not msg.tool_calls:
            return msg.content            # final answer
        messages.append(msg)
        for call in msg.tool_calls:
            args = json.loads(call.function.arguments)
            result = call_tool(call.function.name, args)
            messages.append({
                "role": "tool",
                "tool_call_id": call.id,
                "content": json.dumps(result),
            })
    raise RuntimeError("agent hit the turn limit")

Walk the loop step by step:

Interactive: follow one turn of the agent loop

  1. The harness assembles messages: system prompt + history + the user's request.
  2. The model returns either content (a final answer) or a tool_calls request.
  3. If it's a tool call, the harness runs call_tool(...) — the model itself never executes code.
  4. The harness appends the result as a role: "tool" message and loops back to the model.
  5. The model sees the tool result and continues until it produces a final content answer.

d) Guardrails

Validate what leaves the system. A guardrail is just a deterministic check on the model's output.

ALLOWED = {"accepted", "rejected"}

def guard_label(raw: str):
    label = raw.strip().lower()
    if label not in ALLOWED:
        return None, "unsafe output: " + raw[:40]
    return label, None

5. The Evaluation Harness

You cannot improve what you cannot measure. An evaluation harness runs your model (or your whole agent) against a fixed set of tasks and scores it, so every prompt change, model swap, or tool addition is a measurable decision.

A minimal eval loop is surprisingly little code:

def run_evals(cases, solve, score):
    results = [score(solve(q), expected) for q, expected in cases]
    passed = sum(1 for r in results if r)
    return {
        "pass": passed,
        "total": len(cases),
        "mean": sum(results) / len(results),
    }

# cases = [("9.11 or 9.9?", "9.9"), ("capital of Nepal?", "Kathmandu"), ...]
# run_evals(cases, solve=ask_model, score=exact_match)

For real work, use a battle-tested framework instead of reinventing the wheel. They all accept DeepSeek because it's OpenAI-compatible:

ToolFocus
lm-evaluation-harness (EleutherAI)Standard benchmarks (MMLU, GSM8K, HellaSwag, …)
Inspect AI (UK AI Safety Institute)Safety + capability evaluations with scorers
OpenAI EvalsCustom function-based evals
LangSmith / Braintrust / PhoenixEvals + tracing + observability in one
💡 Evaluate on your tasks: public benchmarks tell you about the model. Your own eval set — real user queries with expected answers — tells you about your system. Build the second one.

6. Getting the Full Benefits

Here's where DeepSeek's design pays off, and how to capture it.

Self-host vs API

APISelf-hosted (vLLM)
SetupMinutesGPU + serving infra
Cost at scalePer-tokenFixed GPU cost
PrivacyData leaves your networkData stays yours
Latency controlLimitedFull (batching, quantization, KV cache)

Use the right model for the job

Route reasoning work (math, logic, planning, debugging) to deepseek-reasoner, and everything else (chat, tool calling, structured extraction) to deepseek-chat. The reasoner thinks longer and costs more tokens — spend them where they earn their keep.

Distill, then serve cheap

A well-known pattern: use a large reasoning model (like R1) to generate high-quality traces, then fine-tune a smaller model on those traces. You keep most of the quality at a fraction of the cost — and because DeepSeek's models are open-weight, this is fully legal and practical.

Observe everything

Instrument the harness, not just the model: tokens in/out, latency per hop, tool-call success rate, cache hit rate, and cost per request. When quality regresses, the trace tells you whether it was the prompt, the retrieval, or the tool.

💡 Full benefits checklist: (1) structured output, (2) tools + a robust agent loop, (3) guardrails, (4) your own eval set, (5) tracing/cost metrics, (6) the right chat-vs-reasoner split, and (7) caching for repeated prompts.

7. A Complete Minimal Agent

Putting it together: a small but real harness that has tools, a loop, structured output, and a guardrail.

import json
from openai import OpenAI

client = OpenAI(api_key="sk-...", base_url="https://api.deepseek.com")

TOOLS = [
    {
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Get current weather for a city",
            "parameters": {
                "type": "object",
                "properties": {"city": {"type": "string"}},
                "required": ["city"],
            },
        },
    }
]

def call_tool(name, args):
    if name == "get_weather":
        return {"city": args["city"], "temp_c": 22, "condition": "sunny"}
    raise ValueError("unknown tool: " + name)

def run_agent(prompt, max_turns=8):
    messages = [{"role": "user", "content": prompt}]
    for _ in range(max_turns):
        resp = client.chat.completions.create(
            model="deepseek-chat", messages=messages, tools=TOOLS,
        )
        msg = resp.choices[0].message
        if not msg.tool_calls:
            return msg.content
        messages.append(msg)
        for call in msg.tool_calls:
            result = call_tool(call.function.name,
                               json.loads(call.function.arguments))
            messages.append({
                "role": "tool",
                "tool_call_id": call.id,
                "content": json.dumps(result),
            })
    raise RuntimeError("agent hit the turn limit")

if __name__ == "__main__":
    answer = run_agent("What is the weather in Kathmandu?")
    print(answer)

That's the whole pattern: a loop that alternates between the model and real tools, with the harness — not the model — in control of every side effect.

8. Gotchas & Best Practices

💡 Summary: the model gives you intelligence; the harness gives you reliability. Build the loop, constrain the output, guard the edges, and measure everything — and you'll get far more value out of DeepSeek (or any model) than a raw completion ever could.
Copied to clipboard!