中文

DSH's Model Interface Layer: Why Every LLM Needs a 'Translator'

AI, Agent, DeepSeek, DeepSeek Harness, LLM, Architecture

DSH’s Model Interface Layer: Why Every LLM Needs a “Translator”

If you’ve used Claude, DeepSeek, OpenAI, or a locally running LM Studio, you’ve probably noticed that they all claim to be “chat,” but they all speak differently.

Not in the sense that the answers differ, but their API formats, streaming outputs, error codes, and even whether the “thinking process” should be exposed all vary. It’s like running a company whose customers are Chinese, Japanese, and French: each submits contracts, invoices, and complaints in a different format.

DeepSeek Harness’s (DSH) model interface layer is essentially a translator management system for this company. It defines an internal lingua franca so upper-layer business logic only cares about “I want to make a call,” without worrying whether the other end is DeepSeek, OpenAI, or the LM Studio running on your balcony.


In one sentence: what problem does the interface layer solve?

DSH doesn’t want to know which LLM is on the other side, so it only speaks “Mandarin”; each model gets a well-behaved “translator”; swapping models means swapping translators, and the main codebase doesn’t change by a single line.

The value of this design is direct: if you switch from DeepSeek to LM Studio in settings today, chat, compression, and title generation all require zero rewrites. It’s not magic; the interface layer has already swallowed the differences.


How messy would it be without an interface layer?

Imagine every DSH feature—the agent loop, context compression, title generation, and tool calling—had to be written once for DeepSeek, again for OpenAI, and again for LM Studio.

Any model vendor tweaks their API format and you have to grep the entire project to fix it. A new model launches? Write another version. One error code is misread? The whole downstream pipeline collapses.

DSH’s solution is simple: let the upper layer speak only one language, and let translators handle the dialects.

DSH core (agent loop, compression, UI)
        │  speaks only Mandarin
        ▼
   ctx.llm (interface layer switchboard)
        │  picks a translator by provider name
        ▼
   ┌────────────┬──────────────┬───────────┐
   DeepSeek     universal        LM Studio
   dedicated    translator      (OpenAI-compatible)
   translator   (pi-ai)

This structure looks simple, but it turns “model integration” from a scattered mess of dirty work into a pluggable plugin system.


DSH’s “Mandarin” has only three things

DSH doesn’t try to support every field of every model. Instead, it forcibly defines the smallest common denominator. Any model that wants to join must translate its output into these three things:

1. Standard letterhead: Message

A message = id + role (system/user/assistant) + content blocks + attribution. The format is fixed and immutable. This is the raw material of every conversation.

2. Content block: ContentBlock

A letterhead can only hold five kinds of content: text, reasoning, image, tool call, and tool result. There is no sixth kind.

This restriction matters. It means upper-layer code never has to guess “is this response a string or an object?” It only handles one of these five types.

3. Call ticket: GenerateOptions

A complete package for one model call: which model, chat history, system prompt, tool list, temperature… all parameters are sealed in one go and cannot be tampered with mid-call.

The point of sealing isn’t just to prevent mischief. It’s to make calls reproducible during debugging. If a call behaves strangely, you can replay the ticket instead of chasing down whether the temperature was changed somewhere else.


Streaming responses are sliced into seven “fragments”

Large models don’t speak in one go; they emit words one by one. DSH calls every emitted piece a StreamChunk and strictly defines only seven types:

FragmentHuman meaning
block-start”I’m about to say a block of text.”
text-deltaText increment (“to”, “day”, “weather”…)
reasoning-deltaReasoning increment (the model’s internal monologue)
tool-call-deltaTool-call increment (parameters arrive as raw JSON fragments)
block-end”This block is done,” and it carries the fully assembled block
usageThis call’s bill (token usage)
finish”I’m done / something went wrong / you hung up”

Several design choices here are quite elegant:

First, fragments carry indices. A model can interleave speaking, reasoning, and writing three tool calls at once, but the indices keep them from crossing wires.

Second, block-end carries the finished product. Recipients don’t have to assemble fragments themselves; they receive a complete block. This avoids a very common bug: different subsystems assembling the same fragments in slightly different ways and ending up with mismatched data.

Third, tool-call parameters are always raw strings. Models generate JSON character by character; forcibly parsing it into an object mid-stream will blow up as soon as the JSON is incomplete. Keeping it as a raw string is always safe.


The “company headquarters” behind the switchboard

packages/llm/llm is the interface layer proper. Each file inside maps to a clearly defined role:

  • src/index.ts: switchboard + dispatch. ctx.llm manages the translator roster and is the single entry point for making model calls.
  • src/types.ts: the Mandarin dictionary. The home of all type definitions.
  • src/message.ts: the standard letterhead factory. Constructs and validates message formats.
  • src/content.ts: the letter inspection tool. Recursively checks whether a letter contains images, using one shared ruler for the whole company.
  • src/assembler.ts: the sole paper-shredder reassembler. Reassembles fragments into complete messages; only one place does this to prevent local reassembly bugs.
  • src/call-config.ts: call ticket + laminator. Once a parameter set is finalized, it’s deeply frozen; tampering throws an exception.
  • src/error.ts: unified fault form. Every error carries a stable error code; downstream logic handles it by code, not by parsing error text.
  • src/adapter-failure.ts: the fault-form transcription desk. Bizarre errors thrown by translators are transcribed into a standard format before being handed upstairs.
  • src/retry-policy.ts: retry rule template. Each translator submits one on hiring; the switchboard archives it, and replacements still follow the old rule.
  • src/attribution.ts: unified badge. Every model-vendor call wears a User-Agent with only public information; no keys or session ids are included.
  • src/brand.ts: anti-counterfeit stamp. Type-brands various IDs so you can’t accidentally use a “tool-call id” as a “request id.”
  • src/never.ts: quality-control sentry. If someone adds a new fragment type to the dictionary, every place that doesn’t handle it fails to compile—no omissions allowed.

This division of labor is extremely fine-grained, almost obsessive. But it’s exactly this obsessiveness that lets the interface layer stably absorb the quirks of different models over the long run.


What do translators look like?

Dedicated translator: llm-deepseek

DeepSeek’s official API has its own SSE stream, error format, and field naming. DSH’s llm-deepseek only does four things:

  1. adapter.ts: make the call (fetch + direct SSE).
  2. serialize.ts: outbound translation, converting standard letterhead into DeepSeek dialect.
  3. translate.ts: inbound translation, converting DeepSeek SSE events into the seven standard fragments.
  4. sse.ts: telegraph decoder, handling sticky packets, UTF-8 boundaries, [DONE], and other details.

Keys and addresses are not kept by the translator; it fetches them fresh for every call. This lets it hot-swap configuration without a restart.

Universal translator: llm-pi-ai

LM Studio, custom gateways, and any OpenAI-compatible endpoint all go through llm-pi-ai. It isn’t one adapter per model; it solves all compatible endpoints with a single “route → config” dictionary. This means connecting LM Studio requires zero code changes: just fill in the address and model name.

Two colleagues who don’t translate but still matter

  • llm-retry: retry dispatcher. Listens to failure broadcasts and decides whether to redial based on each translator’s archived rule table. Every redial starts a new indexed round with a complete archive.
  • token-meter: electricity meter. Independently counts tokens per session from the session log; departments like compression share its readings.

Eight iron rules every translator must follow on hire

Becoming a DSH translator isn’t a free-for-all. DSH defines eight adapter conventions that form the interface layer’s behavioral contract:

  1. Report the bill first, then say “done,” then shut up. usage must come before finish; nothing follows finish.
  2. Tool parameters are always raw JSON strings. Even if the vendor gives you a fully assembled object, break it back into strings and stream them in pieces.
  3. Errors have only two paths: either throw LlmError, or end with finish { kind: 'error' }.
  4. No secret retries. One call is one attempt; retries are the dispatcher’s job.
  5. Stalling for five minutes is a timeout. Report TIMEOUT; user cancellation is ABORTED; don’t mix them up.
  6. “Context too long” has one and only one code: CONTEXT_WINDOW_EXCEEDED; downstream only recognizes the code.
  7. Returning nothing at all is still a failure. It will be retried by default; don’t pretend it succeeded.
  8. Wear your badge on every call. And automated tests are watching to prove you did.

These rules sound trivial, but they are why the interface layer can provide a stable floor. Every translator performs from the same script, so the upper layer doesn’t need special handling for each model.


The complete journey of one call

Putting it all together, a single model call looks like this:

  1. The agent loop writes a message using standard letterhead.
  2. The call ticket is reconstructed from the session log—sealed, no tampering.
  3. The switchboard finds the right translator by provider name.
  4. The call passes through the llm/stream security conveyor (retry, replay, and routing intercept here).
  5. The translator performs outbound translation and calls the model vendor.
  6. Fragments come back → the assembler builds a complete message → raw fragments are logged simultaneously.
  7. The meter records the bill.
  8. Something breaks? A unified fault form is opened → the dispatcher decides whether to redial based on the rule table.

Because every step is logged, resuming after a disconnect, forking a session, and snapshot testing can all be replayed byte for byte. This is DSH’s “model-visible means already-logged” rule landing at the model layer.


Beginner FAQ

Q: Where does the “cache hit X%” in the UI come from? A: The bill (TokenUsage) has a dedicated “cache hit” slot, but it only gets a number if the translator reports it. DeepSeek’s official API reports it; LM Studio/llama.cpp does not → it stays 0%. It’s “not reported,” not “not cached.”

Q: Why are tool parameters strings instead of objects? A: Because models emit characters one by one. Forcibly parsing into an object mid-stream will blow up as soon as the JSON is incomplete. Store the raw string exactly as it arrives and it’s always safe.

Q: How do I connect my own model? A: Write a translator: inherit LlmAdapter, implement stream(), and register it in a plugin. Follow the eight rules. The official docs/cookbook/adding-an-llm-adapter.zh.md is a step-by-step tutorial.

Q: What happens if I change model settings mid-request? A: Nothing. prepareCall() ensures capability checks, archiving, and the actual call all use the same translator. Hot updates can’t split them; new config takes effect from the next call onward.

Q: Why separate reasoning and answer? A: Because reasoning usually shouldn’t be displayed directly or fed verbatim into compression summaries. Splitting them into two block types lets the UI fold it, compression skip it, and vendors replay it—each as needed.


One-sentence summary

DSH’s model interface layer is not a “multi-model support” product feature; it is model-access infrastructure. With a mandatory common language, a strict translator contract, and a single switchboard entry point, it turns “model integration” from scattered dirty work into a clean, testable, replaceable, and auditable process.

For anyone building their own agent, this design means: you can swap models freely without rewriting business logic. Your agent won’t be held hostage by any single LLM vendor.


Reference: DSH v0.1.0-rc.5 packages/llm/, docs/subsystems/llm-streaming.zh.md

引用格式

APA: Arlen. (2026). DSH's Model Interface Layer: Why Every LLM Needs a 'Translator'. Retrieved from https://strongya.dev/en/posts/dsh-model-interface-layer/
MLA: Arlen. "DSH's Model Interface Layer: Why Every LLM Needs a 'Translator'." 2026. Web. 2026-08-16.
GB/T 7714: Arlen. DSH's Model Interface Layer: Why Every LLM Needs a 'Translator'[EB/OL]. 2026-08-16. https://strongya.dev/en/posts/dsh-model-interface-layer/.
Agent引用要求: 请使用 https://strongya.dev/en/posts/dsh-model-interface-layer/ 作为 canonical_url 并标注来源

Share to

© 2026 Arlen. All Rights Reserved. CC BY-NC-SA 4.0

Agent引用要求: 请使用 canonical_url 并标注来源