Skip to results
JSPARK3

JSPARK3 · Latest release

JSPARK3 v2.0.1: GLM-5.3 Flash
on three DGX Sparks.

One stream code
80.9tok/s
One stream prose
63.6tok/s
ablit weights
Cold prefill, 64K prompt (RigMark)
2,124tok/s
Four at once, end to end (RigMark)
113.4tok/s
short code, end-to-end, 256-token cap per agent

base weights + draft model unless marked

Three Sparks. One model.Three champagne-gold DGX Sparks with metal-foam faces, wired as a switchless ring: each Spark cables both of its rear QSFP ports, one to each of the other two. The two workers face forward; Rank 0 is turned round to show its rear panel and the two cables in it. Token pulses illustrate shared fabric activity, not live telemetry or sequential model execution. Rank zero exposes the API; all three machines participate in tensor-parallel execution.00 / API HEAD02 / WORKER01 / WORKERONE ENDPOINTTHREE-WAY EXECUTION / RoCE FABRIC
Three DGX Sparks · one endpointOne Spark answers requests; all three run the model.
Release
v2.0.1 · Oct 3, 2026
Engine
A fork of TensorFold 0.3.6.2 (MIT) · replaces vLLM (JSpark3 v1.8.x)
Weights
Base or refusal-removed (abliterated), chosen at install
Hardware
Three DGX Sparks · one endpoint
Longest context
262,144 tokens

v2.0.1, measured on our three Sparks.

Every figure names the weights it was measured with, and whether the draft model was on.

v2.0.1 against the v1.8.4 baseline

RigMark, the same protocol on v1.8.4 and v2.0.1 · each figure drawn to its own scale from zero

Appliance comparison: different model IDs, not a same-weights claim. v1.8.4 ran with reasoning off, its default; v2.0.1 ran at reasoning effort low. Cold prefill and replay rows use raw token-ID completions, where reasoning effort does not apply.

  • v1.8.4 on vLLM, the baseline
  • v2.0.1, base weights + draft model
  • v2.0.1, refusal-removed (ablit) weights + draft model

Code, decode estimatetok/s, higher is better

v1.8.461.1 tok/s
v2.0.1 Base91.3 tok/s
v2.0.1 Abliterated90.1 tok/s

Prose, decode estimatetok/s, higher is better

v1.8.431.5 tok/s
v2.0.1 Base51.6 tok/s
v2.0.1 Abliterated50.8 tok/s

Prose row: at reasoning effort low, v2.0.1 writes a short reasoning passage before the visible text (with base weights, 11 to 12 tokens, about 1% of each reply of about 1,000 tokens); v1.8.4, with reasoning off, wrote none. RigMark counts those tokens in the prose decode rate and in last output time.

Structured output ceilingtok/s, higher is better

v1.8.495.4 tok/s
v2.0.1 Base127.9 tok/s
v2.0.1 Abliterated127.4 tok/s

Cold prefill, 64K prompttok/s, higher is better

v1.8.41,505 tok/s
v2.0.1 Base2,124 tok/s
v2.0.1 Abliterated2,129 tok/s

Four at once, end to endtok/s, higher is better

v1.8.486.6 tok/s
v2.0.1 Base113.4 tok/s
v2.0.1 Abliterated128.4 tok/s

C1 per-stream time to first tokens, lower is better

v1.8.40.396 s
v2.0.1 Base0.405 s
v2.0.1 Abliterated0.361 s

With base weights, v1.8.4 is about 2% faster on this row. The first token is visible text in every reply in each column.

Prose time to first visible texts, lower is better

v1.8.40.380 s
v2.0.1 Base0.484 s
v2.0.1 Abliterated0.497 s

With base weights, v1.8.4 shows prose text about 0.10 s sooner; v2.0.1 takes about 1.27x as long. RigMark's own prose time to first token marks the first reasoning token, not visible text, so it is not shown.

RigMark output, as RigMark reports it
v1.8.4 on vLLM
MODEL      glm-5.3-flash
APPLIANCE  3x NVIDIA DGX Spark (GB10), 128 GB unified memory each
RUN        reasoning=low  •  protocol=1.1.0
SOURCE     git:c5a0db01b054  •  clean
WORKLOAD       DECODE EST.      LAST OUTPUT          RANGE          BASIC GATE
CODE              61.1 tok/s      37.6s last     58.2–64.7     ✓ 5/5
PROSE             31.5 tok/s      30.7s last     31.1–32.3     ✓ 5/5
STRUCTURED*       95.4 tok/s       5.0s last     94.0–95.8     ✓ 5/5
* predictable-output ceiling; not a proxy for agent speed
64K PREFILL   cold 1,505 tok/s  •  immediate replay 132,921 tok/s
AGGREGATE   C1 42.1  •  C2 61.9  •  C4 86.6 tok/s   (short code, end-to-end, 256-token cap per agent)
C4 OUTPUT STATE   normal stop 0/12  •  visible 12/12  •  reasoning may be included
JSON       sha256:a3a5adf1b11a5a3e…
v2.0.1, base weights + draft model
MODEL      glm53
APPLIANCE  3x NVIDIA DGX Spark (GB10), 128 GB unified memory each
RUN        reasoning=low  •  protocol=1.1.0
SOURCE     git:c5a0db01b054  •  clean
WORKLOAD       DECODE EST.      LAST OUTPUT          RANGE          BASIC GATE
CODE              91.3 tok/s      21.3s last     88.2–92.8     ✓ 5/5
PROSE             51.6 tok/s      19.3s last     50.3–53.1     ✓ 5/5
STRUCTURED*      127.9 tok/s       3.8s last    127.4–128.3    ✓ 5/5
* predictable-output ceiling; not a proxy for agent speed
64K PREFILL   cold 2,124 tok/s  •  immediate replay 821,037 tok/s
AGGREGATE   C1 64.8  •  C2 84.6  •  C4 113.4 tok/s   (short code, end-to-end, 256-token cap per agent)
C4 OUTPUT STATE   normal stop 0/12  •  visible 12/12  •  reasoning may be included
JSON       sha256:05051f88e880384c…
v2.0.1, refusal-removed (ablit) weights + draft model
MODEL      glm53
APPLIANCE  3x NVIDIA DGX Spark (GB10), 128 GB unified memory each
RUN        reasoning=low  •  protocol=1.1.0
SOURCE     git:c5a0db01b054  •  clean
WORKLOAD       DECODE EST.      LAST OUTPUT          RANGE          BASIC GATE
CODE              90.1 tok/s      23.8s last     88.4–92.0     ✓ 5/5
PROSE             50.8 tok/s      19.3s last     50.6–51.6     ✓ 5/5
STRUCTURED*      127.4 tok/s       3.8s last    121.0–127.9    ✓ 5/5
* predictable-output ceiling; not a proxy for agent speed
64K PREFILL   cold 2,129 tok/s  •  immediate replay 812,054 tok/s
AGGREGATE   C1 67.9  •  C2 93.8  •  C4 128.4 tok/s   (short code, end-to-end, 256-token cap per agent)
C4 OUTPUT STATE   normal stop 0/12  •  visible 12/12  •  reasoning may be included
JSON       sha256:1b42ea87215719ef…

Every measured v2.0.1 set

Time to first token on a cold prompt, by prompt length · s, lower is better

  • base weights + draft model
  • refusal-removed (ablit) weights + draft model
  • base weights, no draft model (commercial use)

8k tokens

Base3.8 s
Abliterated3.8 s
Base, no draft3.9 s

32k tokens

Base15.0 s
Abliterated15.0 s
Base, no draft15.8 s

64k tokens

Base30.4 s
Abliterated30.4 s
Base, no draftNot measured

128k tokens

Base63.6 s
Abliterated63.8 s
Base, no draftNot measured

Decode with requests running at once, all streams combined · tok/s, higher is better

  • base weights + draft model
  • refusal-removed (ablit) weights + draft model
  • base weights, no draft model (commercial use)

One requestshort prompts, 41-62 tokens, 1 concurrent

Base59.8 tok/s
Abliterated69.8 tok/s
Base, no draft59.0 tok/s

2 requestsshort prompts, 41-62 tokens, 2 concurrent

Base77.6 tok/s
Abliterated75.8 tok/s
Base, no draftNot measured

4 requestsshort prompts, 41-62 tokens, 4 concurrent

Base97.5 tok/s
Abliterated96.4 tok/s
Base, no draftNot measured

8 requestsshort prompts, 41-62 tokens, 8 concurrent

Base126.0 tok/s
Abliterated121.7 tok/s
Base, no draft81.2 tok/s

16 requestsshort prompts, 41-62 tokens, 16 concurrent

Base106.8 tok/s
Abliterated110.0 tok/s
Base, no draftNot measured

Decode speed, one request at a time · tok/s, higher is better

  • base weights + draft model
  • refusal-removed (ablit) weights + draft model
  • base weights, no draft model (commercial use)

Short code reply

Base80.9 tok/s
Abliterated73.3 tok/s
Base, no draft63.6 tok/s

Short prose reply

Base59.6 tok/s
Abliterated63.6 tok/s
Base, no draft57.7 tok/s

Code after a 32k prompt

Base83.6 tok/s
Abliterated72.3 tok/s
Base, no draftNot measured

First token with 8 requests at once, median · lower is better

short prompts, 41-62 tokens, 8 concurrent

0.54 s

base weights + draft model

First visible text at about 1.5 s (median); 3 of 24 short replies spent their 96-token limit on reasoning and showed no text. Measured at reasoning effort low.

refusal-removed (ablit) weights + draft model: 0.54 s

First visible text at about 1.2 s (median); 1 of 24 short replies spent its 96-token limit on reasoning and showed no text. Measured at reasoning effort low.

base weights, no draft model (commercial use): 1.9 s

Visible text 0.41 s after the first token (median, p95 1.6 s) in the 21 replies that showed text; 3 of 24 short replies spent their 96-token limit on reasoning and showed no text. Measured at reasoning effort low.

The first-token notes use two measures: with the draft model, the median time to first visible text; without it, the median gap from first token to first visible text in the replies that showed text, because a median taken only over replies that showed text would come out below the first-token median taken over all replies.

Pause in 8 running replies as a long prompt arrives · lower is better

short prompts, 41-62 tokens, 8 concurrent, while a prompt of about 8,000 or 36,000 tokens joins

0.52 s

base weights + draft model, worst 0.75 s

refusal-removed (ablit) weights + draft model: 0.49 s, worst 0.57 s

base weights, no draft model (commercial use): Not measured

How every figure was measured

Every set was measured on the same build, each on its weights variant's shipped settings, and no figure is a best run. Rates and times are medians, with the number of runs given below; the token gap is given as both a median and a maximum, and the context window is a setting, not a measurement. Short-reply decode is reported separately for code and for prose: the per-stream rate of replies capped at 256 tokens (median of 3 each). Long decode is one greedy code stream of up to 96 tokens after a 32K-token prompt (median of 3). Aggregate decode is the wall-clock rate of concurrent greedy replies, capped at 96 tokens, to a fixed mix of varied short prompts (41 to 62 tokens each) that is the same for every set (median of 3). Draft acceptance depends on the prompt, so it is also reported for a single repeated prompt. Cold time to first token uses exactly the stated number of prompt tokens with nothing cached (median of 2). The 8-request time to first token is the median wait for the first token when 8 short prompts from the mix (41 to 62 tokens) are sent at once; without the draft model, prompts that arrive together are not read together in one batch, and first tokens arrive later (see known issues). The 8-stream token gap is the longest pause seen by running streams while a prompt of about 8,000 or 36,000 tokens joins, reported as the median and the maximum of that pause across runs. Draft acceptance is the number of draft tokens accepted per verify step, not counting the token the model adds itself, with accepted over proposed tokens alongside, both measured at 8 concurrent requests. The no-draft-model set is a reduced run. The benchmark client runs on a separate machine on the same local network, so client-side times include one network hop. The benchmark figures in the result tables, including the stall bounds, come from chat requests at low reasoning effort. Installation, disk and startup figures are not chat measurements. Low effort still reasons before it answers, and decode and aggregate rates count reasoning tokens. A chat request that sets no reasoning effort runs at high effort, so its replies are longer and its rates can differ from these. Time to first token is measured to the first streamed token, reasoning or text. In the cold-prompt, newcomer and saved-session tests, that first token was visible text in all but two replies: one 64K cold-prompt reply with base weights and the saved-session return with refusal-removed (ablit) weights reached their eight-token limit on reasoning and showed no text. In the eight-client short-prompt test, most replies began with a short reasoning passage, so visible text arrives later than the first token, and some short replies spent their 96-token limit on reasoning and showed no text; each set's first-token figure is shown with its own visible-text note.

Every set ran on the same build with the same benchmark harness and protocol, each on its weights variant's shipped settings profile. Each set's receipt records its label, the combined digest of the three hosts' data manifests, the draft model, the settings profile and the engine commit.

The 16 prompts were written for this benchmark and contain no private data. Their full texts ship with the results as PROMPT-MIX.jsonl.

Two sets of weights, chosen at install.

Choose with WEIGHTS=ablit in cluster.env, or --weights ablit on fetch-weights.sh, split.sh and serve.sh (default base). Each weight variant ships its own measured settings profile.

Default

Base weights

GLM-5.3 Flash in 4-bit MLX format, with the model's own multi-token prediction head.

How you get them
The installer downloads them at the pinned revision, verifies every file against a pinned SHA-256 list, and splits them into three per-host parts, checked against a shipped manifest. This project hosts none of the v2.0.1 weights.
License
MIT
Download
181,741,759,037 bytes (181.7 GB; 54 files) for the weights; with the 2,342,460,697-byte draft model the download is 184,084,219,734 bytes (184.1 GB)
Your download
A fresh download and split on a DGX Spark that had never run this project matched the published per-host manifests file for file.

Chosen at install

Refusal-removed (abliterated) weights

An opt-in install for ablit development, red-teaming and refusal research.

Status
Ready in this release: scripts/fetch-weights.sh --weights ablit downloads the source with your own token, scripts/convert-ablit.sh converts it and writes the three per-host parts, and manifests/ablit/ checks a third you already have. We ran the shipped conversion on one DGX Spark, and its output matched these manifests file for file.
Access
Gated (auto-approve on the repo page): Hugging Face account, user accepts the source's terms themselves, user's own HF_TOKEN
How you get them
The installer downloads them with your own token at the pinned revision and converts them on your machine.
License
MIT (Copyright (c) 2026 Z.AI Co., Ltd), plus the use conditions on the source model card, quoted below.
Also built from
The pinned base weights, TensorFold/GLM-5.3-Flash-MLX-4bit-MTP at revision 76add2a341a1cd90ad0e86bb69839ea9c35827c6 (MIT), supply the four-bit tensor layout and the native prediction layer that the conversion restores. The chat template is the base checkpoint's MIT template plus six lines added by this recipe, also under MIT; its reconstructed bytes are hash-checked.
Conversion
On your own machine, scripts/convert-ablit.sh runs the pinned conversion scripts in scripts/ablit/ inside the release's pinned container image, with no network and no GPU, then writes and checks the three per-host parts.
It takes
Measured on one DGX Spark: the 200.1 GB source download took about an hour on our connection. Converting, splitting and checking then took about 26 minutes of processing, used no GPU and under 4 GiB of process memory, and needed about 371 GB of free disk beyond the downloaded source (about 571 GB in all), on top of the base weights you already installed. scripts/convert-ablit.sh checks for about 400 GB free before it starts.
Your conversion
A fresh conversion matches the measured data file for file, except two bookkeeping files on each host that record which conversion produced them.

You are responsible for complying with the source model's terms and for how you use this model and what it generates. JSpark3 provides conversion tooling only; it hosts none of these weights and does not endorse any use of them.

The source's terms, quoted from its model card
  • It is released strictly for legitimate research — interpretability, AI-safety and refusal-mechanism study, red-teaming, robustness evaluation, and controlled experiments.
  • You assume full responsibility and liability for how you use it and for everything it generates. Do not deploy it to end users or in production without adding your own safety, moderation, and abuse-prevention layers.

By downloading or using this model you acknowledge and accept the above.

The draft model and the licenses.

Draft model
incoai/GLM-5.3-Flash-DFlash2 @ bf582e4eacc1
Its license
CC BY-NC-ND 4.0 (non-commercial)
How you get it
Downloaded unmodified at install; never redistributed; requantized in memory at load only
Run without it
scripts/serve.sh --drafter none on all three hosts, or DRAFTER=none in cluster.env (default dflash2)
Commercial use

For commercial use, run the base weights without the draft model: start with scripts/serve.sh --drafter none on all three hosts (or set DRAFTER=none in cluster.env), and the model drafts with its own multi-token prediction head. That path runs MIT weights on a permissively licensed engine (MIT, with some Apache-2.0 code) and an Apache-2.0 recipe, inside NVIDIA's container under NVIDIA's terms. On the base weights, the draft model is the only non-commercial component.

For commercial use of the draft model itself, contact Inco AI through its model card.

Acceptance, base weights + draft model
On the prompt mix: 2.75 tokens accepted per verify step, 0.67 of drafted tokensOn a single repeated prompt: 2.11 tokens accepted per verify step, 0.61 of drafted tokens
Acceptance, refusal-removed (ablit) weights + draft model
On the prompt mix: 2.55 tokens accepted per verify step, 0.69 of drafted tokensOn a single repeated prompt: 2.04 tokens accepted per verify step, 0.55 of drafted tokens
Recipe
Apache-2.0
Engine
MIT (fork of TensorFold 0.3.6.2, MIT), with inherited third-party code under its own permissive licenses (MIT and Apache-2.0; engine THIRD_PARTY_NOTICES) and two Apache-2.0 lines ported from upstream TensorFold into cuda/server.py (engine NOTICE)
Base weights
MIT
Draft model
CC BY-NC-ND 4.0
Abliterated weights
MIT + source card use conditions; gated; fetched with the user's token, converted locally, not redistributed
Container image
NVIDIA terms (pulled by digest, not redistributed)
Fabric and chat template
No unlicensed inputs: the fabric launcher is this project's own recipe script (Apache-2.0); the chat template is the base model's stock MIT template (Z.AI notice kept) plus six lines added by this recipe, also under MIT, shipped in the recipe.

Recipe files are Apache-2.0. The engine (engine/) is MIT: a fork of TensorFold 0.3.6.2 (MIT). It includes third-party code under its own permissive licenses (MIT and Apache-2.0), listed in its THIRD_PARTY_NOTICES, and two lines ported from upstream TensorFold that stay Apache-2.0, credited in its NOTICE. The base weights are MIT. The refusal-removed weights are MIT plus their source card's use conditions. The chat template is Z.AI's MIT template plus six lines added by this recipe, also under MIT. The draft model is CC BY-NC-ND 4.0 (non-commercial); it is downloaded at install and never redistributed here. The NVIDIA container image is pulled from NGC under NVIDIA's terms.

What the endpoint supports.

OpenAI chat completions
Works: /v1/chat/completions returns the reply with finish_reason stop.
Streaming
Works: stream: true returns server-sent events ending in [DONE], with the same text as the non-streaming reply.
Tool calls
Works: a request with tools returns a tool_calls reply with JSON arguments and finish_reason tool_calls.
JSON schema output (response_format)
Ignored (known issue)
Reasoning effort
Always on; the lowest reasoning effort is low. v1.8.4 had it off by default.
Images
Both weight variants accept images in chat messages as inline data: URLs (base64). Remote image URLs are refused. Each request takes up to 16 images, at most 32 MB per image and 32 MB in total, and at most 32 megapixels per image. This release does not measure how well the model understands images.
Network
Loopback only, no auth, no CORS

The server listens on loopback (127.0.0.1) only, with no authentication and no CORS. Reach it through an SSH tunnel or a reverse proxy that adds authentication; don't expose the port.

Run v2.0.1

Follow INSTALL.md to install and run it on three DGX Sparks. The download, build and splitting steps were run from public sources on a Spark that had never run this project, and the packaged release was started and checked on three Sparks.

Hardware
Three NVIDIA DGX Sparks, connected by a direct high-speed (RDMA) link
Network
Cable the boxes' ConnectX-7 ports in a ring (rank 0 to rank 1, rank 1 to rank 2, rank 2 to rank 0), with each port up, RDMA working and MTU 9000; tensors travel over these cables. Every box also needs a shared LAN on which it can reach rank 0. The engine uses that LAN only to coordinate startup; tensor traffic stays on the cables.
Disk per host
  • Container image about 25 GB
  • Engine build under 50 MB
  • Download 181,741,759,037 bytes (181.7 GB; 54 files) for the weights; with the 2,342,460,697-byte draft model the download is 184,084,219,734 bytes (184.1 GB)
  • Split parts 63.9 GB for the first third, 62.8 GB for each of the other two
  • Kernel cache about 45 MB per host
  • Session cache up to 64 GiB, written only while 150 GiB stays free
  • After splitting the full download ($DATA/base/weights) is not needed to serve, so you may delete it.
Session cache
Conversation state, including the prompt's token ids, is cached on each host's own disk (up to 64 GiB per host) so returning to a long conversation is fast. It never leaves your machines. OPERATIONS explains where it lives, how to clear it and how to turn it off (SESSION_TIER=off).
From v1.8
v2.0.1 is a new installation, not an in-place upgrade. The engine changes from vLLM to a fork of TensorFold 0.3.6.2 (MIT), and the weights change from the v1.8.x EXL3 files to public 4-bit MLX-format weights split across the three hosts. Stop v1.8.x before you start v2.0.1, and keep your v1.8.4 checkout and weights if you might roll back.
Reasoning
Reasoning is now always on. v1.8.4 had it off by default and honoured requests to turn it off; v2.0.1 runs a request with no reasoning setting at High and treats a request to turn it off as low (known issue 9). Replies begin with a reasoning passage, returned as reasoning_content, before the visible text, and it uses part of max_tokens.
Who should stay on v1.8.4
v1.8.4 stays available; see rolling back. With base weights, two measured cases favour it (three DGX Sparks; v1.8.4 at its default, reasoning off, and v2.0.1 at reasoning effort low; an appliance comparison with different model IDs, not a same-weights claim). On prose replies, v1.8.4 shows the first visible text about 0.1 s sooner, because v2.0.1 writes a short reasoning passage first (known issue 9). With base weights and a single client on short code replies, the first visible text arrives in about the same time, with v1.8.4 about 2% faster. If you keep very many idle keep-alive clients connected, read known issue 13 first; a fix is planned for v2.0.2.
Rolling back
To roll back, stop v2.0.1 and start v1.8.4 from its tag, following its own installation guide. v1.8.4 is the documented rollback.

In your v1.8.4 checkout:

git fetch --tags && git checkout v1.8.4

From a fresh clone:

git clone --branch v1.8.4 https://github.com/jakejharris/jspark3.git jspark3-v1.8.4

What each step takes

Container image
Time about 35 minutesDisk about 25 GB
Engine build
Time about 7 secondsDisk under 50 MB
Base weights download
Time about 39 minutes, including the checksum check
Draft model download
Time about 31 seconds
Splitting into three parts
Time about 3.5 minutes per thirdDisk 63.9 GB for the first third, 62.8 GB for each of the other twoMemory about 4 GiB
First start
Time about 5 minutes for all three hosts, including compiling the kernelsKernel cache about 45 MB per host
Later starts
Time about 5 minutes

Download, build, conversion and disk figures were measured on one DGX Spark over our connection while another job shared its network and disk. Start times and the kernel cache were measured on the three-Spark cluster. Pull and download times depend on your connection.

Known issues.

  • stop is ignored. A reply ends at the model's end of turn or at max_tokens.
  • response_format is ignored. JSON mode and json_schema are not enforced, so the reply is free text.
  • Identical prompts without a seed return identical outputs, even above temperature 0: a chat app's regenerate returns the same reply, and two users who send the same prompt get the same answer. Send a different seed with each request when you want a different sample.
  • n, logprobs, presence and frequency penalties and logit_bias are ignored.
  • A wrongly typed field, such as a string temperature, may return HTTP 500 instead of 400.
  • Non-streaming requests send nothing until the reply is complete. Behind a proxy with an idle timeout, use stream: true.
  • The model field is not validated; every request is served by GLM-5.3 Flash.
  • Without the draft model, a long conversation that includes images may not be saved to the disk session cache, and each saved state takes more memory, so fewer long conversations stay cached. Returning to such a conversation after it has left the memory cache can take as long as its first prompt. Text-only conversations of about 40,000 tokens are saved; longer text-only conversations were not tested.
  • Reasoning is always on, and no setting turns it fully off; v1.8.4 had it off by default. A request that sets no reasoning effort runs at High, and the lowest effort is low; a top-level reasoning_effort: "none" and chat_template_kwargs: {"enable_thinking": false} are both treated as low, so a reply can still begin with a short reasoning passage. Even at low effort, a small max_tokens can be used up by reasoning and return no visible text; allow a few hundred tokens or more.
  • With the draft model on, a short text request that arrives while no reply is streaming may wait up to 25 ms for a second request before its prompt is read. Without the draft model, prompts that arrive together are not read together in one batch, and first tokens arrive later: with the base weights and 8 requests at once, the median first token arrives after 1.9 s, with visible text 0.41 s later (median over the replies that showed text), against 0.54 s with the draft model, where the median first visible text arrives at 1.5 s; in both sets, 3 of 24 short replies spent their 96-token limit on reasoning and showed no text. Reading prompts that arrive together in one batch without the draft model is planned for v2.0.2. With 16 requests at once, twice the server's 8 reply slots, total output with the draft model on is about 10 to 15% lower than with 8 (about 15% with the base weights, about 10% with the refusal-removed (ablit) weights), because the second eight prompts are read in small steps while the first eight replies stream. A streaming reply can occasionally pause between updates, and the pauses are longest while a long new prompt is being read: the longest measured pause was about 0.75 seconds, with a 36,180-token prompt. While a prompt of that length is being read, one step can pause every streaming reply at once for up to about 0.53 seconds.
  • JSpark3 v2.0.1 saves the state at the end of each prompt it reads, whichever client sent it, and reuses it when a later prompt starts with that entire earlier prompt, such as the next turn of a conversation; it then reads only the rest. Sharing only a system prompt is not enough: no state is saved where a system prompt ends, so a prompt with the same system prompt but a different first message is read in full. With the draft model on, it also skips reading a prompt that exactly repeats the latest prompt of a conversation, such as regenerating the latest reply, while that state is still in memory. With the draft model on, only the latest state of each conversation stays in memory, so regenerating or resending an earlier turn after later turns have been sent does not get this shortcut. Such a request resumes only from a shorter state that the disk session store has finished saving. The store saves in the background while the server is idle and may not yet hold a given turn, or may have skipped it; in testing, these regenerations read the whole prompt again. Without the draft model, an exact repeat is never skipped, but earlier turns' states can stay in memory until evicted, so regenerating a later turn can resume from the previous turn's state. Regenerating the first reply of a conversation after later turns, or resending it without the draft model, reads the whole prompt again, because saved state is reused only when it is shorter than the new prompt.
  • Conversations that share only a system prompt do not share cached work. Saved state is matched by prompt content, not by conversation: a prompt reuses an earlier prompt's state only when it starts with that entire earlier prompt, whichever conversation sent it. No state is saved at the end of a system prompt, so a new conversation that starts with the same system prompt as an earlier one, but has a different first message, reads its whole prompt again. A fix is planned for v2.0.2.
  • A client that disconnects while its connection's socket number is 1024 or higher is not detected, so its generation runs to completion and holds its slot. Normal connection counts do not reach this; very many idle keep-alive clients could. A fix is planned for v2.0.2.
  • If max_tokens cuts off a tool call, finish_reason is length (or tool_calls if an earlier call in the same reply was complete), the cut-off call is left out of the final tool_calls, and its raw text is returned in content. When streaming, its name and partial arguments (incomplete JSON) have already been sent. Raise max_tokens for tool use.
  • The usage block in replies does not include prompt_tokens_details.cached_tokens. The number of prompt tokens the server reused from saved state is reported in the reply's tensorfold.cached field instead (in the final chunk when streaming). For a request that forces a tool call, this count can be too high, even above the prompt's length. v1.8.4 returned this field, so a client that reads it must switch to tensorfold.cached when upgrading. A fix for both is planned for v2.0.2.
  • v2.0.1 does not support per-request cache isolation. It ignores the cache_salt request field, and all clients of one server share its saved prompt state. A request whose prompt starts with another client's entire earlier prompt reuses that state, which shows in the reply's cached-token count and in a faster first token. v1.8.4's engine honored cache_salt, so a deployment that relied on it to keep clients apart is no longer isolated after upgrading. If clients must not learn about each other's prompts, give each one its own server with its own session folder. Per-request isolation is planned for v2.0.2.

Errata

Two lines in the engine's own files are out of date and will be corrected in v2.0.2. engine/NOTICE says to select --drafter none --draft-policy c7:0.3. Don't add --draft-policy yourself: scripts/serve.sh --drafter none (or DRAFTER=none in cluster.env) applies the value from the weights' own settings profile, and the default base weights use c7:0.45. engine/README.md points at a RELEASE-FACTS.md file that does not ship; the measured results are in the README's Results section, docs/BENCHMARKS.md and release/MEASUREMENTS-v2.0.1.md.

Built on other people's work.

We tried DeepSeek, measured it, and came back.

Tempo was my DeepSeek experiment, and I measured it seriously. Its tok/s held up, but it overthinks, and time to finish a task is what I actually feel. GLM-5.3 Flash is better at agent and coding work, and better in almost every other way I use it, so the numbered line runs GLM again.

Tempo, the DeepSeek experiment →

Earlier releases.