Run v2.0.1
Follow INSTALL.md to install and run it on three DGX Sparks. The download, build and splitting steps were run from public sources on a Spark that had never run this project, and the packaged release was started and checked on three Sparks.
- Hardware
- Three NVIDIA DGX Sparks, connected by a direct high-speed (RDMA) link
- Network
- Cable the boxes' ConnectX-7 ports in a ring (rank 0 to rank 1, rank 1 to rank 2, rank 2 to rank 0), with each port up, RDMA working and MTU 9000; tensors travel over these cables. Every box also needs a shared LAN on which it can reach rank 0. The engine uses that LAN only to coordinate startup; tensor traffic stays on the cables.
- Disk per host
- Container image about 25 GB
- Engine build under 50 MB
- Download 181,741,759,037 bytes (181.7 GB; 54 files) for the weights; with the 2,342,460,697-byte draft model the download is 184,084,219,734 bytes (184.1 GB)
- Split parts 63.9 GB for the first third, 62.8 GB for each of the other two
- Kernel cache about 45 MB per host
- Session cache up to 64 GiB, written only while 150 GiB stays free
- After splitting the full download ($DATA/base/weights) is not needed to serve, so you may delete it.
- Session cache
- Conversation state, including the prompt's token ids, is cached on each host's own disk (up to 64 GiB per host) so returning to a long conversation is fast. It never leaves your machines. OPERATIONS explains where it lives, how to clear it and how to turn it off (SESSION_TIER=off).
- From v1.8
- v2.0.1 is a new installation, not an in-place upgrade. The engine changes from vLLM to a fork of TensorFold 0.3.6.2 (MIT), and the weights change from the v1.8.x EXL3 files to public 4-bit MLX-format weights split across the three hosts. Stop v1.8.x before you start v2.0.1, and keep your v1.8.4 checkout and weights if you might roll back.
- Reasoning
- Reasoning is now always on. v1.8.4 had it off by default and honoured requests to turn it off; v2.0.1 runs a request with no reasoning setting at High and treats a request to turn it off as low (known issue 9). Replies begin with a reasoning passage, returned as
reasoning_content, before the visible text, and it uses part of max_tokens. - Who should stay on v1.8.4
- v1.8.4 stays available; see rolling back. With base weights, two measured cases favour it (three DGX Sparks; v1.8.4 at its default, reasoning off, and v2.0.1 at reasoning effort low; an appliance comparison with different model IDs, not a same-weights claim). On prose replies, v1.8.4 shows the first visible text about 0.1 s sooner, because v2.0.1 writes a short reasoning passage first (known issue 9). With base weights and a single client on short code replies, the first visible text arrives in about the same time, with v1.8.4 about 2% faster. If you keep very many idle keep-alive clients connected, read known issue 13 first; a fix is planned for v2.0.2.
- Rolling back
- To roll back, stop v2.0.1 and start v1.8.4 from its tag, following its own installation guide. v1.8.4 is the documented rollback.
In your v1.8.4 checkout:
git fetch --tags && git checkout v1.8.4
From a fresh clone:
git clone --branch v1.8.4 https://github.com/jakejharris/jspark3.git jspark3-v1.8.4
What each step takes
- Container image
- Time about 35 minutesDisk about 25 GB
- Engine build
- Time about 7 secondsDisk under 50 MB
- Base weights download
- Time about 39 minutes, including the checksum check
- Draft model download
- Time about 31 seconds
- Splitting into three parts
- Time about 3.5 minutes per thirdDisk 63.9 GB for the first third, 62.8 GB for each of the other twoMemory about 4 GiB
- First start
- Time about 5 minutes for all three hosts, including compiling the kernelsKernel cache about 45 MB per host
- Later starts
- Time about 5 minutes
Download, build, conversion and disk figures were measured on one DGX Spark over our connection while another job shared its network and disk. Start times and the kernel cache were measured on the three-Spark cluster. Pull and download times depend on your connection.