Faster LLM inference.
Same model files, same answers.
Sushila.cpp is a llama.cpp-based engine that does the expensive work once per model, right after the model is released, so every token you generate afterwards costs less. Llama-3.3-70B runs 3.6× faster than vanilla Ollama on the same GPU; on the very same model file, Llama-3.1-70B runs 2.0× faster than stock llama.cpp on a GPU and 2.3× on a CPU, with exactly its output.


What Sushila.cpp does
Decoding is limited by how many bytes of weights the hardware reads per token. Sushila (Scalable Upstream Synthesis for Hybrid Inference in Large-model Acceleration) computes small artifacts once per model so each token reads fewer bytes or the model runs fewer passes.
Landscapes
A precomputed map of each model's output layer: a cheap preview picks a short list of candidate tokens, which are then scored exactly. About 13–15% of the layer is read, with the same top token.
Precomputed draft heads
A small head fitted on the model's own answers proposes several tokens; the full model checks them all in one pass. Accepted tokens are exactly what the model would have produced.
Landscape hunt
Each model gets the stack that suits it. New models first try the existing landscapes, and a new one is searched for only if none fits.
Kernels
Tree verification in llama.cpp and a tuned GPU kernel switch make checking 5–8 drafted tokens 18–38% cheaper.
Sushila.cpp also uses established methods, including EAGLE-3 draft heads, small draft models and fast 4-bit kernels, and combines them with its own. The paper credits each method and reports how much it adds.
Measured speed
Tokens per second, greedy decoding, compared with the stock engine on the same hardware and model file.
| Model | Hardware, engine | Stock | Sushila | Speedup | How |
|---|---|---|---|---|---|
| Llama 3.3 70B, 4-bit | A100, vs. vanilla Ollama | 21.6 | 77.4 | 3.58× | SGLang + precomputed draft head (4.0× on MT-Bench, HumanEval, GSM8K) |
| Llama 3.3 70B, 4-bit | A100, SGLang | 33.8 | 77.4 | 2.29× | precomputed draft head on EAGLE-3 trees |
| Llama 3.1 70B, 4-bit | A100, llama.cpp | 22.2 | 44.9 | 2.02× | 1B draft model, chosen for this model |
| Llama 3.1 70B, 4-bit | CPU, 30 threads | 2.73 | 6.22 | 2.28× | 1B draft model, chosen on day 0 |
| Llama 3.1 8B, 16-bit | A100, SGLang | 89.5 | 193 | 2.29× | precomputed draft head on EAGLE-3 trees |
| Llama 3.1 8B, 4-bit | A100, llama.cpp | 153 | 199 | 1.30× | precomputed draft head, tree verification, kernel setting |
| Llama 3.1 8B, 4-bit | CPU, 30 threads | 19.9 | 24.0 | 1.21× | precomputed draft head with a landscape inside it |
| Qwen2.5 0.5B, 4-bit | CPU, 8 threads | 145.5 | 183.3 | 1.26× | output-layer landscape |
Same model file and exactly the stock output in every row except the Ollama comparison, where Ollama reads a different 4-bit file of the same model (Q4_K_M against AWQ; same GSM8K accuracy). The full method, scripts and raw logs are in the repository.
Download Sushila.cpp
Free and open source under the MIT License, like llama.cpp and Ollama. It reads the same GGUF files, including the
ones Ollama has already downloaded, and has the same tools: llama-cli, llama-server (OpenAI-compatible API) and more.
Ubuntu, Debian, Fedora and others; CPU or NVIDIA GPU. Tested on Ubuntu 22.04.
# 1. tools (Ubuntu/Debian; on Fedora: sudo dnf install gcc gcc-c++ make cmake git curl)
sudo apt-get install -y build-essential cmake git curl
# 2. download
git clone https://github.com/syncaissa/sushila.cpp.git && cd sushila.cpp
# 3. build (for an NVIDIA GPU add -DGGML_CUDA=ON; needs the CUDA toolkit)
cmake -S llama.cpp -B build -DCMAKE_BUILD_TYPE=Release -DLLAMA_CURL=OFF -DLLAMA_USE_PREBUILT_UI=OFF
cmake --build build --config Release -j --target llama-cli llama-server llama-speculative-simple
# 4. run
build/bin/llama-server -m model.gguf -ngl 99 --port 8080macOS 13 or newer, Apple Silicon or Intel. On Apple Silicon the GPU (Metal) is used automatically.
# 1. tools
xcode-select --install
brew install cmake git
# 2. download
git clone https://github.com/syncaissa/sushila.cpp.git && cd sushila.cpp
# 3. build
cmake -S llama.cpp -B build -DCMAKE_BUILD_TYPE=Release -DLLAMA_CURL=OFF -DLLAMA_USE_PREBUILT_UI=OFF
cmake --build build --config Release -j --target llama-cli llama-server llama-speculative-simple
# 4. run
build/bin/llama-server -m model.gguf -ngl 99 --port 8080Windows 10/11 through WSL2 (real Linux inside Windows; NVIDIA GPUs work). A native Windows build is planned.
# 1. in PowerShell as administrator, then restart
wsl --install
# 2. open "Ubuntu" from the Start menu and continue there
sudo apt-get update && sudo apt-get install -y build-essential cmake git curl
git clone https://github.com/syncaissa/sushila.cpp.git && cd sushila.cpp
# 3. build (NVIDIA GPU: install the Windows NVIDIA driver, add -DGGML_CUDA=ON)
cmake -S llama.cpp -B build -DCMAKE_BUILD_TYPE=Release -DLLAMA_CURL=OFF -DLLAMA_USE_PREBUILT_UI=OFF
cmake --build build --config Release -j --target llama-cli llama-server llama-speculative-simple
# 4. run
build/bin/llama-server -m model.gguf -ngl 99 --port 8080Already use Ollama? Reuse its files: ollama show --modelfile llama3.1:8b | grep '^FROM /' prints the model's path.
Full guide: INSTALL.md. Provided as is, without warranty (disclaimer).
Models
We host these files ourselves. Downloads are free with a sushila.ai account (no password: a code is sent to your e-mail), so we can record that you accepted each model's license. Each file is byte-identical to the public release, so check its sha256 after you download it. Models tagged precomputed come with measured Sushila artifacts.
| Model | Quant | Size | License | |
|---|---|---|---|---|
| Llama 3.1 8B Instruct precomputed | Q4_K_M | 4.9 GB | Llama 3.1 Community License | Download |
| Llama 3.1 8B Instruct precomputed | Q8_0 | 8.5 GB | Llama 3.1 Community License | Download |
| Llama 3.3 70B Instruct precomputed | Q4_K_M | 43 GB | Llama 3.3 Community License | Download |
| Llama 3.2 3B Instruct | Q4_K_M | 2.0 GB | Llama 3.2 Community License | Download |
| Llama 3.2 1B Instruct draft model for the 70B |
Q8_0 | 1.3 GB | Llama 3.2 Community License | Download |
| Qwen2.5 7B Instruct precomputed | Q4_K_M | 4.7 GB | Apache-2.0 | Download |
| Qwen3 30B-A3B (MoE) precomputed | Q4_K_M | 19 GB | Apache-2.0 | Download |
Downloads are provided as is, with no warranty; you assume all risks of use (see the disclaimer).
Built with Llama. The Llama models are distributed under their community licenses and Meta's acceptable use policy.
More models: download from Hugging Face
Any GGUF model runs with Sushila.cpp. Download these directly from their publishers.
| Model | License | |
|---|---|---|
| Qwen3 8B | Apache-2.0 | Hugging Face ↗ |
| Qwen3 32B | Apache-2.0 | Hugging Face ↗ |
| Qwen2.5 Coder 7B Instruct | Apache-2.0 | Hugging Face ↗ |
| Mistral 7B Instruct v0.3 | Apache-2.0 | Hugging Face ↗ |
| Gemma 2 9B Instruct | Gemma Terms of Use | Hugging Face ↗ |
| Phi-3 Mini 4K Instruct | MIT | Hugging Face ↗ |
| DeepSeek-R1 Distill Llama 8B | MIT + Llama 3.1 License | Hugging Face ↗ |
| Llama 3.1 70B Instruct | Llama 3.1 Community License | Hugging Face ↗ |
Serverless API
This website and Sushila.cpp are free. To skip the hardware, use our hosted GPUs: an OpenAI-compatible API, billed per token. Because each token costs us less to generate, our rates are lower.
Free
Sushila.cpp, the models above and every script, running on your own machine.
Pay per token
Hosted open models on serverless GPUs. No servers to run and no minimum spend.
Dedicated
Reserved capacity and custom day-0 tuning for your own model. Contact us.
The API is provided as is, with no warranty or guarantee of availability; see the disclaimer, the Terms of Service and the Privacy Policy. We use your e-mail only to contact you about early access.
Disclaimer
Sushila is a research project. Sushila.cpp, the precomputed landscapes and draft heads, the benchmarks, the hosted model files and the serverless API are research software and research results, published so that others can study, reproduce and build on them. They are experimental, may change or stop at any time, and are not a finished commercial product.
Use at your own risk. Sushila.cpp, the model files, scripts, benchmarks, the serverless API and everything else on this website are provided "as is" and "as available", without warranty of any kind, express or implied. This includes, without limitation, any warranty of merchantability, fitness for a particular purpose, accuracy, reliability, availability, security or non-infringement. No warranty is implied or given by anything on this website or in any communication from the Sushila project.
By downloading or using any of it, you assume all risks of that use, including the risk of incorrect, harmful or offensive model output, data loss, hardware or system damage, security issues and costs. You are responsible for checking what the software and models produce before you rely on it, and for complying with each model's license and acceptable use policy.
To the fullest extent permitted by law, the Sushila project, its contributors and its suppliers are not liable for any direct, indirect, incidental, special, consequential or punitive damages, or any loss of data, profits or business, arising from or related to the use of, or inability to use, anything provided here, even if advised of the possibility of such damages.
Speed figures are measurements on specific hardware and settings; your results may differ. Models are made by third parties and are governed by their own licenses; Sushila does not endorse or take responsibility for their content. Some jurisdictions do not allow certain warranty exclusions or liability limits, in which case they apply only as far as the law allows.