Know your model.
Know your machine.
Benchmark GGUF or MLX models, choose 1–3 concurrent jobs calling one model, and save results automatically to your TokFire account.
Run a reproducible test
- Install the desktop app
Mac: download the DMG and drag to Applications. Windows: extract the ZIP and open TokFire Bench.exe. Install Python and llama.cpp separately; MLX/oMLX is Mac-only. Windows account sign-in also requires WebView2.
- Choose GGUF or MLX
Choose one model and select 1, 2 or 3 simultaneous jobs before Run. Both llama.cpp and oMLX show live progress and per-job results. TokFire Bench Pro unlocks up to 20 jobs on the same model. HK$180 once, one activated device. Paid release coming after store approval.
- Connect once, then run
Automatic upload is enabled before Run; public sharing is separate and off by default. Offline or failed uploads remain queued. JSON and readable commentary are always saved locally.
Windows 10/11 Intel or AMD 64-bit. Extract the entire ZIP. Core controls offer 20 languages; Windows-specific guidance and reports are currently in English. This preview has been cross-compiled and unit-tested on macOS; Windows runtime testing is pending.
The download includes a trial launcher. In the extracted native folder, run python3 trial-minicpm.py. It downloads the official Q4_K_M model (~1.56 GB), verifies its checksum, and measures one short run on your Mac. Requires Python and llama-server; no Xcode build is needed for the command-line trial. The CLI trial stays local; trials run in the app follow its upload setting. The full benchmark also supports one model, or up to three in sequence.