Static site · Sign in, upload reports and view live comparisons on the main portal. Open the live portal →
MACOS · WINDOWS
DESKTOP PREVIEWSign in / Register
TOKFIRE BENCH / BY TOKFIRE LABS

Know your model.
Know your machine.

Benchmark GGUF or MLX models, choose 1–3 concurrent jobs calling one model, and save results automatically to your TokFire account.

1–3FREE CONCURRENT JOBS
Tests run in the desktop app. The website stores results and attributed oMLX reference data. TokFire Bench has been tested on an M2 Max; the developer DMG is ad-hoc signed and not notarized. Windows x64 is an unsigned preview; native Windows and GPU validation are still pending.
01 / ON YOUR DEVICE

Run a reproducible test

  1. Install the desktop app

    Mac: download the DMG and drag to Applications. Windows: extract the ZIP and open TokFire Bench.exe. Install Python and llama.cpp separately; MLX/oMLX is Mac-only. Windows account sign-in also requires WebView2.

  2. Choose GGUF or MLX

    Choose one model and select 1, 2 or 3 simultaneous jobs before Run. Both llama.cpp and oMLX show live progress and per-job results. TokFire Bench Pro unlocks up to 20 jobs on the same model. HK$180 once, one activated device. Paid release coming after store approval.

  3. Connect once, then run

    Automatic upload is enabled before Run; public sharing is separate and off by default. Offline or failed uploads remain queued. JSON and readable commentary are always saved locally.

Download macOS DMG · 0.5.2Download macOS sourceWindows x64 ZIP · 0.6.0 Preview

Windows 10/11 Intel or AMD 64-bit. Extract the entire ZIP. Core controls offer 20 languages; Windows-specific guidance and reports are currently in English. This preview has been cross-compiled and unit-tested on macOS; Windows runtime testing is pending.

Start small: MiniCPM5-2B trial

The download includes a trial launcher. In the extracted native folder, run python3 trial-minicpm.py. It downloads the official Q4_K_M model (~1.56 GB), verifies its checksum, and measures one short run on your Mac. Requires Python and llama-server; no Xcode build is needed for the command-line trial. The CLI trial stays local; trials run in the app follow its upload setting. The full benchmark also supports one model, or up to three in sequence.

Read the methodology & limitationsoMLX reference library →Connect the desktop app →