All systems operational
Live status and speed of the Umans Code gateway and its models, refreshed every 30 seconds. For each model we show its output speed and median time to first token.
… / 0production models operational
…gateway uptime · 90 days
0models in testing
…active issues
Gateway
API endpoint
API gateway Operational
api.code.umans.ai
…
service uptime · 24h
…
median TTFT · all models
Production
Models in production
tok/s = output tokens per second · TTFT = time to first token · p50 = median over the last 5 minutes. On each gauge the midpoint is that model's target; the marker sits further right when it's beating target (faster TTFT, higher throughput).
Model details are temporarily unavailable.
History
Gateway uptime
Changelog
Recent events & model lifecycle
Aug 32026
Released pay-per-token: Umans DeepSeek V4 Flash Released
umans-deepseek-v4-flash-0731 joins the lineup as the cheapest way we serve real agentic work: $0.14 / $0.28 / $0.028 per 1M (input / output / cache read), a 1M context window, thinking at low effort by default (dial up high or max when a task deserves more). It is the new default for new chats and CLI setups. Founding users pay the 10x cheaper cache rate until Monday, August 10, 2026 (see /pricing). Served on our own GPU infrastructure with high availability.
Aug 32026
The V4 Flash lab continues as umans-deepseek-v4-flash-0731-lab Testing
The V4 Flash pilot closed at the pay-per-token release: the production id now bills per token, so the seat-gated pilot on it ended rather than charge anyone by surprise. The lab reopens on the new umans-deepseek-v4-flash-0731-lab id with a smaller cohort - free, seat-gated, same experimental capacity as before. The model keeps serving as umans-deepseek-v4-flash-0731 regardless: the model stays.
Aug 12026
Generally available: Umans Kimi K3 Released
The Kimi K3 Labs pilot closed and umans-kimi-k3 is generally available: no Labs seat required anymore - plan and pay-per-token keys alike get the 1M context window, native vision, and max-effort reasoning at $3.00 / $15.00 / $0.30 per 1M (input / output / cache read).
Jul 312026
Released pay-per-token: Umans Kimi K3 Released
umans-kimi-k3 opened to pay-per-token (wallet and service-account keys): Moonshot's largest open model, a 1M context window, native vision, and max reasoning effort by default, billed at $3.00 / $15.00 / $0.30 per 1M (input / output / cache read).
Jul 162026
Playground closed: Umans DeepSeek V4 Pro DSpark Testing
The DSpark test window closed after two days. What we took from it: DSpark speculative decoding lets us serve more users at once while keeping each session fast enough, the DeepSeek V4 architecture is now mature enough to serve at scale, and the model itself is solid. V4 Pro is not joining the lineup though: it is still a preview build, and the issues testers hit (DSML leaks, language bleed, long-context artifacts) are model-side. DeepSeek confirmed a better version is coming this month, so we would rather roll the learnings into that. Thanks to everyone who tested.
Jul 142026
Playground opened: Umans DeepSeek V4 Pro DSpark Testing
umans-deepseek-v4-pro-dspark entered the playground for a short, seat-gated test window: DeepSeek V4 Pro served from the original weights with DSpark speculative decoding. Experimental and temporary; not for production.
Jul 22026
Playground closed: Umans GLM 5.2 NVFP4 Testing
The short NVFP4 test window ended after four days. Thanks to everyone who pushed it and shared findings.
Jun 292026
Playground opened: Umans GLM 5.2 NVFP4 Testing
umans-glm-5.2-nvfp4 entered the playground for a short, low-capacity test window. Experimental and temporary; not for production.
Jun 242026
Retired: Umans GLM 5.1 Retired
umans-glm-5.1 was retired in favour of GLM 5.2. Requests to the old id now return a clear deprecation error pointing to umans-glm-5.2.
Jun 212026
Released to production: Umans GLM 5.2 Released
umans-glm-5.2 was released as the long-context model, with a 405K context window, after a pre-release period that started Jun 16.
Jun 182026
Retired: Umans Kimi K2.6 Code Retired
umans-kimi-k2.6 was retired and superseded by K2.7. Requests to the old id now return a clear deprecation error pointing to umans-kimi-k2.7.
Jun 122026
Released to production: Umans Kimi K2.7 Code Released
umans-kimi-k2.7 was released as the recommended coding model (also served as umans-coder).
May 132026
Retired: Umans Kimi K2.5 Retired
umans-kimi-k2.5 was retired and superseded by Kimi K2.6.
May 92026
Retired: Umans MiniMax M2.5 Retired
umans-minimax-m2.5 was retired without a direct replacement.
Past models 6 retired
Umans DeepSeek V4 Pro DSpark Retired
umans-deepseek-v4-pro-dspark · DeepSeek
Umans GLM 5.2 NVFP4 Retired
umans-glm-5.2-nvfp4 · GLM
Umans GLM 5.1 Retired
umans-glm-5.1 · GLM
Umans Kimi K2.6 Code Retired
umans-kimi-k2.6 · Moonshot
Umans Kimi K2.5 Retired
umans-kimi-k2.5 · Moonshot
Umans MiniMax M2.5 Retired
umans-minimax-m2.5 · MiniMax