Decode
Research and announcements from the Inco AI team.
Subscribe and Decode with us- September 17, 2026
Splash: A Local Engine Built Around the Model
Splash is our open-source inference engine for Apple silicon, built around the model rather than around a model zoo. On a 48 GB M5 Pro it generates 210 tokens/s on Qwen3.6-35B-A3B, reopens a cached 32K context in 123 ms, and serves four concurrent requests at 357 tokens/s combined.
- September 3, 2026
Inco AI Launches Its Inference Platform, Leading Across Four Open Models on Artificial Analysis
Kimi K3, MiniMax M3, GLM 5.3, and GLM 5.3 Flash lead their respective Artificial Analysis provider leaderboards as the Inco platform enters public beta.
- August 28, 2026
Inco AI launches Day-0 support for GLM 5.3
Inco AI is releasing day-0 support for GLM 5.3 alongside DFlash 2 and NVFP4 checkpoints for faster, more efficient inference.
- August 18, 2026
DFlash 2: Keep Drafting Parallel
DFlash 2 is the successor to our widely deployed parallel drafter: close to 3× the speed of autoregressive decoding, with the same output. Drafters for Qwen3.8-27B and Meta's Muse Glimmer are out today.