What makes Inkling special? Overall Capabilities and Architecture Inference Support Transformers SGLang vLLM Remote Inference with Hugging Face Inference Providers Local Inference with llama.cpp and Unsloth Use Cases Agentic coding with Pi Multi Token Prediction Drafters Multimodal Vision Multimodal Audio Post-training Deploying Inkling and Inkling-Small Deploying Inkling on a cluster Deploying Inkling-Small on Inference Endpoints Benchmark Results Inkling now comes in smaller size 🤗 Inkling-Small is out by Thinking Machines Lab. We have updated this post with performance, and deployment configurations for the Inkling-Small and the Inkling-Small-NVFP4 variants. Here’s the collection with all the Inkling models. We made it easier for you to deploy Inkling-Small with one-click on Inference Endpoints (getting up to 160 TPS). We also ship a real-time voice and image demo where you can interact with the model.