What makes Inkling special? Overall Capabilities and Architecture Inference Support Transformers SGLang vLLM Remote Inference with Hugging Face Inference Providers Local Inference with llama.cpp and Unsloth Use Cases Agentic coding with Pi Multi Token Prediction Drafters Multimodal Vision Multimodal Audio Post-training Deploying Inkling and Inkling-Small Deploying Inkling on a cluster Deploying Inkling-Small on Inference Endpoints Benchmark Results Inkling now comes in smaller size 🤗 Inkling-Small is out by Thinking Machines Lab. We have updated this post with performance, and deployment configurations for the Inkling-Small and the Inkling-Small-NVFP4 variants. Here’s the collection with all the Inkling models. We made it easier for you to deploy Inkling-Small with one-click on Inference Endpoints (getting up to 160 TPS). We also ship a real-time voice and image demo where you can interact with the model.
Welcome Inkling by Thinking Machines
We’re on a journey to advance and democratize artificial intelligence through open source and open science.

Key takeaways
Inkling is a large multimodal LLM with 1T parameters, supporting image, text, and audio inputs, and agentic capabilities. It is available on Hugging Face with day-0 support in transformers and other inference engines.
- Inkling has 1T parameters and 1M context window.
- It supports image, text, and audio inputs.
- Inkling uses relative attention and hybrid attention.
- It includes a short convolution and MoE with shared experts.
Summarised automatically by AI from the original article by Hugging Face Blog. AI can make mistakes, so check the original for details.
Story details
- Published
- By
- ben burtenshaw, merve, Pedro Cuenca
- Format
- Article
- Original
- huggingface.co ↗



