How I prompt ML Intern 1. A model that knows your field 2. A model that draws your character 3. A model that does a new trick 4. A model that fits your device What it cost Make yours Last week, I wanted a small version of the prompt rewriter that ships with Qwen-Image 2.1. The official one is a 9B model that needs about 20 GB of memory and thinks for thousands of tokens before writing a single paragraph. On the Hub, I found only compressed copies of that same 9B model. So I described what I wanted to ML Intern, and the next day I had a 0.8B version that runs on a CPU. It returns valid output 99.7% of the time and uses about a quarter of the teacher's tokens. The compute for the whole project, including having the 9B model label 8,797 example requests, came to USD 16.
The model that didn't exist, so you made it yourself
We’re on a journey to advance and democratize artificial intelligence through open source and open science.

Key takeaways
A user created several models using ML Intern, resulting in smaller, efficient versions of existing models.
- The user created a 0.8B model using ML Intern, which runs on a CPU and is more efficient.
- Five additional models were created in the same way, each published on the Hub with evaluations.
- Prompts for the models were detailed and improved over time, with verified facts included to save resources.
Summarised automatically by AI from the original article by Hugging Face Blog. AI can make mistakes, so check the original for details.
Story details
- Published
- By
- yuvraj sharma, Abubakar Abid
- Format
- Article
- Original
- huggingface.co ↗




