Author uses ML Intern to create a smaller model for rewriting prompts
The author used ML Intern to create a prompt-rewriting model with 0.8 billion parameters that, according to the author, runs on a CPU. The compute costs of the project amounted to 16 USD; the author reports 99.7 % valid outputs and approximately a quarter of the token usage.
The author of a post on the Hugging Face blog described creating a smaller model for rewriting prompts for Qwen-Image 2.1 using ML Intern. According to the author, the original model has 9 billion parameters, requires approximately 20 GB of memory and consumes thousands of reasoning tokens before generating a paragraph. The new variant has 0.8 billion parameters and, according to the author, runs on a CPU.
The original model labeled 8 797 example requests to prepare training data. The author reports that the smaller model returns valid output in 99.7 % of cases and uses approximately a quarter of the number of tokens used by the original model. These are results reported by the author; output validity is not the same as the quality of the rewritten prompt. According to the author, the compute costs of the entire project, including labeling the examples, amounted to 16 USD.
The described process starts with a prompt in HuggingChat with ML Intern enabled. The tool plans the work, requests a budget, performs a small validation run and then handles training, evaluation and publication of the model on Hugging Face. The author recommends specifying the dataset, starting model, training script, measurements before training and spending cap in the prompt. The author published the prompts in the yvrjsharma/ml-intern-prompts repository.
Why it matters
This case shows a concrete process for building a smaller model for a narrow task using examples generated by a larger model. The reported ability to run on a CPU is practical for use with limited hardware. However, assessing whether it can replace the original model requires checking the quality of the rewritten prompts as well as output validity.
Two audiences, two different impacts
What this means
For individuals
A developer who needs to rewrite prompts for Qwen-Image 2.1 gets a description of how to create a smaller variant, along with published prompts to use as a starting point for their own experiment.
For a business
For a team developing specialized models, this case illustrates a way to control compute spending: an approved budget and a small validation run before full training. The reported 16 USD represents the compute costs of this project, not the total costs of a business deployment.
DevelopmentCheck the original
Event sources
only one source so far · 1 publisher, 0 independent. We count feeds from the same owner only once.