BottleCap AI releases ThinkingCap-Qwen3.8-27B model with reduced thinking token consumption
The Czech company BottleCap AI has released the second model in the ThinkingCap series, a modified version of Alibaba's Qwen 3.8-27B model. According to the manufacturer, it consumes 37% fewer thinking tokens with slightly higher accuracy, and it is free on Hugging Face.
The Czech company BottleCap AI, backed by Tomáš Mikolov, Jaroslav Beck, and Stanislav Fořt, has released the second model in its ThinkingCap series. It is a modified version of the freely available Qwen model from the Chinese company Alibaba, specifically the Qwen 3.8-27B version; the new model is named ThinkingCap-Qwen3.8-27B. The first model in the series was released in the summer and was based on the previous Qwen 3.6-27B version.
According to BottleCap AI, compared to the standard Qwen, the new model consumes 37 percent fewer "thinking" tokens, with an average accuracy increase of 0.86 percentage points. According to the source, this figure ranges from 11 to 66 percent depending on the specific benchmark. The model is also said to retain context better than the standard version.
The model is available for free on Hugging Face under the Apache 2.0 license, in GGUF, FP8, and NVFP4 formats. BottleCap AI also offers paid services for the model.
Why it matters
Lower thinking token consumption while maintaining or slightly increasing accuracy means lower costs and faster responses for reasoning tasks built on the Qwen model. Since the model is freely available on Hugging Face under the Apache 2.0 license in multiple formats (GGUF, FP8, NVFP4), both developers and companies can immediately try it or deploy it locally without having to change the provider of the underlying model.
Release card
ThinkingCap-Qwen3.8-27B
BottleCap AI
- Licence
- PolyForm Small Business 1.0.0 and additional permission for personal use.
- Price
- Self-hosting costs. Commercial licensing under the terms of BottleCap AI; no uniform public price is listed.
- Availability
- Hugging Face after accepting the access terms. PolyForm Small Business 1.0.0 license and additional permission for personal use.
- Vendor measurement: token usage and accuracy Average accuracy of 85.8% versus 86.6% for the baseline model; an average 37% reduction in reasoning tokens. These results come from BottleCap AI, not an independent acceptance test. Benefits vary by task.
- Testing the model on your own hardware
- Token usage comparison with the default model
- This is not an unrestricted Apache 2.0 license; check the terms for your use.
- Lower token usage does not mean higher accuracy on every task.
Data corrected according to the BottleCap AI model card on 27 September 2026.
The card summarizes information from the article and any dated corrections, with a link to the original source. It is not our assessment of the model. You can find a dated profile with supporting materials at model page. Model selection and other announcements →
Two audiences, two different impacts
What this means
For individuals
Developers and individuals working with open-source LLMs now have access to a free, downloadable model with lower token consumption at comparable accuracy, which can be deployed locally in several formats.
For a business
According to the manufacturer's claims, companies running reasoning models based on Qwen can reduce thinking token consumption during inference, which translates into operating cost savings, without having to change model providers.
DevelopmentCheck the original
Event sources
only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.