Skip to content
context New models

BottleCap AI releases ThinkingCap-Qwen3.8-27B model with reduced thinking token consumption

only one source so far

The Czech company BottleCap AI has released the second model in the ThinkingCap series, a modified version of Alibaba's Qwen 3.8-27B model. According to the manufacturer, it consumes 37% fewer thinking tokens with slightly higher accuracy, and it is free on Hugging Face.

The Czech company BottleCap AI, backed by Tomáš Mikolov, Jaroslav Beck, and Stanislav Fořt, has released the second model in its ThinkingCap series. It is a modified version of the freely available Qwen model from the Chinese company Alibaba, specifically the Qwen 3.8-27B version; the new model is named ThinkingCap-Qwen3.8-27B. The first model in the series was released in the summer and was based on the previous Qwen 3.6-27B version.

According to BottleCap AI, compared to the standard Qwen, the new model consumes 37 percent fewer "thinking" tokens, with an average accuracy increase of 0.86 percentage points. According to the source, this figure ranges from 11 to 66 percent depending on the specific benchmark. The model is also said to retain context better than the standard version.

The model is available for free on Hugging Face under the Apache 2.0 license, in GGUF, FP8, and NVFP4 formats. BottleCap AI also offers paid services for the model.

What changed

Why it matters

Lower thinking token consumption while maintaining or slightly increasing accuracy means lower costs and faster responses for reasoning tasks built on the Qwen model. Since the model is freely available on Hugging Face under the Apache 2.0 license in multiple formats (GGUF, FP8, NVFP4), both developers and companies can immediately try it or deploy it locally without having to change the provider of the underlying model.

Release card

ThinkingCap-Qwen3.8-27B

BottleCap AI

open weights
Licence
PolyForm Small Business 1.0.0 and additional permission for personal use.
Price
Self-hosting costs. Commercial licensing under the terms of BottleCap AI; no uniform public price is listed.
Availability
Hugging Face after accepting the access terms. PolyForm Small Business 1.0.0 license and additional permission for personal use.
Documented measurements
  • Vendor measurement: token usage and accuracy Average accuracy of 85.8% versus 86.6% for the baseline model; an average 37% reduction in reasoning tokens. These results come from BottleCap AI, not an independent acceptance test. Benefits vary by task.
According to the sources, it is suitable for
  • Testing the model on your own hardware
  • Token usage comparison with the default model
Documented limits
  • This is not an unrestricted Apache 2.0 license; check the terms for your use.
  • Lower token usage does not mean higher accuracy on every task.

Data corrected according to the BottleCap AI model card on 27 September 2026.

The card summarizes information from the article and any dated corrections, with a link to the original source. It is not our assessment of the model. You can find a dated profile with supporting materials at model page. Model selection and other announcements →

Two audiences, two different impacts

What this means

01

For individuals

Developers and individuals working with open-source LLMs now have access to a free, downloadable model with lower token consumption at comparable accuracy, which can be deployed locally in several formats.

What to do Try the ThinkingCap-Qwen3.8-27B model for free on Hugging Face for tasks requiring reasoning.
More practical updates →
02

For a business

According to the manufacturer's claims, companies running reasoning models based on Qwen can reduce thinking token consumption during inference, which translates into operating cost savings, without having to change model providers.

Development
What to decide Consider trying the ThinkingCap-Qwen3.8-27B model or the paid services from BottleCap AI to reduce inference costs for tasks using Qwen.
More business impacts →
Alibaba BottleCap AI Hugging Face open-source Qwen ThinkingCap-Qwen3.8-27B

Check the original

Event sources

only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.

1
Lupa.cz independent context · first detected Tomáš Mikolov a Jaroslav Beck vydali druhý AI model. Opět výrazně snižuje „žravost“