Skip to content
worth noting New models

Alibaba introduced the model Qwen3.8-Max with 2.4 trillion parameters for autonomously completing long-term tasks

only one source so far

Alibaba has published the model Qwen3.8-Max (2.4 trillion parameters, 95 billion active). According to the company, it autonomously created a tool in 16 days (265 commits), reproduced and surpassed the results of a scientific paper, and performed successfully in a Tianchi competition against 526 teams. The weights will be released next week.

Alibaba has published the model Qwen3.8-Max, which the company describes as its most capable language model to date. The model has 2.4 trillion total parameters, of which 95 billion are active per query, and builds on the Qwen3.5 architecture. According to the company, the model is focused on independently completing complex tasks over extended periods, rather than just answering individual queries. The company first announced the model in mid-July as a preview version available through Token Plan, Qoder and QoderWork at 10 % of the standard price; it is the first model in the Qwen-Max class whose weights are to be released, with the release planned for next week.

Alibaba presented three case studies in which the model reportedly worked without human assistance. In the first, the model created the command-line tool oh-my-cli over 16 days — it accepted user requests, converted them into tasks on GitHub, assigned them to itself, wrote code, ran tests and iteratively improved the results; by 30 July 2026, it had accumulated 265 commits, 127 pull requests and 151 issues without a single human intervention. In the second study, the model was given the scientific paper “Unified Data Selection for LLM Reasoning" without any initial code and tasked with reproducing and then improving its results; over approximately five days and 125 hours of compute time, it wrote 7 600 lines of code, ran 33 training jobs on GPUs, reproduced all six main results of the paper and, after testing 18 ideas of its own across four rounds, outperformed the method from the original paper on the AIME24 benchmark by 2.7 points.

In the third study, the model took part in the WWW2025 Multimodal Dialogue Intent Recognition Challenge on the Tianchi platform, where 526 human teams competed. Within 24 hours, the model fine-tuned several Chinese language models along with Qwen2.5-VL-7B for recognizing product screenshots and combined them into a voting system; accuracy across 45 submitted solutions rose from 0.60 to 0.853, which, according to the company, put the model ahead of 458 of the 526 competing human teams.

The source does not contain the full text of the article; the remaining portion is missing. Details can be found in the source article.

What changed

Why it matters

If the described capabilities are independently confirmed, this is a model capable of working on complex tasks (software development, research reproduction) for days without human intervention, available as an open-weight model starting next week. This is relevant to developers and companies considering deploying autonomous AI agents on their own infrastructure, because open weights enable independent verification and local operation instead of reliance on closed APIs. The figures cited (benchmarks, competition results) come from internal tests conducted by Alibaba and have not yet been independently verified.

Two audiences, two different impacts

What this means

01

For individuals

Starting next week, when the model weights are scheduled for release, developers will be able to try an open-weight model claimed to be capable of working independently on tasks such as tool creation or research reproduction for several days.

What to do Monitor the planned release of the model weights next week and try it on your own tasks.
More practical updates →
02

For a business

Companies considering deploying AI agents for long-term autonomous tasks (software development, research reproduction, analyses) will have access to an open-weight model that Alibaba claims can work autonomously for days without human intervention; the figures so far come only from internal manufacturer benchmarks.

Development
What to decide Assess the potential of autonomous agents based on the model Qwen3.8-Max for internal development or research tasks once the weights and independent tests are available.
More business impacts →
Alibaba autonomous agents LLM open-weight Qwen3.8-Max research

Check the original

Event sources

only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.

1
The Decoder (daily AI news) independent context · first detected Alibaba’s open-weight Qwen3.8-Max takes on long-horizon AI tasks with 2.4 trillion parameters