Startup Mostik connected GLM-5.2 and Qwen-3.5 through communication without text output
Startup Mostik has developed a method through which AI models communicate directly using mathematical values in their weights, without text output. Combining GLM-5.2 (753 billion parameters) with Qwen-3.5 (4 billion) produced a hybrid system at one fifth... one twentieth of the price of the full GLM, with performance exactly between…
Russian startup Mostik has developed a method that allows two AI models to pass information directly to each other using mathematical values contained in their weights, without having to generate text output. According to the company, this means in practice that the capabilities of a larger model can be transferred more efficiently to a smaller model. The team demonstrated the technique by connecting the largest version of GLM-5.2 (753 billion parameters) with Qwen-3.5, a model with 4 billion parameters that can run on a mobile device. According to the company, the resulting hybrid system costs one twentieth of the price of the full GLM model, and its performance falls exactly halfway between the two connected models.
According to CEO Sasha Malysheva, who developed the method, it is well known in machine learning that ensembles of multiple models achieve better results than individual models. The usual way to combine models involves sending the output of one model to another as text, which is costly in both time and money. The approach from Mostik is intended to bypass this by having the models communicate without generating text. The company used the same technique to build a model that, according to the article, took the lead in the ARC-AGI 3 competition for AI models, but the company did not disclose details of this solution for competitive reasons.
External commentators quoted in the source describe the approach as promising: Vladimir Arustamian, technical lead at Lovable, says the team has been working on the solution for only a few months and already has a functioning system that he would have expected to take years to develop. Karl Tuyls, a former scientist at Google DeepMind, describes the method as a way to approach the quality of a large model without having to run it throughout the processing loop. Stanislav Smirnov, a recipient of the Fields Medal in 2010 and chief scientist at Mostik, points out that finding a common mathematical language between two models is surprisingly difficult and that no suitable mathematical framework for this exists yet.
Why it matters
If the technique from Mostik became widespread, it would allow large and small models to be combined more cheaply instead of training or running a single monolithic model — according to the company, it is an alternative to the strategy of scaling up models. According to the description in the source, this could give open models an advantage because they could be combined more easily with small specialized models (e.g. for biology or physics) and compete more effectively with closed offerings from Anthropic and OpenAI. For now, however, it is a research demonstration with no released code, product or price list for public use.
Relevant practical impact
What this means
For a business
If the technique proves successful, it could reduce the long-term costs of running AI systems for companies by combining large and small models, and make open models more competitive with closed offerings from Anthropic and OpenAI. For now, however, it is a research demonstration with no released code or product.
DevelopmentCheck the original
Event sources
only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.