A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.
169
published research events4
new research papers today
Latest work
A significant claim from a single source is published only after further confirmation.
The study presents Sieve and Sage: a lightweight module that filters out distracting source material before generating a response. Compared to single-stage approaches, the authors report an improvement in accuracy of up to 69.4 percentage points, Macro-F1 of 55.2 percentage points, and up to a 1.99x speedup.
The study presents FlexRouter, which selects LLM models based on their capabilities and mutual complementarity. It determines the number of models dynamically. According to the authors, in RouterEval tests it achieves a higher probability of at least one correct answer and lower redundancy than the compared methods.
The study presents PrimeSeeker for training search agents. The authors created 9 221 expert trajectories and trained a model with 30 billion parameters. In five test sets they report good results and fewer tool calls; the abstract does not contain exact comparison numbers.
The study presents the ArgGYM benchmark with 12 tasks and 1 440 verified examples. A symbolic system verifies the answers. The tested models handle parts of the solution, but their performance declines with longer dependencies and more complex structures.
A study used SFT and PPO to fine-tune the Qwen3-4B model into AURA-4B. The score for the direction and magnitude of stock returns rose from 20.94 to 43.31, and directional accuracy rose from 62.9 to 65.4. The results apply to the test environment examined.
The AIMS framework uses two agents to coordinate the creation of simulated data and model training for multi-modal ISAC. The authors report better vehicle detection and beam prediction on the DeepSense 6G dataset compared to selected baseline methods; the abstract does not state the magnitude of the improvement.
The study presents TRACE, which supplements an oncology LLM with structured evidence when generating responses. The authors report improvements on ten classification tasks and one question-answering benchmark, including comparisons with RAG and GraphRAG.
The study describes the use of historical data, machine learning, and deep learning to predict transit time in the supply chain. The abstract does not specify a concrete model, data size, or measured accuracy.
According to the study, given the same time budget and the same base model, advanced agentic architectures did not provide an advantage over a single session of a simple coding agent. The results concern current tests of autonomous machine learning tasks.
The study presents the LACE framework for automated heuristic generation. On 36 CO-Bench tasks, it achieved an average score of 0.945, compared to 0.870 for the best benchmarked LLM-based method. The data and resulting heuristics are publicly available under the MIT license.
Over six weeks, researchers reviewed 400 URLs and selected 88 websites with non-consensual intimate content. According to a peer-reviewed study, their infrastructure is provided mainly by the companies Cloudflare, Google, Namecheap, and Proton, and the product WordPress.
The author of the article describes a machine learning system for estimating geomagnetic risk for 66 935 substations in the US with a lead time of 30–60 minutes. It combines solar wind observations, geological data, and grid data. The available portion does not state numerical accuracy results.
Researchers combined AI, acoustic and visual techniques in studying two Etruscan tombs from the 5th century BC. The results suggest that sound and lighting enhanced the impact of the paintings on visitors. The specific AI method is missing from the available excerpt.
Researchers introduced an AI system for games with hidden information. According to the article, it beat top-level Stratego players, outperformed existing models, and required fewer computational resources during training. The study is published in the journal Nature.
The study covered 23 models, five tasks, and more than 5 500 runs. At the same generation budget, agent debate matched or worsened results compared to self-consistency sampling, while taking 1.6x longer and consuming 3.4x more tokens.
The study presents the KUPAS MASTER platform for converting professional experience into AI agent skills. It processed 1 576 files from 20 experts. In the evaluation, the agent achieved a score of 89.58, compared to 79.75 when using RAG over the original source materials and 70.63 for the base model.
The study presents the DASA method, which uses synthetic input representations for LLM adaptation. According to the authors, it achieves comparable or better results than original text data on six models with 1–32 billion parameters. Synthesis is 3.6–4.9× faster than the GRADMM method.
The study compared standard and adversarial training of the GPT-2 Small model. The robust variant had simpler representations, but needed smaller causal circuits only at 90% and 95% fidelity. Below 85%, the standard variant performed better or the results were equal.
AI Radar monitors Czech and international sources every day, looking for changes that truly deserve attention.
MonitorsOfficial AI company blogs, specialist media, and research sources.
Selects and combinesFilters out information noise and combines articles about the same change into a single event.
Summarizes and explainsExplains significant events in English: what happened, why it matters and where the information comes from.
The result is a quick overview of what has actually changed in the AI world, rather than another stream of articles.
Everything you need to navigate the AI world
Today’s briefingSignificant events, verified and explained — without duplicates or repeated reports of the same news.
AI FlashAn ongoing feed of brief updates with a direct link to the original and a visible evidence status.
Practical applicationsWhat new tools and features can do, what you can try and what their actual impact could be.
AI model selectionModel comparison by type of work, capabilities, price and speed.
Research and archiveA separate research overview, topic search and older events by date.
One event, everything that matters
Each row represents one event — not one article. At a glance, you can see its significance, credibility and main point.
▮▮▮ majorOpenAIModels✓ 6
Agent mode is available to all paying users
Until now, the mode was available only on the highest plan; it is now available on all paid tiers without a waitlist.
▮▮▮ major · ▮▮ important · ▮ we're tracking = how significant the change is✓ 6 = how many sources confirm itbold = who is behind the changegray text = a brief summary of what happened
When you open an event, you will find a more detailed summary, an explanation of its significance, and links to the original sources. Quick updates are available separately in AI Flash; important topics move to Today’s Overview once more information is added.
The same news, two practical uses
We first summarize each event in the same way for everyone. Based on those same facts, we then explain what the change means for your own use and what it could mean for how a company operates.
For individualsWhat you can use or try, how the change can help you at work and what to watch out for.
For businessesWhat impact the change could have on processes, costs, risks and other business decisions.
Today’s briefing is the same for everyone. Pages
For individuals and For businesses
can be found in the main navigation — they select only events relevant to the given use case.