A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.
257
published research events
Latest work
A significant claim from a single source is published only after further confirmation.
Research from Google Cloud AI Research addresses the problem where AI agents that optimize their own harness overfit too much to tests. The new RRSI method regulates optimization with a budget of edit changes and a filter against task-specific tricks. Result: generalization is preserved without increasing compute costs.
NASA and IBM released the Lunar Foundation Model, an open-source foundation model for lunar science. Trained on 2 million tile bundles from 17 years of data from the Lunar Reconnaissance Orbiter and other missions. The model specializes in ice and crater detection.
Part of an overview of multiple AI topics; only this event has been covered.
Aleph Alpha created a benchmark covering 967 politically sensitive topics. Tests of the Qwen, DeepSeek, and Kimi models showed that only 17–41% of responses were balanced; the rest repeat state doctrine, evade the question, or refuse to answer. The bias also shows up in unrelated questions. Nvidia Nemotron contains approx…
The ThinkingBox benchmark evaluates AI agents across 507 business processes based on the final state of the systems. According to the authors, 79 853 out of 121 680 attempts with 12 models failed; 67.24% of the failed attempts nevertheless ended without a final tool error and still performed a state change.
In an essay, Benjamin Bratton, Blaise Agüera y Arcas and James Manyika propose Artificial Symbiotic Intelligence: an interconnected system of humans and AI agents. According to the authors, AGI may emerge through their collaboration; this is a conceptual argument, not an announcement of a new model.
Researchers from University of Maryland and AWS introduced LEGO-Anything: an agent progressively creates and edits Blender code based on a photograph. In LEGO-Bench with 208 images, the GPT-6 Astra model achieved accuracy of 53.4 % indoors and 39.6 % outdoors, according to the study.
According to a New York Times report, Anthropic has since autumn 2025 invited dozens of theologians and philosophers to discussions about the possible consciousness and moral behavior of Claude models. The report is based on interviews with 20 participants; this does not prove consciousness of the models.
A profile of MIT researcher Cathy Wu describes her focus on machine learning and reinforcement learning for transportation systems. According to the researcher, these methods could make it easier to compare many design variants; the available text does not provide specific results.
The study's authors interviewed 45 knowledge workers at a US public university and analyzed their conversations with AI. According to the authors, deliberately seeking out conflicting viewpoints using AI can improve problem-solving; the outputs must be assessed using one's own expertise.
According to a study by Mercor, AI models outperform 12 licensed accountants on simplified tasks. In the full test with 160 tasks, the Claude Opus 5.5 model leads with 61.8% of criteria met. Nearly 60% of tasks were not fully solved by any model; closing still requires supervision.
The author of the commentary explains that the AlphaGo system combined neural networks with search through possible moves. He claims that current LLMs lack similar reasoning abilities and that these are needed for reliable results in science and medicine. The text is incomplete.
The ServiceNow CoreAI team describes AutoSynthData, which generates and verifies training tasks for enterprise agents based on the failures of the target model and the successes of a stronger model. The available excerpt does not provide measurement results or a tool release.
The study presents Mem++, which stores entire documents with date and author without a generative model at write time. When queried, it selects time-matching documents. According to the authors, it outperformed the strongest comparison memory system on OrgMemBench by 8.0 to 13.1 points.
The SELF-POT study compares five low-cost models across 350 tasks and counts all calls toward the cost. In repeated evaluation of coding candidates, selection based on public examples increased the number of correct solutions from 376 to 453 out of 500 and reduced API costs by 12–49%.
The study states that none of the 16 Qwen3 model configurations met the predetermined thresholds for four AI agent subtasks. Even 4-bit quantization failed to produce a satisfactory configuration. The authors recommend combining small models with a proven baseline procedure.
The study proposes CTWM, which allocates the context budget according to memory usage frequency. The authors report a token savings of 24.48% on LongMemEval with comparable overall accuracy, and on the Synthetic Graph World a 13.6% decrease in prediction error for the less frequently used half of memory.
The study presents the DSB-DG benchmark with 1 636 question-answer pairs from 50 documents across five professional domains. According to the authors, the accuracy of voice agents decreases with context and dialogue length; the ASR-LLM-TTS architecture performs best.
The authors tested an agent using the Qwen3.6-35B-A3B model with a library of geometric trajectories in a dynamic 2D environment. According to the study, the configuration without reasoning achieved the same goal-reaching success rate as the chain-of-thought variant, while cutting decision time from minutes to seconds.
AI Radar monitors Czech and international sources every day, looking for changes that truly deserve attention.
MonitorsOfficial AI company blogs, specialist media, and research sources.
Selects and combinesFilters out information noise and combines articles about the same change into a single event.
Summarizes and explainsExplains significant events in English: what happened, why it matters and where the information comes from.
The result is a quick overview of what has actually changed in the AI world, rather than another stream of articles.
Use the CS/EN switch to read the same Radar in Czech or English. English content is published after its translation has been checked, so new and older items may appear later.
Everything you need to navigate the AI world
Today’s briefingThe “What is worth attention” selection sits beside Live · AI Flash, followed by research and links to other Radar sections. On mobile, these blocks appear one below another.
AI FlashAn ongoing feed of brief updates with an evidence status. Links lead to a Radar detail page when one is ready, otherwise to the original source. You can also find reset and outage histories here.
Practical applicationsWhat new tools and features can do, what you can try and what their actual impact could be.
Model selectionModel comparison by type of work, capabilities, price and speed.
Research and archiveA separate research overview, topic search and older events by date.
One event, everything that matters
Each row represents one event — not one article. At a glance, you can see its significance, credibility and main point.
Illustrative example, not a current news item.
Importance: ▮▮▮ majorOpenAIModels✓ 6
Agent mode is available to all paying users
Until now, the mode was available only on the highest plan; it is now available on all paid tiers without a waitlist.
▮▮▮ major · ▮▮ important · ▮ we're tracking = how significant the change is✓ 6 = six independent publishers, not the number of articles or feeds✓ official = a clear release, law or incident is substantiated by the relevant authority1 source = no independent confirmation yetbold = who is behind the changegray text = a brief summary of what happened
The detail page contains a fuller summary, its significance and original sources. Practical impact appears in the detail and the For individuals and For businesses views. An AI Flash item reaches the main selection only after it has been expanded and meets the publication rules.
The same news, two practical uses
We first summarize each event in the same way for everyone. Based on those same facts, we then explain what the change means for your own use and what it could mean for how a company operates.
For individualsWhat you can use or try, how the change can help you at work and what to watch out for.
For businessesWhat impact the change could have on processes, costs, risks and other business decisions.
Today’s briefing is the same for everyone. Pages
For individuals and For businesses
can be found in the main navigation — they select only events relevant to the given use case.