A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.
825
published research events
Latest work
A significant claim from a single source is published only after further confirmation.
The research team introduced EnvACE, a method replacing costly interactions with external environments during training with ‘world rehearsal’: The agent alternates between generating actions and playing the environment's role. It achieves strong performance on BFCL-v4, tau²-Bench, VitaBench and FinMCP-Bench. The source code is…
A study of 105 AI agent configurations for microscope control showed that benchmarks were suitable for familiar tasks, but models trained on them did not predict performance on new ones. It covered five LLMs, different RAG parameters, 1949 tests and 49109 RAG retrievals.
The research survey introduces AISAC, a paradigm for AI agents in sensing and communication. The framework contains six phases (observation, contextualization, reasoning, planning, execution, feedback) and five maturity levels. An audit of systems found that none reports more than 1–2 of the nine agentic capability criteria.
The SkillTrace research framework audits skill reuse in LLM agent ecosystems. It extracts three provenance traces (Expression, Implementation, Operational) and represents the flow as a Skill Operational Graph. It achieves AUROC 0.938 and F1 0.898 on 820 transformed positives. An audit of 36 446…
Scientists introduced Implicit Social Context Analysis (MoCA), a task with a dataset benchmark containing 3 108 multimodal instances for studying affect, intent and stance in human communication. Tests showed that current multimodal language models struggle with this task; a framework was proposed…
An arXiv paper on the Unified Agent system for AI agents working across multiple devices over time. It proposes efficient state management for cross-device interactions. Measurements show that it significantly outperforms four comparison approaches and remains robust across different multimodal LLM models. Code and data will be on GitHub.
The study examines how biased conversational interaction in a multi-turn setting affects the expression of cognitive biases in LLMs. It tested eight models on 24 300 jury-validated prompts covering 81 cells of a 9×9 matrix. It found that biased conversation increases the expression of biases in six out of eight…
The research tests five LLMs on localization of numbers, times and dates. Including localization principles in the prompt produced a statistically significant improvement in accuracy compared with direct translation and alternative strategies.
The study proved that marginal matching (a common regularization approach) in factorized generative models does not prevent the latent style variable from containing class information. Although the model achieves near-zero global MMD, a linear probe still recovers the class with 74–100% accuracy (compared with 10%…
The study compares diffusion language models (DLM) and autoregressive models (ARM). DLM achieve higher arithmetic intensity through parallelism but do not scale efficiently to long contexts. Blockwise decoding improves scaling. ARM show better throughput in batch inference. The key to lower DLM latency is reducing sampling…
Researchers from UC Berkeley and Lawrence Berkeley National Lab proposed ARBITRAGE, a framework for step-level speculative decoding. Instead of a fixed acceptance threshold, it uses a router trained to predict when the target model produces a better step. On mathematical benchmarks, they achieved inference speedups of up to 2× with…
An Australian study shows that administrative occupations (70% women) are most exposed to AI automation. Of the 20 most at-risk occupations, 15 are female-dominated; among the least at-risk, 17 are male-dominated. Vacancies in exposed occupations are falling by up to 22% annually.
Stanford University scientists used large-scale genomic models to design new viruses infecting bacteria. All created viruses are closely related to existing ones and have some different properties. The researchers highlight the possibility that similar AI could in future design viruses targeting…
Google DeepMind's WeatherNext improves hurricane forecast accuracy by an average of one day. Three-day forecasts are as accurate as previous two-day forecasts. Trained on a combination of general meteorological and hurricane-specific data, the model handles both track and intensity prediction.
British organization Which? tested AI chatbots on financial advice. The tools correctly explained general concepts but made errors in data — ChatGPT cited nonexistent products, while Copilot used incorrect rates. AI drew from social networks and did not provide current regional information.
The research team proposed MCTS-Report, a framework combining Monte Carlo Tree Search with language models to automatically create professional reports from structured data. It divides generation into atomic actions (chapter planning, chart creation, formulating insights), which it optimizes according to…
The study finds that multimodal language models perform an average of 17.8 points worse on tasks when the question is presented as text in an image rather than plain text. Although models transcribe the text correctly, they cannot use it as an instruction for reasoning. The authors propose prompt-region grounding, which…
The research paper introduces the OCSD method for training language models for agents. It addresses the problem of confounding in supervision by using two structurally identical replay views — with and without the future observation. Tested on ALFWorld, WebShop and Search-QA, the method consistently outperformed strong baseline models.
AI Radar monitors Czech and international sources every day, looking for changes that truly deserve attention.
MonitorsOfficial AI company blogs, specialist media, and research sources.
Selects and combinesFilters out information noise and combines articles about the same change into a single event.
Summarizes and explainsExplains significant events in English: what happened, why it matters and where the information comes from.
The result is a quick overview of what has actually changed in the AI world, rather than another stream of articles.
Use the CS/EN switch to read the same Radar in Czech or English. English content is published after its translation has been checked, so new and older items may appear later.
Everything you need to navigate the AI world
Today’s briefingThe “What is worth attention” selection sits beside Live · AI Flash, followed by research and links to other Radar sections. On mobile, these blocks appear one below another.
AI FlashAn ongoing feed of brief updates with an evidence status. Links lead to a Radar detail page when one is ready, otherwise to the original source. You can also find reset and outage histories here.
Practical applicationsWhat new tools and features can do, what you can try and what their actual impact could be.
Model selectionModel comparison by type of work, capabilities, price and speed.
Research and archiveA separate research overview, topic search and older events by date.
One event, everything that matters
Each row represents one event — not one article. At a glance, you can see its significance, credibility and main point.
Illustrative example, not a current news item.
Importance: ▮▮▮ majorOpenAIModels✓ 6
Agent mode is available to all paying users
Until now, the mode was available only on the highest plan; it is now available on all paid tiers without a waitlist.
▮▮▮ major · ▮▮ important · ▮ we're tracking = how significant the change is✓ 6 = six independent publishers, not the number of articles or feeds✓ official = a clear release, law or incident is substantiated by the relevant authority1 source = no independent confirmation yetbold = who is behind the changegray text = a brief summary of what happened
The detail page contains a fuller summary, its significance and original sources. Practical impact appears in the detail and the For individuals and For businesses views. An AI Flash item reaches the main selection only after it has been expanded and meets the publication rules.
The same news, two practical uses
We first summarize each event in the same way for everyone. Based on those same facts, we then explain what the change means for your own use and what it could mean for how a company operates.
For individualsWhat you can use or try, how the change can help you at work and what to watch out for.
For businessesWhat impact the change could have on processes, costs, risks and other business decisions.
Today’s briefing is the same for everyone. Pages
For individuals and For businesses
can be found in the main navigation — they select only events relevant to the given use case.