A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.
825
published research events
Latest work
A significant claim from a single source is published only after further confirmation.
A research project presents a unified formulation of LLM routing (5 components: context encoders, model encoders, scoring functions, decision rules, learning signals) and the xRouteBench benchmark for various tasks. The open-source LLMRouter infrastructure includes 16+ routers. Experimental results show…
The research introduces FutureBridge, a technique for improving collaboration between large and small language models. Instead of relying on the local preferences of the large model (LLM), it uses supervised token reranking based on how well the tokens enable the small model (SLM) to continue. On five benchmarks of mathematical…
The paper proposes an agentic approach to generating knowledge graphs for HR platforms. It combines LLMs with Wikidata and processes multilingual skills (5 European languages) in five stages: reconciliation, canonicalization, curation, deduplication and recovery. The system maps unstructured text to structured…
The research paper presents a benchmark and framework for analyzing LLM agent trajectories, focused on attributing responsibility for individual steps. It includes 1300+ annotated trajectories from AgentDojo and Agent3Sigma, defines two evaluation tasks and provides baseline results. It releases reusable annotation…
The study compares agent behavior in the Artificial Life model under noisy perception. Agents that are aware of uncertainty have a significantly higher survival rate and fewer critical errors. The results show a qualitative shift from exploratory to cautious behavior as uncertainty increases; explicit collection…
A research team proposes a diagnostic method that separates different sources of failure in multimodal models for speech and emotion analysis. The method compares the emitted answer, option logits and different ways of reading out the hidden state. Across five systems and two emotion datasets, an average improvement was found…
The new Science Edge Evaluation benchmark tested 19 multimodal LLMs on scientific tasks in chemistry, biology and materials science. The best model achieved 48.7% accuracy, and general-purpose models outperformed specialized ones. Adding tools increased accuracy to 52.7%. The study found that models cannot scientifically…
The paper finds that RAG systems ignore user instructions for communication tone (e.g. formal, friendly, simple) because the style of retrieved documents dominates generation. It proposes the TA-RAG framework with four constraints: dishonorable language, readability, adaptation to the recipient, and empathetic framing.…
Researchers introduced the NS-RIS (Newton-Schulz Retraction-based Inference) algorithm for more efficient training of Hidden Quantum Markov Models (HQMMs). NS-RIS demonstrates in practice for the first time that HQMM can outperform classical EM-trained HMM even on data not generated by quantum processes. On synthetic HMM…
OpenAI announced ten mathematical discoveries made with the unreleased Astra model (geometry, cryptography, coding theory). The mathematics community is discussing questions of attribution and AI's impact on the discipline's future. The Leiden Declaration advocates AI as support for human creativity rather than a replacement; leading mathematicians…
A study with 2500+ participants showed that readers cannot distinguish ChatGPT 4.0-generated short stories from human-written ones better than chance and rate AI texts more highly (quality: 1.54 vs 0.97; immersion: 1.42 vs 1.00). However, ratings fall once readers learn that the texts were generated by AI.
WeatherNext, developed by Google DeepMind and Google Research, can forecast hurricanes more accurately than existing models. According to research published in Nature, it provides an average of one additional day of warning compared with existing models — its three-day forecasts are as accurate as…
Climate scientist Zeke Hausfather analyzed eight weeks of Claude Code usage: 1 138 prompts triggered over 14 000 model calls and processed 3.2 billion tokens, consuming 170 kWh (150 Wh per prompt). This is roughly 600 times more than Google and OpenAI figures suggest.
Hugging Face introduced TutorMoments, an evaluation framework based on real tutoring sessions that measures LLM models' ability to choose between support and independence. The result: Models provide too much help; an explicit instruction improves performance, but a gap remains compared with human tutors. It releases the dataset, code and…
Scientists from Stanford University and Arc Institute used Evo 1 and Evo 2 AI models to design 16 entirely new, functional bacteriophages based on genomes from millions of organisms. Of 300 synthesized candidates, 16 proved fully functional.
The research team analyzed over 1600 comments under BBC News videos on YouTube and found that a more effective way to detect AI disinformation is to recognize red herring tactics and diversion of conversations toward polarizing topics rather than traditional AI-written text detection, which generative AI…
A survey of 647 employers showed that communication, teamwork and critical thinking are valued more highly in graduates than AI knowledge. Automation eliminates traditional entry-level roles such as apprenticeships; higher education still prepares students for the old model.
The new PoolBench benchmark evaluated 19 pooling strategies on 3 open language models (Llama, Gemma, Mistral) using 37 693 texts. The W4_hierarchical strategy achieved an AUROC of 0.7799, significantly better than the standard P1_last_token method (0.7640, p=2.0e-36). The research revealed that the choice of construction method has…
AI Radar monitors Czech and international sources every day, looking for changes that truly deserve attention.
MonitorsOfficial AI company blogs, specialist media, and research sources.
Selects and combinesFilters out information noise and combines articles about the same change into a single event.
Summarizes and explainsExplains significant events in English: what happened, why it matters and where the information comes from.
The result is a quick overview of what has actually changed in the AI world, rather than another stream of articles.
Use the CS/EN switch to read the same Radar in Czech or English. English content is published after its translation has been checked, so new and older items may appear later.
Everything you need to navigate the AI world
Today’s briefingThe “What is worth attention” selection sits beside Live · AI Flash, followed by research and links to other Radar sections. On mobile, these blocks appear one below another.
AI FlashAn ongoing feed of brief updates with an evidence status. Links lead to a Radar detail page when one is ready, otherwise to the original source. You can also find reset and outage histories here.
Practical applicationsWhat new tools and features can do, what you can try and what their actual impact could be.
Model selectionModel comparison by type of work, capabilities, price and speed.
Research and archiveA separate research overview, topic search and older events by date.
One event, everything that matters
Each row represents one event — not one article. At a glance, you can see its significance, credibility and main point.
Illustrative example, not a current news item.
Importance: ▮▮▮ majorOpenAIModels✓ 6
Agent mode is available to all paying users
Until now, the mode was available only on the highest plan; it is now available on all paid tiers without a waitlist.
▮▮▮ major · ▮▮ important · ▮ we're tracking = how significant the change is✓ 6 = six independent publishers, not the number of articles or feeds✓ official = a clear release, law or incident is substantiated by the relevant authority1 source = no independent confirmation yetbold = who is behind the changegray text = a brief summary of what happened
The detail page contains a fuller summary, its significance and original sources. Practical impact appears in the detail and the For individuals and For businesses views. An AI Flash item reaches the main selection only after it has been expanded and meets the publication rules.
The same news, two practical uses
We first summarize each event in the same way for everyone. Based on those same facts, we then explain what the change means for your own use and what it could mean for how a company operates.
For individualsWhat you can use or try, how the change can help you at work and what to watch out for.
For businessesWhat impact the change could have on processes, costs, risks and other business decisions.
Today’s briefing is the same for everyone. Pages
For individuals and For businesses
can be found in the main navigation — they select only events relevant to the given use case.