A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.
825
published research events
Latest work
A significant claim from a single source is published only after further confirmation.
IEEE Spectrum is promoting a whitepaper presenting survey results from 700+ visual and physical AI specialists, focused on data bottlenecks in production model deployment and the causes of failures. The whitepaper is freely available to download.
An author from AI2 claims that LLMs have improved at short-form writing (copy), but are stagnating in long-form, factual writing. Organizing knowledge is the compression required for insight; however, models increase entropy in long-form writing. This signals that LLMs are not ready for open-ended scientific problems, although the author remains…
Researchers used the Evo 2 model to design genetic sequences for bacteriophages, some of which proved functional in the laboratory. The study raises the question of whether biotechnology safeguards are evolving quickly enough as AI capabilities grow. The model was developed with safety restrictions…
An independent study by economist Sultan Mehmood tested the JudgeGPT AI tool on 1559 Pakistani judges. The tool combines GPT-4 with a database of 130 000 Pakistani legal opinions and laws and uses RAG to prevent hallucinations. Results: a 6.3% increase in resolved cases without a decline in quality…
Google Research introduces a knowledge profiling framework that distinguishes encoding failures from recall failures. An analysis of frontier LLMs on the WikiProfile benchmark (2150 facts) shows that models encode almost all information but struggle to retrieve it. Chain-of-thought and thinking mode provide a solution.
A research team introduced Cross-Contextual Consistency (C3), a method for measuring the credibility of large language models. It compares model responses to the same task with different, semantically neutral contextual variations. A study of 26 models and six benchmarks (reasoning, factuality…
The research paper introduces LLM Agents Factory, a framework that constructs agents from a collection of over 20K profiles through semantic retrieval or distillation into a more compact model. On the MMLU and BIG-bench benchmarks, it achieves higher accuracy than the non-agent baseline and lower costs than AutoGen.…
A research team benchmarked 8 open-source SLMs with four fine-tuning strategies (zero-shot, prefix tuning, LoRA, full fine-tuning) on 2 083 MIMIC-IV-ED cases. Tasks: triage prediction, specialist referral recommendations and diagnosis. LoRA fine-tuned SLMs outperformed Claude Haiku 4.5 and Claude Sonnet 4.5…
The UserToolBench research benchmark tests LLMs' capabilities in personalized decision-making when using tools. The dataset contains 10 user profiles, 36 tool sets and 1 065 interactions. Experiments show that current models struggle to understand users' latent preferences…
The new benchmark VialectBench measures how well LLMs handle Vietnamese dialects. The research team tested 10 models on 2400 dialectal variants. The average performance drop was 2.82%, with some dialects (PNT3, PNT2) causing degradation of up to 6.17%. The answering task saw the largest drop in performance…
A research team published the ProTAGAD model for anomaly detection in graphs with textual content. The model uses separate topological and textual prototypes to address the cross-domain problem and achieves the best results on 14 benchmarks.
The research analyzed the behavior of 32 language models from 6 families using responses to 10 000 prompts and three distance metrics. It found that models cluster by family, distances between families decrease over time and newer reasoning-oriented models have more compact response clouds.
A research team published SpaHybGen, a framework combining neural networks with analytical planners for robotic grasping. The system generalized to 7 different robotic hands without retraining, with a success rate of 94.3–98.0 %. The code and models are open-source.
An unreleased Anthropic model tested 650 ideas with 60 subagents and raised the lower bound for solutions to the Riemann hypothesis. The company's mathematicians verified the progress and formalized it using Lean tools.
Microsoft Research introduced CARE-X, a research model for radiological analysis that combines generative and discriminative capabilities to create reports, assess the presence of findings and locate them on chest X-rays. The model is not approved for clinical use.
Researchers at Stanford University used AI to design the DNA of bacteriophage ΦX174. The synthetic DNA was assembled in the laboratory and produced a functioning virus. The study in Science demonstrates that AI can learn biological systems well enough to design functional organisms. The aim is to use a similar…
Evo AI models designed 285 new bacteriophages, 16 of which produced functional viruses that infect E. coli. This is the first practical demonstration of generative biology in which AI creates new genetic code rather than merely analyzing it.
A team at the University of Florida healthcare facility developed GatorOnco, an agentic LLM trained on 282 billion tokens of biomedical text. In a blinded evaluation by five oncologists, the model outperformed open-source LLMs and achieved performance comparable to human experts in accuracy and safety, outperforming them in…
AI Radar monitors Czech and international sources every day, looking for changes that truly deserve attention.
MonitorsOfficial AI company blogs, specialist media, and research sources.
Selects and combinesFilters out information noise and combines articles about the same change into a single event.
Summarizes and explainsExplains significant events in English: what happened, why it matters and where the information comes from.
The result is a quick overview of what has actually changed in the AI world, rather than another stream of articles.
Use the CS/EN switch to read the same Radar in Czech or English. English content is published after its translation has been checked, so new and older items may appear later.
Everything you need to navigate the AI world
Today’s briefingThe “What is worth attention” selection sits beside Live · AI Flash, followed by research and links to other Radar sections. On mobile, these blocks appear one below another.
AI FlashAn ongoing feed of brief updates with an evidence status. Links lead to a Radar detail page when one is ready, otherwise to the original source. You can also find reset and outage histories here.
Practical applicationsWhat new tools and features can do, what you can try and what their actual impact could be.
Model selectionModel comparison by type of work, capabilities, price and speed.
Research and archiveA separate research overview, topic search and older events by date.
One event, everything that matters
Each row represents one event — not one article. At a glance, you can see its significance, credibility and main point.
Illustrative example, not a current news item.
Importance: ▮▮▮ majorOpenAIModels✓ 6
Agent mode is available to all paying users
Until now, the mode was available only on the highest plan; it is now available on all paid tiers without a waitlist.
▮▮▮ major · ▮▮ important · ▮ we're tracking = how significant the change is✓ 6 = six independent publishers, not the number of articles or feeds✓ official = a clear release, law or incident is substantiated by the relevant authority1 source = no independent confirmation yetbold = who is behind the changegray text = a brief summary of what happened
The detail page contains a fuller summary, its significance and original sources. Practical impact appears in the detail and the For individuals and For businesses views. An AI Flash item reaches the main selection only after it has been expanded and meets the publication rules.
The same news, two practical uses
We first summarize each event in the same way for everyone. Based on those same facts, we then explain what the change means for your own use and what it could mean for how a company operates.
For individualsWhat you can use or try, how the change can help you at work and what to watch out for.
For businessesWhat impact the change could have on processes, costs, risks and other business decisions.
Today’s briefing is the same for everyone. Pages
For individuals and For businesses
can be found in the main navigation — they select only events relevant to the given use case.