A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.
825
published research events
Latest work
A significant claim from a single source is published only after further confirmation.
The research paper proposes PL-Guard, a neurosymbolic architecture that separates semantic grounding from policy reasoning in guardrails. On the XSTest benchmark, it reduces unsafe responses from 22% to 0.5% but increases refusals of benign queries to 14.4%.
The study introduces CEDAR-GRPO, combining final correctness with rewards for evidence coverage and logic. Four open-weight LLMs are trained to generate and select hypotheses. Improvements average +7.4 points over the baseline, with a maximum of +30.8 points on 11 unseen tasks.
The study identifies rubric interference: when an LLM assesses multiple criteria simultaneously, its verdicts change depending on the combination of criteria (only 1/3 of samples have consistent verdicts). The authors propose SARA, which uses single-rubric assessments as anchors and on-policy self-distillation. Validated on…
Researchers proposed a formal framework (policy algebra) for safe execution of AI agents in enterprises. It defines agent reliability as the ability to achieve a goal without violating constraints on data access, authorization, side effects and budgets. Evaluation: it stops 94.8% of policy violations with 86.9%…
The research paper proposes UniFed-VLM, a method for federated instruction tuning of vision-language models. It addresses heterogeneity in client tasks, modalities and architectures through FedCSA for adapter aggregation and TCoD for knowledge transfer. Code is available on GitHub.
A research team presented the Unwritten benchmark on arXiv to test the ability of multimodal models to recognize words from acoustic signals produced by a pencil and video of its movements without visible ink. While humans achieve accuracy of >80 %, the GPT-4o and Gemini 2.5-Pro models fail with accuracy below 10 %. Paradoxically, the combination…
The new RHMP architectural approach simulates physical fields and separates topological aspects (conservation laws) from geometry learned from data. Tested on 7 physical domains, including fluids, electromagnetism, and CFD. It achieves the best performance particularly in the interplay between topology, geometry, and field structure.
A replication study verified the α-FLOPs formula for predicting AI operation execution time. It confirmed that raw FLOPs are not a sufficient metric but found that the formula falls short on newer hardware — with oscillations and abrupt changes in execution time that it predicts inadequately.…
Scientists created AI-designed bacteriophages capable of killing antibiotic-resistant strains of E. coli. The article points out that biotechnology regulation is not keeping pace with the convergence of AI, CRISPR and synthetic biology, and laws designed for individual technologies cannot address their combinations.
Scientists from Apple ML Research, MIT and UC Berkeley proposed a new semismooth Newton method for solving kernel-based optimal transport problems. It outperforms the computationally expensive SSIPM approach and achieves O(1/√k) global and quadratic local convergence. Experiments show significant speedups…
The Qwen 3.8 27B model scored 52 in the Artificial Analysis Intelligence Index, the same as GPT-5.6 Luna, one point less than GLM-5.2 (753 billion parameters) and DeepSeek V4 Pro 0813 (1.6 billion parameters).
Simon Willison — AI tag (leading independent LLM commentator)
Original source ↗
DiG-bench is a new benchmark of 70 text-based games designed to measure the ability of AI models to discover hidden rules in an environment. A research team from Oxford, Princeton, MIT and other institutions released 21 games publicly, while the remaining 49 are hidden. The games remain unsolvable by current frontier models.
Axiom Math used its AxiomProver AI system to verify the proof of the 246 theorem on prime numbers — a fundamental result in prime number theory. The verification demonstrates the potential of automated proof checking; the system created a library of components for reuse in further research.
Google Research introduces PhotoScan, a neural network that estimates body composition (body fat, A/G and V/S ratios) from 2D smartphone photos. Trained on 35 000+ UK Biobank records and validated in clinical cohorts. Accuracy comparable to DXA scans in predicting insulin resistance…
The research paper proposes a method for evaluating continuous learning in agent systems without labeled benchmarks. A stronger teacher model provides sparse, correct corrections to a student with a learning harness; improvement relative to the teacher correlates with actual improvement. The method is validated on cybersecurity tasks and more…
The research article identifies 11 specific mechanisms that generate and amplify misunderstanding in AI-mediated communication. It consolidates findings from nine disciplines, models 8 analytical layers and provides an evidence matrix with 9 analyzed dialogue cases. It maps where in the communication process…
The research survey integrates federated learning with large language models for decentralized training without centralizing data. It analyzes communication efficiency, security and privacy in pre-training, fine-tuning and practical applications. It identifies remaining security and robustness challenges.
A research team presents CLAIR-Fin, a nine-agent framework for mitigating hallucinations in AI responses to financial questions. On a dataset of 500 questions, it increases credibility from 0.78 to 0.89, with an abstention rate of 5.4%. It outperforms the HyDE and Graph-RAG baselines.
AI Radar monitors Czech and international sources every day, looking for changes that truly deserve attention.
MonitorsOfficial AI company blogs, specialist media, and research sources.
Selects and combinesFilters out information noise and combines articles about the same change into a single event.
Summarizes and explainsExplains significant events in English: what happened, why it matters and where the information comes from.
The result is a quick overview of what has actually changed in the AI world, rather than another stream of articles.
Use the CS/EN switch to read the same Radar in Czech or English. English content is published after its translation has been checked, so new and older items may appear later.
Everything you need to navigate the AI world
Today’s briefingThe “What is worth attention” selection sits beside Live · AI Flash, followed by research and links to other Radar sections. On mobile, these blocks appear one below another.
AI FlashAn ongoing feed of brief updates with an evidence status. Links lead to a Radar detail page when one is ready, otherwise to the original source. You can also find reset and outage histories here.
Practical applicationsWhat new tools and features can do, what you can try and what their actual impact could be.
Model selectionModel comparison by type of work, capabilities, price and speed.
Research and archiveA separate research overview, topic search and older events by date.
One event, everything that matters
Each row represents one event — not one article. At a glance, you can see its significance, credibility and main point.
Illustrative example, not a current news item.
Importance: ▮▮▮ majorOpenAIModels✓ 6
Agent mode is available to all paying users
Until now, the mode was available only on the highest plan; it is now available on all paid tiers without a waitlist.
▮▮▮ major · ▮▮ important · ▮ we're tracking = how significant the change is✓ 6 = six independent publishers, not the number of articles or feeds✓ official = a clear release, law or incident is substantiated by the relevant authority1 source = no independent confirmation yetbold = who is behind the changegray text = a brief summary of what happened
The detail page contains a fuller summary, its significance and original sources. Practical impact appears in the detail and the For individuals and For businesses views. An AI Flash item reaches the main selection only after it has been expanded and meets the publication rules.
The same news, two practical uses
We first summarize each event in the same way for everyone. Based on those same facts, we then explain what the change means for your own use and what it could mean for how a company operates.
For individualsWhat you can use or try, how the change can help you at work and what to watch out for.
For businessesWhat impact the change could have on processes, costs, risks and other business decisions.
Today’s briefing is the same for everyone. Pages
For individuals and For businesses
can be found in the main navigation — they select only events relevant to the given use case.