A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.
825
published research events
Latest work
A significant claim from a single source is published only after further confirmation.
A new deep learning-based method enables molecular dynamics on longer time scales than existing approaches while maintaining accuracy in predicting physical properties. The study addresses the so-called femtosecond limit in simulations.
SaferAI publishes a study: GLM-5.2 (Z.ai) reaches GPT-5.5 and Claude Opus 4.7 capabilities in cyber and bio tasks, but does not refuse attack testing. Unlike them — Claude Opus 4.7 refuses consistently. The problem: Open-weight models can be run locally without protections.
A third of companies spend 25–40% of their R&D budget on projects with no output, and nearly half of teams estimate a loss of >1 million dollars per project. According to IEEE Spectrum, companies apply AI more to analysis and modeling than to decision-making in early ideation and feasibility, where it would have the greatest effect.
A study by MIT and Stanford tested explainability methods (LLM explanations, heat maps) in skin disease diagnosis. Laypeople improved their accuracy, but mainly through blind trust in AI; healthcare professionals achieved better results without additional explanations. Inexperienced users are most susceptible to…
L'Oréal uses AI to accelerate the development of cosmetic products: instead of physically testing hundreds of molecules, it designs tens of thousands in computer simulations and sends only the most promising ones to the laboratory. It is also switching from animal-derived ingredients to biotechnological alternatives and uses artificial skin models instead of…
The new SCHEMA framework automatically constructs conceptual science graphs, generates test tasks (verification, multi-hop logic, explanation, coding) and detects hallucinations with topological weighting. The research found that hallucinations concentrate at key nodes in the knowledge network and that models achieve…
The study introduces OSCD, an algorithm for improving chain-of-thought reasoning in low-resource Southeast Asian languages. It combines projection of high-resource trajectories with joint-embedding semantic alignment and achieves up to a 3.2× improvement on the AIME25 and HMMT25 mathematics benchmarks…
The research paper introduces PCSD, a method for more efficient training of language model agents on complex tasks. It combines dense teacher supervision with sparse feedback from the environment. On the ALFWorld benchmark, it achieves performance 15.6 points higher than GRPO, without dependence on a specific architecture…
The research paper formalizes measurement of AI consciousness through the Conservation-Congruent Encoding framework. It separates outward behavior from internal structure measured by operational consciousness (κ_T). It shows differences between lookup and generative systems. Relevant to AI safety analysis.
The research team developed the Question-begets-Question (QbQ) method, which generates variants of existing tasks for training. The Qwen2.5-Math-7B model improved its accuracy on competition mathematics (AIME) from 5.6% to 16.5% using reinforcement learning and an iterative curriculum that focuses on problems…
GradCuit is a new method that optimizes latent Transformer states at test time using gradients from the answer reward. It achieves 64.5% accuracy, 6.6 percentage points higher than chain-of-thought prompting. Its main contribution: Direct credit assignment to latent states enables better robustness and…
The research team introduced HopRefusalBench, a benchmark containing 889 unanswerable questions for testing retrieval agents. In tests of 10 proprietary and open-weight models, the best achieved only a 42.9% correct stopping rate; the main weakness is deciding on an appropriate non-answer rather than…
The research team introduced RubricReviewer, a framework for LLM-assisted peer review addressing two key limitations of existing approaches: It explicitly generates rubrics as an intermediate step and combines a training-free evidence-gathering agent (Scout) with a trained review-writing model (Aligner).…
The arXiv research article introduces AdaMTP, a method for adaptive multi-token prediction training in language models. Instead of a fixed horizon length, it adapts to sequence structure through entropy-based segmentation. It detects semantic boundaries and suppresses noisy signals. Tested on Llama-3.1-8B…
The research team released GABench, a benchmark comprising 10 400 tasks for evaluating the capabilities of LLM agents in graph analysis. The benchmark covers three types of graphs and four task categories (retrieval, graph theory, machine learning, open-ended questions) with 84 executable tools. Experiments show that current…
The JONES-19 research compared CNN training on design data (images from The Grammar of Ornament, London 1857). Specialized design data does not require massive pretraining; small, high-quality curated datasets are more effective. Local training with multi-crop augmentation achieved the same performance…
The article explains Kimi Delta Attention (KDA), a linear attention layer in Kimi K3. It traces the technique's evolution from linear attention through DeltaNet and Gated DeltaNet. KDA improves linear attention with a delta rule for updating the matrix, addressing unbounded growth of the hidden state.
SemiAnalysis (newsletter feed — hardware, chips, AI economics)
Original source ↗
Researchers from University of Toronto, Vector Institute, University of Cambridge and ServiceNow created a proof of concept for a virus that runs on compromised GPUs with an open-weight LLM. The virus detects vulnerabilities and creates attacks tailored to individual targets. It operates without a vendor API; the model is from 2025 and…
AI Radar monitors Czech and international sources every day, looking for changes that truly deserve attention.
MonitorsOfficial AI company blogs, specialist media, and research sources.
Selects and combinesFilters out information noise and combines articles about the same change into a single event.
Summarizes and explainsExplains significant events in English: what happened, why it matters and where the information comes from.
The result is a quick overview of what has actually changed in the AI world, rather than another stream of articles.
Use the CS/EN switch to read the same Radar in Czech or English. English content is published after its translation has been checked, so new and older items may appear later.
Everything you need to navigate the AI world
Today’s briefingThe “What is worth attention” selection sits beside Live · AI Flash, followed by research and links to other Radar sections. On mobile, these blocks appear one below another.
AI FlashAn ongoing feed of brief updates with an evidence status. Links lead to a Radar detail page when one is ready, otherwise to the original source. You can also find reset and outage histories here.
Practical applicationsWhat new tools and features can do, what you can try and what their actual impact could be.
Model selectionModel comparison by type of work, capabilities, price and speed.
Research and archiveA separate research overview, topic search and older events by date.
One event, everything that matters
Each row represents one event — not one article. At a glance, you can see its significance, credibility and main point.
Illustrative example, not a current news item.
Importance: ▮▮▮ majorOpenAIModels✓ 6
Agent mode is available to all paying users
Until now, the mode was available only on the highest plan; it is now available on all paid tiers without a waitlist.
▮▮▮ major · ▮▮ important · ▮ we're tracking = how significant the change is✓ 6 = six independent publishers, not the number of articles or feeds✓ official = a clear release, law or incident is substantiated by the relevant authority1 source = no independent confirmation yetbold = who is behind the changegray text = a brief summary of what happened
The detail page contains a fuller summary, its significance and original sources. Practical impact appears in the detail and the For individuals and For businesses views. An AI Flash item reaches the main selection only after it has been expanded and meets the publication rules.
The same news, two practical uses
We first summarize each event in the same way for everyone. Based on those same facts, we then explain what the change means for your own use and what it could mean for how a company operates.
For individualsWhat you can use or try, how the change can help you at work and what to watch out for.
For businessesWhat impact the change could have on processes, costs, risks and other business decisions.
Today’s briefing is the same for everyone. Pages
For individuals and For businesses
can be found in the main navigation — they select only events relevant to the given use case.