A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.
224
published research events4
new research papers today
Latest work
A significant claim from a single source is published only after further confirmation.
The study compared standard and adversarial training of the GPT-2 Small model. The robust variant had simpler representations, but needed smaller causal circuits only at 90% and 95% fidelity. Below 85%, the standard variant performed better or the results were equal.
The SCLATE study compares 10 agent configurations across 10 models and 7 task sets. Added memory did not reliably improve performance. According to the authors, fine-tuning the Qwen3.5-4B model increased the success rate on the SWE-bench Verified set by 16.7 percentage points.
The study found that activation steering methods often degrade text fluency and are less effective on instruction-tuned models. Simple text metrics correlated strongly with costly LLM-as-judge evaluation.
Research published in Nature Machine Intelligence describes a sparse Vision Transformer framework for analyzing signals from high-energy neutrino detectors. It combines self-supervised pretraining (masked-autoencoder and relational voxel-level objects) with fine-tuning. On the FASERCAL detector, pretraining improved detection…
Google Research has introduced Diffusion Controller, a lightweight add-on network for controlling image generation in diffusion models. It achieves a 90% win rate against the baseline model and can also be attached to closed models without access to their weights. It unifies existing inference-time and fine-tuning approaches into…
Anthropic deployed 950 Claude agents to search genomic databases and identified a new ART system with CRISPR-like repeating sequences. The discovery was made in 21.5 hours, but it is as yet unpeer-reviewed work, and its functional usability for genetic engineering is unclear.
Microsoft Research presents Quine, an AI system for biology combining foundation models with scientific tools and data. Designed to overcome the experimental limitations of biological research by connecting insights across modalities and scales.
Research shows that fine-tuning vision-language models on narrow tasks induces emergent misalignment – undesirable behavior in unrelated tasks. Across 15 models, visual misinformation, dangerous generation, and vulnerability to visual jailbreaks were observed. Mitigation strategies including prompt…
A research study examines how hybrid attention affects the multilingualism of large language models. It found that cross-lingual alignment is tied to the order of attention layers, and alternative layer orderings achieve up to 2.5x faster learning. The authors recommend starting with a full-attention layer.
PowerBench, a benchmark for testing the capabilities of LLM agents in autonomous information retrieval and reasoning in energy systems, is being released to the scientific community. The dataset contains 761 devices, 13.35 million hourly telemetry records, and 24 939 operational documents. The best evaluated models…
A research team introduced GenoMorph, a framework combining a DNA foundation model with adaptive latent computation for predicting genetic diseases. Instead of memorizing gene-disease associations, it uses pathway-based reasoning. On the benchmark, it achieves an F1 of 0.9725 (compared to 0.7863 for BioReason) and reduces latency by…
A research paper introduces the NLPG method for improving language agents without modifying the model's parameters. The method diagnoses errors in execution traces, propagates feedback through a graph, and converts failures into local natural language corrections. Tested on 6 benchmarks with an average improvement of 8.71…
The research paper presents a practical framework for designing human-AI collaboration that takes into account the human, the AI system, the task, the organization, and society. It includes a case study of a collaborative assembly system with a cobot and recommendations for effective, human-centred, and responsible design.
A research team introduced a framework that uses word embeddings from scientific literature to filter candidate compositions in materials discovery. On average, the framework eliminates 74.27 % of candidates with an error of 1.93 % relative to experiments and outperforms expert-chosen descriptors.
Research proposes BB-EDGE, a statistical framework for continuously valid evaluation of LLM models. It represents the leaderboard as a directed graph with certified edges and uses empirical-Bernstein e-processes with block factorization to ensure anytime-valid family-wise error rate (FWER) control without…
MIT researchers investigated the impact of algorithmic monoculture in hiring, where the same algorithm makes decisions across all companies. Their analysis shows that while monoculture creates information echoes that hinder talent discovery, ensemble methods can mitigate its effects. The results challenge earlier concerns about…
The new PhysFieldBench benchmark, with 24 tasks and 1160 examples, evaluates how multimodal models interpret physical fields. The best MLLM achieved a normalized score of 29.3 (close to chance). The solution is supervised fine-tuning with chain-of-thought supervision followed by reinforcement learning for the best…
A research team released the EEGAgentBench benchmark for systematically evaluating LLM agents in EEG analysis. The benchmark covers six applications with signals ranging from 2 seconds to 23 hours and includes 10 deterministic analytical tools. 29 models from 15 families were tested; the results show limitations of agents in long-term…
AI Radar monitors Czech and international sources every day, looking for changes that truly deserve attention.
MonitorsOfficial AI company blogs, specialist media, and research sources.
Selects and combinesFilters out information noise and combines articles about the same change into a single event.
Summarizes and explainsExplains significant events in English: what happened, why it matters and where the information comes from.
The result is a quick overview of what has actually changed in the AI world, rather than another stream of articles.
Use the CS/EN switch to read the same Radar in Czech or English. English content is published after its translation has been checked, so new and older items may appear later.
Everything you need to navigate the AI world
Today’s briefingThe “What is worth attention” selection sits beside Live · AI Flash, followed by research and links to other Radar sections. On mobile, these blocks appear one below another.
AI FlashAn ongoing feed of brief updates with an evidence status. Links lead to a Radar detail page when one is ready, otherwise to the original source. You can also find reset and outage histories here.
Practical applicationsWhat new tools and features can do, what you can try and what their actual impact could be.
Model selectionModel comparison by type of work, capabilities, price and speed.
Research and archiveA separate research overview, topic search and older events by date.
One event, everything that matters
Each row represents one event — not one article. At a glance, you can see its significance, credibility and main point.
Illustrative example, not a current news item.
Importance: ▮▮▮ majorOpenAIModels✓ 6
Agent mode is available to all paying users
Until now, the mode was available only on the highest plan; it is now available on all paid tiers without a waitlist.
▮▮▮ major · ▮▮ important · ▮ we're tracking = how significant the change is✓ 6 = six independent publishers, not the number of articles or feeds✓ official = a clear release, law or incident is substantiated by the relevant authority1 source = no independent confirmation yetbold = who is behind the changegray text = a brief summary of what happened
The detail page contains a fuller summary, its significance and original sources. Practical impact appears in the detail and the For individuals and For businesses views. An AI Flash item reaches the main selection only after it has been expanded and meets the publication rules.
The same news, two practical uses
We first summarize each event in the same way for everyone. Based on those same facts, we then explain what the change means for your own use and what it could mean for how a company operates.
For individualsWhat you can use or try, how the change can help you at work and what to watch out for.
For businessesWhat impact the change could have on processes, costs, risks and other business decisions.
Today’s briefing is the same for everyone. Pages
For individuals and For businesses
can be found in the main navigation — they select only events relevant to the given use case.