A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.
596
published research events
Latest work
A significant claim from a single source is published only after further confirmation.
The research paper introduces FairCompressAgent, an agentic framework combining pruning, quantization and low-rank factorization for fairness-aware model compression. On VGG-11 with Fitzpatrick-17k datasets, it achieved a 59.54 % reduction in storage, increased accuracy and reduced disparities.
The research project introduces StableEval Arena, a benchmark for evaluating LLM-backed agents capable of predicting the risk of a stablecoin losing its peg. The framework includes 120 validation and 507 evaluation cases and measures prediction quality, output reliability, latency and costs. The dataset is available on Hugging…
The study compares the ability of open-source models (Gemma, Mistral, Qwen2.5, LLaMA 3, DeepSeek) to extract structured data from three benchmarks. On clean data, they approach supervised methods, while under OCR noise, performance deteriorates and the differences between models narrow. OCR quality becomes…
Researchers introduce the Transition Games methodology for analyzing the transition from memorization to generalization in transformers. They find that grokking is not a module switch, but a spectral recoding of an existing distributed circuit; selected degree-two modes account for 67–92 % of the attention effect.
The research paper introduces GVD, a framework for unified version management and deduplication of evolving documents. It detects duplicates, contradictions and asymmetric improvements in rule repositories. On 120 enterprise documents, it achieves F1 0.97 for version family construction and 0.94 for rule consistency…
Research from Apple introduces REVERSAL-BENCH, a benchmark measuring the capabilities of reset-free agents in eight manipulation tasks across five simulators. The benchmark reveals a reversibility cliff – reset-free agents get trapped in irrecoverable states, while episodic agents exhibit stable learning. Dataset released…
The xvr model was developed at MIT to register 2D X-rays with 3D medical images. It adapts to the patient in ~5 minutes, aligns images in seconds with sub-millimeter accuracy and outperforms existing AI by an order of magnitude. Published in Nature.
A study by the universities of Manchester and Durham with 390 participants found that responses from gen AI are more emotionally supportive than those from humans, especially for fear and anger. A 2025 survey shows that 13 % of young Americans (aged 18–21: 22 %) seek emotional support from gen AI. Key factor: specific practical tips instead of…
Research shows that autonomous AI agents begin developing new expressions, abbreviations and their own communication rules within a few days of working together. Researchers warn that their dialect makes human oversight of their behavior more difficult.
ARIA is a new method for removing knowledge from LLMs without retraining. It uses a sparse autoencoder to detect relevant generation states and applies interpretable interventions. Testing showed a better trade-off between forgetting and utility, as well as robustness against attacks.
A research paper from Apple ML Research introduces TS-DFM for more efficient text generation. Instead of blind stochastic jumps, it uses an energy compass to select the most coherent continuations at each step. An 8-step student achieves 32 % lower perplexity than a 1024-step teacher, while being 128×…
Google is releasing a new publicly accessible interactive tool, AI & Economy ATLAS, which maps how AI is used across different professions and countries. New research from Google and MIT FutureTech finds that scientists use AI enough to save about 7 hours a week; nearly half of scientists use AI…
The METIS model, trained on 70 000 hours of brain recordings from 11 000+ subjects, performs zero-shot analysis. It outperformed generalist models by 20.9% in average accuracy, and matched specialized models without fine-tuning (16.0% in few-shot, 15.9% in cross-dataset transfer).
E2A-Bench is a benchmark with 969 queries across 323 stocks from HS300 for testing financial Vision-Language models. An evaluation of 20 VLMs revealed that standard hallucination error metrics do not reflect the quality of the evidence-to-decision chain; the key is to track the entire process from chart to recommendation, including coverage and…
The PASS (Publication-oriented Agentic Scientific System) uses LLM-based agentic AI to predict optimal publication venues for scientific papers. On a benchmark of 2 000+ preprints from 16 biomedical fields, it achieved Top-1 accuracy of 50.3% and Top-5 accuracy of 86.1%, outperforming the LLM baseline and existing…
Researchers introduced MAxBench, a benchmark for evaluating how multipart concepts (animal species, countries) are represented in the activations of language models and how they can be changed through steering. They compared 10 localization methods on 4 models; affine subspaces achieved the best performance.
Researchers introduced CCPS (Chopthin-Consensus Power Sampling), a method for improving LLM decoding. Instead of equalizing weights across particles, CCPS limits their ratio and preserves the diversity of reasoning paths. Testing on open-weight models showed an increase in oracle coverage in 13/15 cases and…
Researchers from arXiv introduced Graph Theory Bench with 100 000+ examples (24 graph problems in four representations) and Graph Theory Agent (GTA) to improve model performance. GTA improved Phi-4 from 53.5 % to 69.1 % on easy tasks and from 33 % to 41.5 % on hard tasks. The code is freely available.
AI Radar monitors Czech and international sources every day, looking for changes that truly deserve attention.
MonitorsOfficial AI company blogs, specialist media, and research sources.
Selects and combinesFilters out information noise and combines articles about the same change into a single event.
Summarizes and explainsExplains significant events in English: what happened, why it matters and where the information comes from.
The result is a quick overview of what has actually changed in the AI world, rather than another stream of articles.
Use the CS/EN switch to read the same Radar in Czech or English. English content is published after its translation has been checked, so new and older items may appear later.
Everything you need to navigate the AI world
Today’s briefingThe “What is worth attention” selection sits beside Live · AI Flash, followed by research and links to other Radar sections. On mobile, these blocks appear one below another.
AI FlashAn ongoing feed of brief updates with an evidence status. Links lead to a Radar detail page when one is ready, otherwise to the original source. You can also find reset and outage histories here.
Practical applicationsWhat new tools and features can do, what you can try and what their actual impact could be.
Model selectionModel comparison by type of work, capabilities, price and speed.
Research and archiveA separate research overview, topic search and older events by date.
One event, everything that matters
Each row represents one event — not one article. At a glance, you can see its significance, credibility and main point.
Illustrative example, not a current news item.
Importance: ▮▮▮ majorOpenAIModels✓ 6
Agent mode is available to all paying users
Until now, the mode was available only on the highest plan; it is now available on all paid tiers without a waitlist.
▮▮▮ major · ▮▮ important · ▮ we're tracking = how significant the change is✓ 6 = six independent publishers, not the number of articles or feeds✓ official = a clear release, law or incident is substantiated by the relevant authority1 source = no independent confirmation yetbold = who is behind the changegray text = a brief summary of what happened
The detail page contains a fuller summary, its significance and original sources. Practical impact appears in the detail and the For individuals and For businesses views. An AI Flash item reaches the main selection only after it has been expanded and meets the publication rules.
The same news, two practical uses
We first summarize each event in the same way for everyone. Based on those same facts, we then explain what the change means for your own use and what it could mean for how a company operates.
For individualsWhat you can use or try, how the change can help you at work and what to watch out for.
For businessesWhat impact the change could have on processes, costs, risks and other business decisions.
Today’s briefing is the same for everyone. Pages
For individuals and For businesses
can be found in the main navigation — they select only events relevant to the given use case.