A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.
825
published research events
Latest work
A significant claim from a single source is published only after further confirmation.
The new ContextWeave benchmark evaluates memory in long-running agent workflows. It reproduces 1005 tasks from real work sequences of 14 users. The best memory configuration increased Workspace Score from 68.08 to 78.20 and Preference Score from 41.50 to 70.60. The finding: Rich, actionable memory supports…
The arXiv research article introduces a mathematical framework for evaluating LLM configurations under a budget constraint. It proposes algorithms optimizing the hypervolume-per-cost index with a logarithmic budget bound and exponentially decreasing error in Pareto identification.
The scientific study compared five approaches to feature selection for predicting opioid dependence from EHR data: recurrence enrichment, NTK-motivated early gradient sensitivity, LightGBM-SHAP, Elastic Net and LLM-guided semantic selection. NTK sensitivity achieved the best balance of accuracy and stability…
The arXiv research paper reproduces the claimed successes of self-distillation (SD) and explains its failures on QA, mathematics, coding and agentic tool use. The method works on simple tasks; on harder ones, training loss decreases but accuracy does not improve. The cause is ‘PI bias’: The teacher…
The arXiv study presents Redundancy-Adjusted Artificial Age Score, a mathematical framework for analyzing whether AI systems can operate indefinitely without unbounded growth in structural age. The main result: When redundancy conditions are met, a system can pass through infinitely many cycles with bounded age-related…
The research paper introduces DLR-Lock, a method protecting pretrained models against unauthorized adaptation. It replaces MLP layers with deep low-rank residual networks that increase memory during the backward pass and make fine-tuning harder. It preserves the model's original capabilities.
The Apple and UC Santa Barbara research team released DeepAmbigQA, a dataset of 3 600 questions for testing LLMs. A study with GPT-5 showed incomplete answers: 0.13 exact match on ambiguous questions and 0.21 on unambiguous ones.
The OncoTriad-QA benchmark combines radiological, pathological and genomic data with 86.1 thousand questions from 9 thousand patients. OncoVLM outperforms MedGemma-4B by 10.7 points in multimodal oncology analysis.
A research paper from arXiv introduces DDRSR (Deep Divide-and-Reduce), a method that improves the approach to symbolic regression. The new method addresses problems with the previous method, AI Feynman – it expands the options for decomposing expressions, avoids brute-force search and increases practical usability. It includes…
The research team publishes AS-FedBridge, a framework for federated learning on resource-constrained devices. It combines artificial neural networks and spiking neural networks, which offer greater energy efficiency. The framework introduces a bridge with a Pseudo-Spike interface for transforming continuous signals into…
The ANCHOR-RE research framework integrates ontology-guided reasoning and verification rules into LLM inference for biomedical relation extraction. It improved performance on three benchmarks: SemRepGS 0.654→0.676, DDI 0.769→0.872, ChemProt 0.939→0.941; accuracy on post-cutoff data was 69%.
OptR is a new method that minimizes errors in INT2 quantization of large language models' KV-cache by optimizing the post-attention output. Tested on three models and five benchmarks, it improves long contexts.
The research team converted 21 of the 28 attention layers in Qwen3-0.6B-Base to KDA linear attention. After conversion, the model fixated on the answer position (choosing ‘A’ 81% of the time) rather than following the content, and accuracy fell to 25–29%. Format-targeted KL + SFT + DPO raised C-Eval by +12.48 points. Code and weights released.
The research paper proposes a method for aligning vision-language models at inference time through trajectory learning and MCMC. It shows improved accuracy on multimodal datasets without large computational costs.
The research paper introduces GoT-CD, which discovers causal graphs from data through LLM-based Graph-of-Thoughts. It achieves the best F1 score on three benchmarks but shows that structural accuracy does not guarantee correct fairness audits — on the Alzheimer's dataset, 5 of 8 graphs found no path from…
The MemArena benchmark tests personal AI assistants with memory management on edge devices using open-weight models. The research involves 50 agents communicating over 15 days. Key findings: the memory backend is more important for accuracy than model size; all systems struggle with permission-aware access…
The article introduces TabletCraft, an open-source system for bidirectional translation between Akkadian and English. The ByT5 model trained on 116K samples achieves 49.1 BLEU for Akkadian→English and 48.5 BLEU for English→Akkadian (the first published result in the reverse direction). The system integrates…
The Apple ML Research team identified a problem with outlier tokens in diffusion transformers (DiTs) for image generation — a few high-norm tokens reduce quality. They proposed Dual-Stage Registers (DSR), with training-time and test-time registers. Testing on ImageNet and in text-to-image…
AI Radar monitors Czech and international sources every day, looking for changes that truly deserve attention.
MonitorsOfficial AI company blogs, specialist media, and research sources.
Selects and combinesFilters out information noise and combines articles about the same change into a single event.
Summarizes and explainsExplains significant events in English: what happened, why it matters and where the information comes from.
The result is a quick overview of what has actually changed in the AI world, rather than another stream of articles.
Use the CS/EN switch to read the same Radar in Czech or English. English content is published after its translation has been checked, so new and older items may appear later.
Everything you need to navigate the AI world
Today’s briefingThe “What is worth attention” selection sits beside Live · AI Flash, followed by research and links to other Radar sections. On mobile, these blocks appear one below another.
AI FlashAn ongoing feed of brief updates with an evidence status. Links lead to a Radar detail page when one is ready, otherwise to the original source. You can also find reset and outage histories here.
Practical applicationsWhat new tools and features can do, what you can try and what their actual impact could be.
Model selectionModel comparison by type of work, capabilities, price and speed.
Research and archiveA separate research overview, topic search and older events by date.
One event, everything that matters
Each row represents one event — not one article. At a glance, you can see its significance, credibility and main point.
Illustrative example, not a current news item.
Importance: ▮▮▮ majorOpenAIModels✓ 6
Agent mode is available to all paying users
Until now, the mode was available only on the highest plan; it is now available on all paid tiers without a waitlist.
▮▮▮ major · ▮▮ important · ▮ we're tracking = how significant the change is✓ 6 = six independent publishers, not the number of articles or feeds✓ official = a clear release, law or incident is substantiated by the relevant authority1 source = no independent confirmation yetbold = who is behind the changegray text = a brief summary of what happened
The detail page contains a fuller summary, its significance and original sources. Practical impact appears in the detail and the For individuals and For businesses views. An AI Flash item reaches the main selection only after it has been expanded and meets the publication rules.
The same news, two practical uses
We first summarize each event in the same way for everyone. Based on those same facts, we then explain what the change means for your own use and what it could mean for how a company operates.
For individualsWhat you can use or try, how the change can help you at work and what to watch out for.
For businessesWhat impact the change could have on processes, costs, risks and other business decisions.
Today’s briefing is the same for everyone. Pages
For individuals and For businesses
can be found in the main navigation — they select only events relevant to the given use case.