A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.
230
published research events4
new research papers today
Latest work
A significant claim from a single source is published only after further confirmation.
A study on 500 synthetic oncology cases found that eight LLM configurations failed to preserve the partial support label in 83.3-100% of relevant cases. According to the authors, the MTB-AuditAgent framework reduced excessive refusal to 6.7% and achieved an accuracy of 91.2%.
The study combines classifier output with lexical evidence. According to the authors, it reduces the area under the risk-coverage curve by 15.8% on BANKING77, 15.1% on CLINC150, and 11.8% on HWU64 compared to a learned approach using only semantic information.
The XOR-Trellis study presents two techniques for LLM weight quantization: simple parallel reconstruction and sensitivity-aware optimization of the model. According to the authors, they enable accurate compression at very low bit counts without the Hadamard transform; the abstract does not provide measurements.
The study presents Lingtai, a layer for tracking concept signals during LLM inference without training probes. The signals relate to predictive uncertainty, but in tests they do not provide a stable indicator of correctness. According to the authors, the overhead during code generation is 0.7–1.6% per token.
The study adapted the Ragas framework for Romanian and created the AdminRo-Eval dataset. The authors report 96% agreement with human evaluation for Faithfulness when decomposing the evaluation with the Gemini 2.5 Pro model, and 90% for Answer Relevance when comparing answers.
The study presents the generation of synthetic health data using Fuzzy Cognitive Maps running exclusively on CPU. On the Heart Disease dataset, the authors report an accuracy of up to 0.81 and AUROC of up to 0.90; the results are compared with the TVAE and Gaussian Copula methods.
The study organizes RAG research according to retrieval efficiency, robustness and safety, interactive procedures, and multi-step reasoning. It compares methods, architectures, and evaluation approaches and describes persistent reliability and scaling challenges.
A study across four tasks found that the geometric deviation of the Adam optimizer from the natural gradient descent method increases under poor conditioning. For a small neural network, it reached approximately 10³, but did not worsen the final loss value.
A study shows that concurrently generating dependent tokens from the distributions of individual positions does not reproduce the training distribution. On the ScanAndAdd task, it measured a total variation deviation 29× above the sampling noise level, even though per-sample metrics reached 1.0.
The study on English and French tested an auxiliary language classifier and per-language k-means targets. The phone-ABX error rate dropped from 11.6% to 10.4%, while the sWUGGY score rose from 52.1% to 56.7%. The interventions in the first iteration had the largest impact.
Researchers from MIT CSAIL, Google, and Northeastern University have introduced InstructMesh. The tool connects the TRELLIS system and the GPT-4 model and allows parts of a 3D design to be marked for editing before printing. They demonstrated the functionality on a mug, a whistle, and other objects.
Graphite compared model texts with 10 000 articles from before the launch of ChatGPT. According to the study, it found 13 000 phrases at least twice as common in generated texts; the Claude Opus 5.5 model used "this matters" 116 times more often.
Expert analysis points out that research on relationships with conversational AI has not verified all six conditions of emotional attachment. A survey of 7 027 people in four countries did not examine a key condition: seeking safety and comfort during genuine distress.
The study introduces CypherTurn with 721 conversations and 5 927 steps over 7 knowledge graphs. In evaluating 15 models, the best model achieved 64.7% query execution correctness; full-conversation correctness remained below 5%. A higher action budget did not resolve the gap in autonomous operation.
The study presents the Talk2Agent benchmark for voice instructions to agents. On 32 hours of human speech, the proposed metric showed a Pearson correlation with task success that was 0.246 higher than WER/CER, without requiring task execution.
The AREX-2 study trains an agent based on the Qwen3.8-27B model to iteratively improve its solving of programming tasks. The authors report a score of 81.8 on MLE-bench Lite and 84.0 on BrowseComp, along with further improvement as the number of rounds increases.
According to the authors, continuously updated probes during training reduce harmfulness and improve the truthfulness of models while preserving usefulness. The method shows a better safety-to-usefulness ratio than DPO and inference-time steering, and it preserves the ability to inspect internal representations.
The study presents the Alignment Forecasting method and the ALIGNMENTFORECASTBENCH benchmark with more than 5 000 questions. The method predicts fine-tuning risks better than the compared approaches; the benefit of data filtering for open-ended conversations remains unclear.
AI Radar monitors Czech and international sources every day, looking for changes that truly deserve attention.
MonitorsOfficial AI company blogs, specialist media, and research sources.
Selects and combinesFilters out information noise and combines articles about the same change into a single event.
Summarizes and explainsExplains significant events in English: what happened, why it matters and where the information comes from.
The result is a quick overview of what has actually changed in the AI world, rather than another stream of articles.
Use the CS/EN switch to read the same Radar in Czech or English. English content is published after its translation has been checked, so new and older items may appear later.
Everything you need to navigate the AI world
Today’s briefingThe “What is worth attention” selection sits beside Live · AI Flash, followed by research and links to other Radar sections. On mobile, these blocks appear one below another.
AI FlashAn ongoing feed of brief updates with an evidence status. Links lead to a Radar detail page when one is ready, otherwise to the original source. You can also find reset and outage histories here.
Practical applicationsWhat new tools and features can do, what you can try and what their actual impact could be.
Model selectionModel comparison by type of work, capabilities, price and speed.
Research and archiveA separate research overview, topic search and older events by date.
One event, everything that matters
Each row represents one event — not one article. At a glance, you can see its significance, credibility and main point.
Illustrative example, not a current news item.
Importance: ▮▮▮ majorOpenAIModels✓ 6
Agent mode is available to all paying users
Until now, the mode was available only on the highest plan; it is now available on all paid tiers without a waitlist.
▮▮▮ major · ▮▮ important · ▮ we're tracking = how significant the change is✓ 6 = six independent publishers, not the number of articles or feeds✓ official = a clear release, law or incident is substantiated by the relevant authority1 source = no independent confirmation yetbold = who is behind the changegray text = a brief summary of what happened
The detail page contains a fuller summary, its significance and original sources. Practical impact appears in the detail and the For individuals and For businesses views. An AI Flash item reaches the main selection only after it has been expanded and meets the publication rules.
The same news, two practical uses
We first summarize each event in the same way for everyone. Based on those same facts, we then explain what the change means for your own use and what it could mean for how a company operates.
For individualsWhat you can use or try, how the change can help you at work and what to watch out for.
For businessesWhat impact the change could have on processes, costs, risks and other business decisions.
Today’s briefing is the same for everyone. Pages
For individuals and For businesses
can be found in the main navigation — they select only events relevant to the given use case.