A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.
253
published research events
Latest work
A significant claim from a single source is published only after further confirmation.
The study benchmarks the ability of LLMs to respond to ad hominem attacks in dialogues. An analysis of a corpus of presidential debates shows that LLMs focus on logical defenses and cannot strategically employ ethical counterattacks. Safety fine-tuning limits their argumentative flexibility.
The research paper introduces the UO-FIE method for classifying Chinese context-hypothesis pairs into nine ordinal factivity intervals. It combines direct supervision with a utility-oriented approach based on the Qwen3.5-9B model with LoRA. It achieved 1st place in the fine-tuning track of the FIE2026 competition with a macro utility score…
A scientific study evaluates the multimodal AI agent SkinAgent for supporting skin care. The system combines visual analysis, database information, and safety checks. It achieved 88.85% accuracy in determining skin type and 84.59% in assessing acne severity. No safety violations were detected during testing, but…
A study compared ASR models across Ghanaian languages (Twi, Dagbani, Ewe). Fine-tuning Qwen3-ASR-0.6B on a ~90k corpus reduced WER for Ewe from 109.3% to 64.8%. KasaHealth application (50 users): 100% chat-approval, 72% Good translation, 92% would-recommend. Domain data proved to be the main limitation, not…
A research team developed a two-dimensional framework for analyzing framing: salience (how to write) vs. selection (what to choose). They trained LLM annotators on 10 000 headlines, then applied the classifier to 902 111 headlines from 25 French media outlets (2022–2025). They found unequal salience for Jewish, right-wing, and Muslim…
A research study presents a multilevel framework for hate speech detection combining DistilBERT embeddings, BiLSTM, and an attention mechanism with LIME explanation. It achieves an F1-score of 96.78–99.53% on binary classification, and 94.99–97.00% on multi-class classification on the Davidson and SMHS datasets.
Scientists from the MIT McGovern Institute published a study on a new NLP tool for predicting suicide risk from texts of conversations with crisis counselors. The tool analyzes 49 risk factors and was trained on de-identified data from ~16 000 conversations with Crisis Text Line. Published in the Journal of…
Google research on a multi-agent framework for autonomous generation of long-form videos. The framework addresses issues with identity drift and cascading errors in the AI pipeline. It builds on the Gemini and Veo models. The research papers will appear at the COLM 2026 and EMNLP 2026 conferences.
An interim report from the Forecasting Research Institute analyzes 339 experts and finds that serious scientists, economists, and policymakers significantly underestimated the pace of AI development. AI achieved IMO gold in July 2025 (5 years earlier than experts' estimate), virology in April 2025 (5–9 years earlier), Anthropic achieved approx…
Anthropic's lab in the Bay Area, with the help of the Claude model, discovered a new enzyme system in bacteriophage DNA that has CRISPR-like properties. The analysis took about 21 hours using ~950 agents and 210 million tokens. The physical experiments were carried out by humans at BSL-1/BSL-2 safety level.
Epoch AI and MIT: prices for fixed AI performance are falling 5–13x per year. Example: the o3 model (75% GPQA Diamond) cost 30 cents per question in 2025, GPT-5.6 achieves the same performance for 0.0004 cents (1/725 of the price). MIT: pure algorithmic progress is only 3x per year; the rest is hardware and competition.
Evaluation of five open 7B-8B LLMs for Turkish in local mode on an NVIDIA RTX 3050. Benchmark: 100 questions from an industrial R&D report, accuracy 49–75%. Key takeaway: none of the retrieval methods is better than the TF-IDF baseline; model, strategy, and hardware must be assessed separately.
The study examines the effect of policy distillation (OPD) on training quality in reinforcement learning. Models initialized with distillation achieve higher final performance than direct reinforcement learning. The results show that reverse-KL OPD is more suitable before reinforcement-learning training, while forward-KL outperforms it afterward. Alignment…
The new AEWM approach edits the erroneous state instead of simulating the environment. It achieves 70.5% macro-F1 on the Action Judge benchmark, improving agent performance by 3.2–6.7 points across six benchmarks.
A research team introduced ChipMEM, a memory system for LLM-based agents working with electronic design automation (EDA) tools. It combines procedural memory with Bayesian estimates and distills skills only after successful verification. On the RTLRewriter-Bench benchmark: 39/54 correct designs vs. 35/54 without memory…
The research paper presents a methodology for identifying cases where different LLMs disagree, in order to focus expert attention on the most problematic points when creating annotation guidelines. The Rationale Labeling method achieved 64.9% accuracy compared to the traditional approach (57.8%) and shortened the review time from months…
The study introduces the Drift Contract method of spectral updates for local learning (lr = epsilon/RMS(input)). On CIFAR-10 MLP: 48.9% accuracy vs 46.6% for local Adam (width 512), robustness to depth 12–48. With standard normalization, the benefits shift from local training to global.
NADI 2026 is the seventh edition of the shared task focused on Arabic dialects and the second edition dedicated to speech processing. It includes five tasks: automatic speech recognition (ASR), dialect identification (SDID), text-to-speech synthesis (TTS), speech translation (SLT), and spoken language understanding (SLU). 21 teams from at least 13 countries participated…
AI Radar monitors Czech and international sources every day, looking for changes that truly deserve attention.
MonitorsOfficial AI company blogs, specialist media, and research sources.
Selects and combinesFilters out information noise and combines articles about the same change into a single event.
Summarizes and explainsExplains significant events in English: what happened, why it matters and where the information comes from.
The result is a quick overview of what has actually changed in the AI world, rather than another stream of articles.
Use the CS/EN switch to read the same Radar in Czech or English. English content is published after its translation has been checked, so new and older items may appear later.
Everything you need to navigate the AI world
Today’s briefingThe “What is worth attention” selection sits beside Live · AI Flash, followed by research and links to other Radar sections. On mobile, these blocks appear one below another.
AI FlashAn ongoing feed of brief updates with an evidence status. Links lead to a Radar detail page when one is ready, otherwise to the original source. You can also find reset and outage histories here.
Practical applicationsWhat new tools and features can do, what you can try and what their actual impact could be.
Model selectionModel comparison by type of work, capabilities, price and speed.
Research and archiveA separate research overview, topic search and older events by date.
One event, everything that matters
Each row represents one event — not one article. At a glance, you can see its significance, credibility and main point.
Illustrative example, not a current news item.
Importance: ▮▮▮ majorOpenAIModels✓ 6
Agent mode is available to all paying users
Until now, the mode was available only on the highest plan; it is now available on all paid tiers without a waitlist.
▮▮▮ major · ▮▮ important · ▮ we're tracking = how significant the change is✓ 6 = six independent publishers, not the number of articles or feeds✓ official = a clear release, law or incident is substantiated by the relevant authority1 source = no independent confirmation yetbold = who is behind the changegray text = a brief summary of what happened
The detail page contains a fuller summary, its significance and original sources. Practical impact appears in the detail and the For individuals and For businesses views. An AI Flash item reaches the main selection only after it has been expanded and meets the publication rules.
The same news, two practical uses
We first summarize each event in the same way for everyone. Based on those same facts, we then explain what the change means for your own use and what it could mean for how a company operates.
For individualsWhat you can use or try, how the change can help you at work and what to watch out for.
For businessesWhat impact the change could have on processes, costs, risks and other business decisions.
Today’s briefing is the same for everyone. Pages
For individuals and For businesses
can be found in the main navigation — they select only events relevant to the given use case.