A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.
846
published research events
Latest work
A significant claim from a single source is published only after further confirmation.
A study on arXiv found that LLMs can identify their own generated text in programming and longer reflective tasks, but detection fails for short answers, and models often label them as “more human” than actual student texts.
The Rushes dataset (44 226 decisions from 8 167 users) shows that frontier LLMs, including GPT-5, do not outperform a simple popularity-based baseline method when predicting individual choices in AI-generated stories — models tend toward majority preferences instead of personalization.
Researchers have described a “split-knowledge attack" on RAG systems — combining individually harmless documents into a harmful association that filters such as LlamaGuard fail to detect. The proposed defense, TopoGuard, uses a document similarity graph and, according to the authors, detects significantly more attacks with low latency.
The ShriNep team described the RAKSHAK system (DeBERTa-v3-base with rationale distillation from Qwen2.5-14B and data from Jigsaw) for classifying toxic intent in World of Tanks chat. In the GameTox shared task, it achieved a Macro F1 of 0.5883 and ranked 7th out of 35 teams.
A study on arXiv shows that activation steering can smoothly control eight Jungian cognitive functions on Llama-3.1-8B instead of using the Big Five framework, with a dataset of over 2100 character narratives.
Researchers introduced SCoPE, a module for Emotion Recognition in Conversations (ERC) that builds on the emotional history of a particular speaker and combines it with multimodal signals. According to the authors, the model outperforms current state-of-the-art approaches on the IEMOCAP dataset.
A research study examining the relationship between the emotional tone (valence) of text and morality. The team created a dataset of 500 annotations from six people for the moral valence of actions and consequences, ranging from -1 to 1. A regularized logistic regression model achieved Matthew's correlation coefficient 0.764 for binary classification. The results suggest that valence is useful for estimating the morality of text.
Research introduces CAMeR, a memory system for LLM agents that combines gating at the word level (Jaccard) and the embedding level (cosine similarity) with adaptive weights. CAMeR-Bench, a new benchmark (76 memories, 100 rounds) in 8 thematic groups, shows 1.6× greater discrimination between frequently used and never…
Researchers discovered that routing in Mixture-of-Experts models (Phi-3.5-MoE, Gemma-4-27B-A4B, Qwen3.5-35B-A3B) follows Huffman coding principles — allocating sparse expert resources to common tokens and more diverse expert committees to rare tasks. They propose Subset Difference Pruning to…
Researchers propose Pulsar Attention to optimize language model inference. The method reduces FLOPs by 3.3× compared with Star Attention and achieves a 4.7% gain over dense attention on sequences up to 128K tokens in Llama-3.1-8B.
A research team in an arXiv study trained a 4B vision-language model to detect violations of 19 UI quality principles (WCAG 2.2, deceptive design, perception/cognition) on a synthetic dataset of 10 thousand pages. The model achieved 84% micro-F1 (up from 36%); 13 of the 19 principles exceeded 80% F1. The resulting critic can…
A research preprint proposes the SGRE method for protecting LLMs from knowledge distillation by modifying reasoning chains in responses. The method preserves accuracy and naturalness of the text while making them harder to copy. Tests demonstrated reduced distillation effectiveness without a loss of quality.
A research paper describes a method for detecting differences between model behavior during testing and ordinary deployment. The authors locate internal representations of these differences and test editing model behavior. The method works in 10 of 12 cases, but does not guarantee deployment safety.
A research team examines how to teach small language models (0.6–20B parameters) to generate MiniZinc code from verbal descriptions. They found that syntax errors dominate failures; they propose collecting errors from multiple runs and using them to train repairs. With fine-tuning, they achieved 98% execution accuracy…
A research paper proposes CSPF to improve evaluation of non-verifiable tasks by combining multiple reward models. The method decomposes evaluator signals into shared and specific representations under human preference supervision. On the LM-Arena and PPE datasets, it achieved the best results compared with…
The new Instruct-FD benchmark tests how well full-duplex spoken dialogue systems follow speech-control instructions. A study of six state-of-the-art systems found that even the best model follows only 64.4% of instructions. They particularly struggle with proactive behaviors such as backchanneling and interruption.
Researchers replaced simple lexical matching (Grievance Dictionary) with contextual models for detecting grievances in text. They showed that the original benchmark is biased (all "random" items are lexicon-negative), while the contextual approach improves accuracy, particularly for negated and…
The research paper describes a system for classifying toxicity in gaming chats. It combines an ensemble of compact transformers (DeBERTa-v3-base, XLM-RoBERTa-base) with a linguistically informed mediator. The system ranked 3rd in Macro F1 and 1st in accuracy among participants in the EEUCA 2026 competition. It addresses…
AI Radar monitors Czech and international sources every day, looking for changes that truly deserve attention.
MonitorsOfficial AI company blogs, specialist media, and research sources.
Selects and combinesFilters out information noise and combines articles about the same change into a single event.
Summarizes and explainsExplains significant events in English: what happened, why it matters and where the information comes from.
The result is a quick overview of what has actually changed in the AI world, rather than another stream of articles.
Use the CS/EN switch to read the same Radar in Czech or English. English content is published after its translation has been checked, so new and older items may appear later.
Everything you need to navigate the AI world
Today’s briefingThe “What is worth attention” selection sits beside Live · AI Flash, followed by research and links to other Radar sections. On mobile, these blocks appear one below another.
AI FlashAn ongoing feed of brief updates with an evidence status. Links lead to a Radar detail page when one is ready, otherwise to the original source. You can also find reset and outage histories here.
Practical applicationsWhat new tools and features can do, what you can try and what their actual impact could be.
Model selectionModel comparison by type of work, capabilities, price and speed.
Research and archiveA separate research overview, topic search and older events by date.
One event, everything that matters
Each row represents one event — not one article. At a glance, you can see its significance, credibility and main point.
Illustrative example, not a current news item.
Importance: ▮▮▮ majorOpenAIModels✓ 6
Agent mode is available to all paying users
Until now, the mode was available only on the highest plan; it is now available on all paid tiers without a waitlist.
▮▮▮ major · ▮▮ important · ▮ we're tracking = how significant the change is✓ 6 = six independent publishers, not the number of articles or feeds✓ official = a clear release, law or incident is substantiated by the relevant authority1 source = no independent confirmation yetbold = who is behind the changegray text = a brief summary of what happened
The detail page contains a fuller summary, its significance and original sources. Practical impact appears in the detail and the For individuals and For businesses views. An AI Flash item reaches the main selection only after it has been expanded and meets the publication rules.
The same news, two practical uses
We first summarize each event in the same way for everyone. Based on those same facts, we then explain what the change means for your own use and what it could mean for how a company operates.
For individualsWhat you can use or try, how the change can help you at work and what to watch out for.
For businessesWhat impact the change could have on processes, costs, risks and other business decisions.
Today’s briefing is the same for everyone. Pages
For individuals and For businesses
can be found in the main navigation — they select only events relevant to the given use case.