Skip to content

Research separate from news

What could become important next

A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.

846 published research events

Latest work

A significant claim from a single source is published only after further confirmation.

Research only one source so far

Study: language models, including GPT-5, perform no better at personalizing preferences than simple statistical methods

The Rushes dataset (44 226 decisions from 8 167 users) shows that frontier LLMs, including GPT-5, do not outperform a simple popularity-based baseline method when predicting individual choices in AI-generated stories — models tend toward majority preferences instead of personalization.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Researchers have described a new type of attack on RAG systems and a defense using graph theory (TopoGuard)

Researchers have described a “split-knowledge attack" on RAG systems — combining individually harmless documents into a harmful association that filters such as LlamaGuard fail to detect. The proposed defense, TopoGuard, uses a document similarity graph and, according to the authors, detects significantly more attacks with low latency.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

New SCoPE method combines a speaker's emotional history with multimodal signals for emotion recognition in conversations

Researchers introduced SCoPE, a module for Emotion Recognition in Conversations (ERC) that builds on the emotional history of a particular speaker and combines it with multimodal signals. According to the authors, the model outperforms current state-of-the-art approaches on the IEMOCAP dataset.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Valence and morality in natural language: a pilot annotation study

A research study examining the relationship between the emotional tone (valence) of text and morality. The team created a dataset of 500 annotations from six people for the moral valence of actions and consequences, ranging from -1 to 1. A regularized logistic regression model achieved Matthew's correlation coefficient 0.764 for binary classification. The results suggest that valence is useful for estimating the morality of text.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

CAMeR: Adaptive memory in LLM agents with hybrid word and embedding gating

Research introduces CAMeR, a memory system for LLM agents that combines gating at the word level (Jaccard) and the embedding level (cosine similarity) with adaptive weights. CAMeR-Bench, a new benchmark (76 memories, 100 rounds) in 8 thematic groups, shows 1.6× greater discrimination between frequently used and never…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Mixture-of-Experts routing follows Huffman coding principles, study shows

Researchers discovered that routing in Mixture-of-Experts models (Phi-3.5-MoE, Gemma-4-27B-A4B, Qwen3.5-35B-A3B) follows Huffman coding principles — allocating sparse expert resources to common tokens and more diverse expert committees to rare tasks. They propose Subset Difference Pruning to…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Detecting user interface principle violations with reinforcement learning

A research team in an arXiv study trained a 4B vision-language model to detect violations of 19 UI quality principles (WCAG 2.2, deceptive design, perception/cognition) on a synthetic dataset of 10 thousand pages. The model achieved 84% micro-F1 (up from 36%); 13 of the 19 principles exceeded 80% F1. The resulting critic can…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Protecting language models from knowledge distillation by editing logical schemas

A research preprint proposes the SGRE method for protecting LLMs from knowledge distillation by modifying reasoning chains in responses. The method preserves accuracy and naturalness of the text while making them harder to copy. Tests demonstrated reduced distillation effectiveness without a loss of quality.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Auditing evaluation–deployment mismatch in fine-tuned language models

A research paper describes a method for detecting differences between model behavior during testing and ordinary deployment. The authors locate internal representations of these differences and test editing model behavior. The method works in 10 of 12 cases, but does not guarantee deployment safety.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Learn2Zinc: Fine-tuning small language models to translate natural language into MiniZinc

A research team examines how to teach small language models (0.6–20B parameters) to generate MiniZinc code from verbal descriptions. They found that syntax errors dominate failures; they propose collecting errors from multiple runs and using them to train repairs. With fine-tuning, they achieved 98% execution accuracy…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Shared-private fusion of reward models for preference evaluation

A research paper proposes CSPF to improve evaluation of non-verifiable tasks by combining multiple reward models. The method decomposes evaluator signals into shared and specific representations under human preference supervision. On the LM-Arena and PPE datasets, it achieved the best results compared with…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Instruct-FD: How well full-duplex dialogue systems follow instructions for speech control

The new Instruct-FD benchmark tests how well full-duplex spoken dialogue systems follow speech-control instructions. A study of six state-of-the-art systems found that even the best model follows only 64.4% of instructions. They particularly struggle with proactive behaviors such as backchanneling and interruption.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Detecting toxicity in gaming chats using a neuro-symbolic approach

The research paper describes a system for classifying toxicity in gaming chats. It combines an ensemble of compact transformers (DeBERTa-v3-base, XLM-RoBERTa-base) with a linguistically informed mediator. The system ranked 3rd in Macro F1 and 1st in accuracy among participants in the EEUCA 2026 competition. It addresses…

arXiv cs.CL (Computation and Language / NLP) Original source ↗