Skip to content

Research separate from news

What could become important next

A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.

253 published research events

Latest work

A significant claim from a single source is published only after further confirmation.

Research only one source so far

Benchmark of LLM argumentative behavior: defense against ad hominem attacks

The study benchmarks the ability of LLMs to respond to ad hominem attacks in dialogues. An analysis of a corpus of presidential debates shows that LLMs focus on logical defenses and cannot strategically employ ethical counterattacks. Safety fine-tuning limits their argumentative flexibility.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

UO-FIE method for ordinal factivity inference wins FIE2026 competition

The research paper introduces the UO-FIE method for classifying Chinese context-hypothesis pairs into nine ordinal factivity intervals. It combines direct supervision with a utility-oriented approach based on the Qwen3.5-9B model with LoRA. It achieved 1st place in the fine-tuning track of the FIE2026 competition with a macro utility score…

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

Evaluation of the SkinAgent AI agentic framework for safe skin care support

A scientific study evaluates the multimodal AI agent SkinAgent for supporting skin care. The system combines visual analysis, database information, and safety checks. It achieved 88.85% accuracy in determining skin type and 84.59% in assessing acne severity. No safety violations were detected during testing, but…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Domain adaptation of automatic speech recognition for adolescent healthcare communication in Ghanaian languages

A study compared ASR models across Ghanaian languages (Twi, Dagbani, Ewe). Fine-tuning Qwen3-ASR-0.6B on a ~90k corpus reduced WER for Ewe from 109.3% to 64.8%. KasaHealth application (50 users): 100% chat-approval, 72% Good translation, 92% would-recommend. Domain data proved to be the main limitation, not…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Two-dimensional framing analysis of French news headlines

A research team developed a two-dimensional framework for analyzing framing: salience (how to write) vs. selection (what to choose). They trained LLM annotators on 10 000 headlines, then applied the classifier to 902 111 headlines from 25 French media outlets (2022–2025). They found unequal salience for Jewish, right-wing, and Muslim…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Explainable DistilBERT-BiLSTM-Attention framework for hate speech detection

A research study presents a multilevel framework for hate speech detection combining DistilBERT embeddings, BiLSTM, and an attention mechanism with LIME explanation. It achieves an F1-score of 96.78–99.53% on binary classification, and 94.99–97.00% on multi-class classification on the Davidson and SMHS datasets.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

MIT developed an NLP tool for estimating suicide risk from texts

Scientists from the MIT McGovern Institute published a study on a new NLP tool for predicting suicide risk from texts of conversations with crisis counselors. The tool analyzes 49 risk factors and was trained on de-identified data from ~16 000 conversations with Crisis Text Line. Published in the Journal of…

MIT News – Artificial intelligence Original source ↗
Research only one source so far

Automating coherent long-form video generation

Google research on a multi-agent framework for autonomous generation of long-form videos. The framework addresses issues with identity drift and cascading errors in the AI pipeline. It builds on the Gemini and Veo models. The research papers will appear at the COLM 2026 and EMNLP 2026 conferences.

Google Research Blog Original source ↗
Research only one source so far

Analysis: AI experts underestimate the pace of progress in the field

An interim report from the Forecasting Research Institute analyzes 339 experts and finds that serious scientists, economists, and policymakers significantly underestimated the pace of AI development. AI achieved IMO gold in July 2025 (5 years earlier than experts' estimate), virology in April 2025 (5–9 years earlier), Anthropic achieved approx…

The Decoder (daily AI news) Original source ↗
Research confirmed by 4 independent sources

Anthropic's lab discovers a new enzyme with the help of Claude

Anthropic's lab in the Bay Area, with the help of the Claude model, discovered a new enzyme system in bacteriophage DNA that has CRISPR-like properties. The analysis took about 21 hours using ~950 agents and 210 million tokens. The physical experiments were carried out by humans at BSL-1/BSL-2 safety level.

Anthropic News · and 4 more sources Original source ↗
Research only one source so far

Research: prices for fixed AI performance are falling faster than for any other technology

Epoch AI and MIT: prices for fixed AI performance are falling 5–13x per year. Example: the o3 model (75% GPQA Diamond) cost 30 cents per question in 2025, GPT-5.6 achieves the same performance for 0.0004 cents (1/725 of the price). MIT: pure algorithmic progress is only 3x per year; the rest is hardware and competition.

The Decoder (daily AI news) Original source ↗
Research only one source so far

Evaluation of open models for Turkish domain documents in local deployment

Evaluation of five open 7B-8B LLMs for Turkish in local mode on an NVIDIA RTX 3050. Benchmark: 100 questions from an industrial R&D report, accuracy 49–75%. Key takeaway: none of the retrieval methods is better than the TF-IDF baseline; model, strategy, and hardware must be assessed separately.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Policy Distillation in Preparation for Reinforcement Learning

The study examines the effect of policy distillation (OPD) on training quality in reinforcement learning. Models initialized with distillation achieve higher final performance than direct reinforcement learning. The results show that reverse-KL OPD is more suitable before reinforcement-learning training, while forward-KL outperforms it afterward. Alignment…

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

Agent-Editing World Model for Improving LLM Agents

The new AEWM approach edits the erroneous state instead of simulating the environment. It achieves 70.5% macro-F1 on the Action Judge benchmark, improving agent performance by 3.2–6.7 points across six benchmarks.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

ChipMEM: verified memory for chip-design LLM agents

A research team introduced ChipMEM, a memory system for LLM-based agents working with electronic design automation (EDA) tools. It combines procedural memory with Bayesian estimates and distills skills only after successful verification. On the RTLRewriter-Bench benchmark: 39/54 correct designs vs. 35/54 without memory…

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

Targeting expert effort using disagreement between LLM models in annotation projects

The research paper presents a methodology for identifying cases where different LLMs disagree, in order to focus expert attention on the most problematic points when creating annotation guidelines. The Rationale Labeling method achieved 64.9% accuracy compared to the traditional approach (57.8%) and shortened the review time from months…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Drift Contract: Spectral Updates for Robust Local Learning

The study introduces the Drift Contract method of spectral updates for local learning (lr = epsilon/RMS(input)). On CIFAR-10 MLP: 48.9% accuracy vs 46.6% for local Adam (width 512), robustness to depth 12–48. With standard normalization, the benefits shift from local training to global.

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

NADI 2026: second edition of the Arabic speech processing shared task

NADI 2026 is the seventh edition of the shared task focused on Arabic dialects and the second edition dedicated to speech processing. It includes five tasks: automatic speech recognition (ASR), dialect identification (SDID), text-to-speech synthesis (TTS), speech translation (SLT), and spoken language understanding (SLU). 21 teams from at least 13 countries participated…

arXiv cs.CL (Computation and Language / NLP) Original source ↗