Skip to content

Research separate from news

What could become important next

A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.

846 published research events

Latest work

A significant claim from a single source is published only after further confirmation.

Research only one source so far

Auditing data pipelines for language model alignment and evaluation

The arXiv study introduces an efficient method for auditing LLM training data through data influence scoring without retraining. The authors applied the approach to HelpSteer2 and Anthropic HH-RLHF datasets, identifying types of labeling errors and safety problems in benchmarks.

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

Dynamic fact-checking evaluation remains contaminated by claims after models' cutoff dates

An arXiv preprint shows that dynamic benchmarks for multimodal fact-checking are less uncontaminated than assumed. Approximately 17–29% of claims published after models' cutoff dates contain previously available knowledge, increasing system performance by up to 11 points and distorting evaluation.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

QFedPolyp: Federated learning with quantization for polyp segmentation

The scientific paper proposes QFedPolyp, a federated learning framework with 8-bit quantization for polyp segmentation. It achieves a 4-fold reduction in communication between servers while maintaining accuracy and increasing inference speed. Tested on the Kvasir-SEG and CVC-ClinicVideoDB datasets.

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

StanceFlip: A Benchmark for Predicting Stance Changes in Multimodal Conversations

A new scientific benchmark, StanceFlip, for analyzing dynamic stance changes in multimodal dialogues. It includes the extraction of a six-component stance (subject, target, emotion, sentiment, stance, reason) and the attribution of changes. The authors propose the ConStaFF framework with a Thought-of-Stance methodology that achieves state-of-the-art…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Memory-efficient audio synthesis in Siri Expressive Voices

Apple published a technical study on the audio synthesis architecture in Siri Expressive Voices. The system uses Diffusion Transformers with separate temporal and depth components on the AMX chip, just ~21MB of memory and execution 16× faster than real time. The MOS score improved by +0.28 (4.15 vs 3.87).

Apple Machine Learning Research Original source ↗
Research only one source so far

Study: vision-language models often “correct" the transcription of damaged text instead of transcribing it faithfully

A new study introduces the FaithC4 benchmark (1 455 documents, English, Chinese, Korean) and shows that general-purpose vision-language models performing OCR on damaged text often replace an unreadable word with a more probable alternative instead of transcribing it faithfully, while specialized OCR models are significantly…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

New regularization method UOWReg reduces bias in self-supervised learning

Researchers introduced UOWReg, a regularization technique for self-supervised learning and JEPA models intended to suppress the encoding of biases based on sensitive attributes. According to the authors, it reduces violations of the Equalized Odds metric on the CelebA benchmark while maintaining comparable classification accuracy.

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

Researchers introduced HierFlow, a training-free method for automatically designing agent workflows for LLMs

Researchers introduced HierFlow, a method for automatically designing workflows for multi-agent systems with large language models without training. According to the authors, it outperforms comparison approaches on question-answering, mathematics, and code generation tasks while maintaining efficiency.

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Study proposes evaluating LLMs based on mutual agreement in preferences among models

A new study on arXiv proposes evaluating LLMs based on how other models express preferences for their responses in anonymous voting, instead of comparing them with a correct answer. However, the authors caution that this reflects agreement in preferences among models, not objective correctness or agreement with human judgment.

arXiv cs.CL (Computation and Language / NLP) Original source ↗