Skip to content

Research separate from news

What could become important next

A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.

596 published research events

Latest work

A significant claim from a single source is published only after further confirmation.

Research only one source so far

FairCompressAgent framework for fairness-aware model compression on FPGA

The research paper introduces FairCompressAgent, an agentic framework combining pruning, quantization and low-rank factorization for fairness-aware model compression. On VGG-11 with Fitzpatrick-17k datasets, it achieved a 59.54 % reduction in storage, increased accuracy and reduced disparities.

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

StableEval Arena: a benchmarking framework for predicting stablecoin stability using AI agents

The research project introduces StableEval Arena, a benchmark for evaluating LLM-backed agents capable of predicting the risk of a stablecoin losing its peg. The framework includes 120 validation and 507 evaluation cases and measures prediction quality, output reliability, latency and costs. The dataset is available on Hugging…

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

Key-value extraction from documents: a benchmark of open-source LLMs under OCR noise

The study compares the ability of open-source models (Gemma, Mistral, Qwen2.5, LLaMA 3, DeepSeek) to extract structured data from three benchmarks. On clean data, they approach supervised methods, while under OCR noise, performance deteriorates and the differences between models narrow. OCR quality becomes…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Where grokking takes place: distributed utility and Fourier recoding without a module switch

Researchers introduce the Transition Games methodology for analyzing the transition from memorization to generalization in transformers. They find that grokking is not a module switch, but a spectral recoding of an existing distributed circuit; selected degree-two modes account for 67–92 % of the attention effect.

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

GVD: A Framework for Version Management and Deduplication in Document Repositories

The research paper introduces GVD, a framework for unified version management and deduplication of evolving documents. It detects duplicates, contradictions and asymmetric improvements in rule repositories. On 120 enterprise documents, it achieves F1 0.97 for version family construction and 0.94 for rule consistency…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

REVERSAL-BENCH: a reversibility benchmark for reset-free agent learning

Research from Apple introduces REVERSAL-BENCH, a benchmark measuring the capabilities of reset-free agents in eight manipulation tasks across five simulators. The benchmark reveals a reversibility cliff – reset-free agents get trapped in irrecoverable states, while episodic agents exhibit stable learning. Dataset released…

Apple Machine Learning Research Original source ↗
Research only one source so far

Generative AI provides emotional support more effectively than humans, research finds

A study by the universities of Manchester and Durham with 390 participants found that responses from gen AI are more emotionally supportive than those from humans, especially for fear and anger. A 2025 survey shows that 13 % of young Americans (aged 18–21: 22 %) seek emotional support from gen AI. Key factor: specific practical tips instead of…

The Conversation — Artificial Intelligence Original source ↗
Research only one source so far

Machine unlearning at inference time: The ARIA method

ARIA is a new method for removing knowledge from LLMs without retraining. It uses a sparse autoencoder to detect relevant generation states and applies interpretable interventions. Testing showed a better trade-off between forgetting and utility, as well as robustness against attacks.

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

TS-DFM: 128× faster text generation through energy-guided distillation

A research paper from Apple ML Research introduces TS-DFM for more efficient text generation. Instead of blind stochastic jumps, it uses an energy compass to select the most coherent continuations at each step. An 8-step student achieves 32 % lower perplexity than a 1024-step teacher, while being 128×…

Apple Machine Learning Research Original source ↗
Research only one source so far

E2A-Bench: A benchmark for evaluating financial Vision-Language models

E2A-Bench is a benchmark with 969 queries across 323 stocks from HS300 for testing financial Vision-Language models. An evaluation of 20 VLMs revealed that standard hallucination error metrics do not reflect the quality of the evidence-to-decision chain; the key is to track the entire process from chart to recommendation, including coverage and…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Positioning scientific publications in the literature using AI with autonomous agents

The PASS (Publication-oriented Agentic Scientific System) uses LLM-based agentic AI to predict optimal publication venues for scientific papers. On a benchmark of 2 000+ preprints from 16 biomedical fields, it achieved Top-1 accuracy of 50.3% and Top-5 accuracy of 86.1%, outperforming the LLM baseline and existing…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

MAxBench: a benchmark for recovering multinomial concepts in language models

Researchers introduced MAxBench, a benchmark for evaluating how multipart concepts (animal species, countries) are represented in the activations of language models and how they can be changed through steering. They compared 10 localization methods on 4 models; affine subspaces achieved the best performance.

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

Chopthin-Consensus Power Sampling: a method for diversity in LLM inference

Researchers introduced CCPS (Chopthin-Consensus Power Sampling), a method for improving LLM decoding. Instead of equalizing weights across particles, CCPS limits their ratio and preserves the diversity of reasoning paths. Testing on open-weight models showed an increase in oracle coverage in 13/15 cases and…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Graph Theory Agent and a benchmark for graph reasoning in LLMs

Researchers from arXiv introduced Graph Theory Bench with 100 000+ examples (24 graph problems in four representations) and Graph Theory Agent (GTA) to improve model performance. GTA improved Phi-4 from 53.5 % to 69.1 % on easy tasks and from 33 % to 41.5 % on hard tasks. The code is freely available.

arXiv cs.AI (Artificial Intelligence) Original source ↗