Skip to content

Research separate from news

What could become important next

A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.

846 published research events

Latest work

A significant claim from a single source is published only after further confirmation.

Research only one source so far

Portable latency prediction for optimizing LLM deployment on heterogeneous edge devices

The research team presents a framework for more accurate latency prediction when deploying LLMs on mobile and edge devices. On Pixel 8, it improved decoding accuracy from R² 0.957 to 0.973, and on Pixel 8 Pro, prefill latency from -1.383 to 0.966. The method combines static parameters with dynamic hardware telemetry.

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Towards standardized testing of personal LLM agents over time

A research paper proposes a new framework for evaluating personal LLM agents. The authors identify four conditions for effective testing: explicit temporal interventions, persistent user state, cross-component effects, and configuration variation. After auditing published benchmarks, they found no protocol…

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

CEL: A library and benchmark for AI explanation methods

Introduces CEL, a library with 18 datasets and 14 counterfactual explanation methods. It includes a unified evaluation protocol with metrics for validity, coverage, sparsity, and plausibility. The first comprehensive benchmark for systematically comparing methods within a unified framework.

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

Improving compositional generation in diffusion models through intrinsic reward

The research paper introduces TILT, a method for improving image generation from text descriptions. It addresses diffusion model failures on complex prompts involving multiple concepts. The method uses the model's intrinsic reward without external supervision. Results on the T2ICompBench benchmark show improvement without reducing quality…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

RED-PIM: Reducing transformer data movement through processing in memory

RED-PIM research optimizes transformer processing through algorithm–architecture co-design. The method reduces inter-bank data movement from O(N²) to O(N) and shrinks cache matrices from N×N to d×d. On real data, it improves inference by 99.60 % for long documents and 13.44 % for shorter…

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

Trajectory-aware agents for temporal decision-making

Researchers introduced TLM for processing long, temporally structured texts using LLM agents. Traditional RAG fragments chronological context; TLM preserves it through LGCM and SHAP feedback. It was tested on medical questions and financial event prediction, with better results.

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Multi-horizon consistency as the geometry of latent dynamics

An arXiv preprint examines how the weight of multi-horizon latent consistency affects video predictor geometry. On Moving-MNIST, increasing lambda from 0 to 0.8 reduced latent-space expansivity (L20 from 4.96 to 1.01) and prediction error. The same effect does not hold in other domains (Pendulum…

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

Industrial tokenization for diagnostics: A federated architecture for integrating heterogeneous data into LLMs

The research paper presents the concept of Industrial Tokenization — a way of structuring heterogeneous industrial data from sensors, maintenance and diagnostic systems into standardized units for interpretation using large language models. The pilot implementation includes vibration diagnostics.

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

From frame recognition to event-level confirmation: Failure analysis of gesture interaction in public spaces

The research analyzes 8 records from deployments of gesture interaction in kiosks. It identifies 20 failures arising from the difference between individual-frame recognition and the stability of interaction events. It defines 6 failure classes and proposes a runtime abstraction for event confirmation.

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Research into toxic behavior in the Mastodon community

A scientific paper from arXiv uses machine learning to examine how toxic behavior develops and spreads in the decentralized social network Mastodon, which has no uniform moderation standards. The goal is to understand toxicity trends and their impact on community health.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Convergence of stochastic low-rank adaptation

The research paper improves the theoretical analysis of convergence for LoRA (Low-Rank Adaptation). It shows that finding a stationary point requires only O(ε⁻⁴) gradient evaluations instead of an exponential number. It proposes the LoRA-NSGDM and LoRA-STORM algorithms with improved oracle complexity for the stochastic setting.

arXiv cs.LG (Machine Learning) Original source ↗