Skip to content

Research separate from news

What could become important next

A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.

825 published research events

Latest work

A significant claim from a single source is published only after further confirmation.

Research only one source so far

ContextWeave: A benchmark for memory in long-running agent workflows

The new ContextWeave benchmark evaluates memory in long-running agent workflows. It reproduces 1005 tasks from real work sequences of 14 users. The best memory configuration increased Workspace Score from 68.08 to 78.20 and Preference Score from 41.50 to 70.60. The finding: Rich, actionable memory supports…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Efficient evaluation of LLM configurations on a limited budget

The arXiv research article introduces a mathematical framework for evaluating LLM configurations under a budget constraint. It proposes algorithms optimizing the hypervolume-per-cost index with a logarithmic budget bound and exponentially decreasing error in Pareto identification.

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

Comparing feature selection methods for predicting opioid dependence from electronic health records

The scientific study compared five approaches to feature selection for predicting opioid dependence from EHR data: recurrence enrichment, NTK-motivated early gradient sensitivity, LightGBM-SHAP, Elastic Net and LLM-guided semantic selection. NTK sensitivity achieved the best balance of accuracy and stability…

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

Self-distillation with privileged information fails on difficult tasks

The arXiv research paper reproduces the claimed successes of self-distillation (SD) and explains its failures on QA, mathematics, coding and agentic tool use. The method works on simple tasks; on harder ones, training loss decreases but accuracy does not improve. The cause is ‘PI bias’: The teacher…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

A model of long-term AI system persistence without cumulative deterioration

The arXiv study presents Redundancy-Adjusted Artificial Age Score, a mathematical framework for analyzing whether AI systems can operate indefinitely without unbounded growth in structural age. The main result: When redundancy conditions are met, a system can pass through infinitely many cycles with bounded age-related…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Protecting pretrained weights with deep low-rank residual networks

The research paper introduces DLR-Lock, a method protecting pretrained models against unauthorized adaptation. It replaces MLP layers with deep low-rank residual networks that increase memory during the backward pass and make fine-tuning harder. It preserves the model's original capabilities.

Apple Machine Learning Research Original source ↗
Research only one source so far

OncoTriad-QA: A benchmark for oncology diagnosis

The OncoTriad-QA benchmark combines radiological, pathological and genomic data with 86.1 thousand questions from 9 thousand patients. OncoVLM outperforms MedGemma-4B by 10.7 points in multimodal oncology analysis.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

New DDRSR method for symbolic regression

A research paper from arXiv introduces DDRSR (Deep Divide-and-Reduce), a method that improves the approach to symbolic regression. The new method addresses problems with the previous method, AI Feynman – it expands the options for decomposing expressions, avoids brute-force search and increases practical usability. It includes…

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

AS-FedBridge: federated learning with mixed artificial and spiking neurons

The research team publishes AS-FedBridge, a framework for federated learning on resource-constrained devices. It combines artificial neural networks and spiking neural networks, which offer greater energy efficiency. The framework introduces a bridge with a Pseudo-Spike interface for transforming continuous signals into…

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

ANCHOR-RE: A neuro-symbolic framework for biomedical relation extraction

The ANCHOR-RE research framework integrates ontology-guided reasoning and verification rules into LLM inference for biomedical relation extraction. It improved performance on three benchmarks: SemRepGS 0.654→0.676, DDI 0.769→0.872, ChemProt 0.939→0.941; accuracy on post-cutoff data was 69%.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Interface damage during KDA linearization: Diagnosing and repairing Qwen3-0.6B-Base

The research team converted 21 of the 28 attention layers in Qwen3-0.6B-Base to KDA linear attention. After conversion, the model fixated on the answer position (choosing ‘A’ 81% of the time) rather than following the content, and accuracy fell to 25–29%. Format-targeted KL + SFT + DPO raised C-Eval by +12.48 points. Code and weights released.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

MemArena: A benchmark for personal assistants with memory management on edge devices

The MemArena benchmark tests personal AI assistants with memory management on edge devices using open-weight models. The research involves 50 agents communicating over 15 days. Key findings: the memory backend is more important for accuracy than model size; all systems struggle with permission-aware access…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

TabletCraft: bidirectional neural translation of Akkadian and cuneiform rendering

The article introduces TabletCraft, an open-source system for bidirectional translation between Akkadian and English. The ByT5 model trained on 116K samples achieves 49.1 BLEU for Akkadian→English and 48.5 BLEU for English→Akkadian (the first published result in the reverse direction). The system integrates…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Apple research on controlling outlier tokens in diffusion transformers

The Apple ML Research team identified a problem with outlier tokens in diffusion transformers (DiTs) for image generation — a few high-norm tokens reduce quality. They proposed Dual-Stage Registers (DSR), with training-time and test-time registers. Testing on ImageNet and in text-to-image…

Apple Machine Learning Research Original source ↗