ML & Research
A new open dataset named LAION-BVD has been released, providing 10 million hours of video content for multimodal pre-training. The dataset, compiled from 80 million downloaded videos, aims to increase accessibility for researchers in video, audio, and image modalities.
Aug 26, 2026 · 2 min read
8 connections in the Atlas
ML & Research
Researchers at MIT have introduced CrysVCD, an AI-driven framework that integrates fundamental chemical rules into the initial stages of material design. This approach significantly increases the proportion of stable, usable materials generated by AI models, reducing the time and computational resources typically spent on screening unstable designs.
Aug 26, 2026 · 2 min read
2 connections in the Atlas
ML & Research
A new study on logic puzzles found that readily available AI assistance can hinder long-term human skill development, despite improving short-term performance. Participants who frequently used AI assistance performed worse on tasks after the assistance was removed, especially when the cost to access AI was low.
Aug 25, 2026 · 3 min read
3 connections in the Atlas
ML & Research
Researchers have introduced SWE Refactor Bench, a new benchmark designed to assess the capability of coding agents to perform long-horizon, whole-repository stack migrations. The benchmark addresses limitations in existing evaluations by verifying both the completion of the migration and the behavioral correctness of the updated code, rather than solely relying on test pass rates.
Aug 25, 2026 · 2 min read
6 connections in the Atlas
ML & Research
A new benchmark called ConceptGuard evaluates how well Large Language Models (LLMs) can selectively remove harmful concepts without affecting useful knowledge. Current evaluation methods are insufficient, often using separate sets of facts that do not capture the nuance of true unlearning. ConceptGuard aims to provide a more accurate assessment by focusing on conceptual removal.
Aug 21, 2026 · 2 min read
7 connections in the Atlas
ML & Research
A new framework, Pandora's AI Model Routing Box, addresses the challenge of efficiently allocating queries across diverse AI models, considering the inherent cost of estimating each model's potential value. This approach formalizes the trade-off between inexpensive, noisy estimators and accurate, costly ones, aiming to optimize overall system performance and efficiency.
Aug 21, 2026 · 2 min read
3 connections in the Atlas
ML & Research
Researchers from Navers Lab and Einsia.AI Tsinghua University have released AI4AI-Bench, a new benchmark designed to assess large language model agents' capacity to design training algorithms. This benchmark specifically targets the feasibility of recursive self-improvement in AI systems by focusing on algorithmic innovation rather than data collection or hyperparameter tuning.
Aug 21, 2026 · 2 min read
6 connections in the Atlas
ML & Research
Researchers directly measured the influence of a single training example on a GPT-2 model by conducting 24 counterfactual pre-training runs. The experiment involved training 32 GPT-2 models with 124 million parameters each on the OpenWebText dataset to quantify how one altered data point affects the final model.
Aug 20, 2026 · 3 min read
5 connections in the Atlas
ML & Research
Current artificial intelligence benchmarks focus on capability, measuring average output rather than consistency, according to new research. The study proposes that "precision," or the tightness of output concentration around a target across repeated requests, is the true frontier metric for distinguishing AI systems.
Aug 20, 2026 · 2 min read
5 connections in the Atlas
ML & Research
Researchers have introduced ADEPT, a reinforcement learning framework designed to improve the dexterity of high degree-of-freedom robots. The system pre-trains robotic policies on generic manipulation tasks in simulation, then uses this learned behavior as a foundation for more complex, contact-rich tasks in real-world environments.
Aug 20, 2026 · 2 min read
4 connections in the Atlas
ML & Research
A new paper on arXiv details a method for distributing large language model inference across multiple Intel AI PCs, enabling these devices to collectively run models that exceed the memory capacity of any single machine. This technique, called pipelined sharding, leverages OpenVINO to pre-compile model layers into optimized graphs for efficient, distributed execution.
Aug 20, 2026 · 2 min read
5 connections in the Atlas
ML & Research
Researchers have introduced SPADE, a self-play reinforcement learning framework where a single large language model functions as both an environment designer and a reasoning agent. This approach allows the LLM to generate executable, long-horizon training environments and then learn to act within them, fostering continuous self-improvement.
Aug 20, 2026 · 2 min read
5 connections in the Atlas