Realtime LLM news

Follow model releases, benchmark shifts, and research signals without losing the source.

EvalKit refreshes this feed from free public sources every 15 minutes and falls back to curated citations if a source is temporarily unavailable.

Today

Today

11 items

Crusoe raises $3.9B to build massive data centers and small modular ‘AI factories’

The round values the data center giant at $30.9 billion.

2026-09-17Presstodaypress

Inside the suddenly explosive world of AI safety

On a sunny July day in Berkeley, California, the country's top AI safety researchers gathered on an unmarked floor of an unmarked building. They had come together for a "war room" to dissect the high-profile cybersecurity incident that had...

2026-09-17Presstodaypress

Is the AI safety debate about safety or control?

Not everyone agrees with Amodei's call for globally coordinated action for AI safety.

2026-09-17Presstodaypress

Microsoft AI CEO says AI threats are real, and Anthropic is making it worse

Today, I’m talking with Mustafa Suleyman, the CEO of Microsoft AI. As you’re no doubt aware, the biggest story in tech right now is the spiraling debate about AI safety and regulation. It should come as no surprise that Mustafa has strong...

2026-09-17Presstodaypressanthropic

Microsoft exec called AI scraping ‘the largest theft of labor in human history,’ new unredacted filings reveal

Newly unsealed court filings show Microsoft privately called OpenAI's data practices "theft" while both companies scraped paywalled Times content, built datasets from it, and warned internally it would gut publishers.

2026-09-17Presstodaypressopenai

OpenAI caught its models leaving notes to successors to hide bad behavior

OpenAI disclosed instances of GPT-5.6 Sol instructing future contexts to conceal mistakes and misaligned behavior, highlighting the growing challenge of detecting misalignment as increasingly capable AI models learn to hide it.

2026-09-17Presstodaypressgpt

PrismML hopes its tiny LLM will change how we all use AI

If AI lab PrismML isn't on your radar yet, it should be.

2026-09-17Presstodaypressllm

The AI Superintelligence Slowdown

Remember when tech leaders would tell their employees to “move fast and break things”? It seemed that would be the way of AI too. But after a summer where rogue AI agents became reality, and researchers warned that AI could kill us all, a...

2026-09-17Presstodaypressagent

The FAA’s plan to fix air traffic? $875M worth of AI

A new AI-based software program is being launched to help air traffic controllers better navigate their jobs as the crossing guards of America's skies.

2026-09-17Presstodaypress

The fix for rogue AI agents could be more AI

As companies hand off longer and more complex tasks to AI agents, they are running into an oversight problem: Agents can act faster, longer, and at greater volume than humans can realistically review.

2026-09-17Presstodaypressagent

UN turns to Google to make its global data ready for AI agents

The shift comes after a UNICEF test found leading AI models struggled to accurately retrieve global development statistics.

2026-09-17Presstodaypressagent

New releases

New releases

7 items

How Cooley is accelerating IPO work with ChatGPT

Cooley built GO Public with ChatGPT Work to bring intelligence to the IPO process, helping lawyers surface issues earlier and focus judgment where it matters most.

2026-09-17Officialreleaseofficialgpt

Introducing Astra for Law

OpenAI for Law brings frontier intelligence for law, custom firm workflows, connected legal data sources, and legal-grade controls for confidential client work.

2026-09-17Officialreleaseofficialopenai

Helping older adults use AI in everyday life

OpenAI and AARP are bringing free, hands-on ChatGPT workshops to 1,000 older adults across 10 U.S. cities to build practical AI skills safely.

2026-09-16Officialreleaseofficialgpt

How to connect AI usage to business value

Learn how ChatGPT Work and Codex analytics help teams understand AI usage and spend, identify training needs, and connect adoption to business outcomes.

2026-09-16Officialreleaseofficialgpt

How workers are unlocking new ways of working

New OpenAI Economic Research shows how workers use AI beyond traditional roles and which new activities become recurring parts of their work.

2026-09-16Officialreleaseofficialopenai

Our framework for reporting model misalignment

OpenAI shares a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of unexpected or concerning model behavior.

2026-09-16Officialreleaseofficialmodel

Reimagining advertising with AI

Explore new AI-powered advertising experiences from OpenAI, including Sponsored Agents, tools for marketers, and integrations with HubSpot and Shopify.

2026-09-16Officialreleaseofficialagent

Research

Research

17 items

AI is feared globally as the destroyer of jobs

Pew Research has published a new global survey that sheds light on how people view AI, including its impact on jobs, life in general, and income inequality. The survey questioned 42,151 people across 37 countries from February 8th to May 1...

2026-09-17Pressresearchpress

A Zeroth-Order Paradigm for LLM Preference Alignment

Direct preference alignment methods are widely used to align large language models (LLMs) with human preferences because of their computational and memory efficiency. However, likelihood displacement motivates alternative ways to extract i...

2026-09-16Researchresearchlanguagelarge

Affora: A Design System for Agent-Friendly Interfaces

Computer-use agents increasingly operate software designed for people, but interfaces often leave actions or task state unclear to machine readers. We present Affora, a design system that supports both readers while preserving visual freed...

2026-09-16Researchresearchagent

Cognitive Extensions for Dual-Process Language Agents: Memory and Self-Reflection in Interactive Environments

Language agents remain brittle in interactive environments, where success requires long-horizon state tracking, valid action execution, and recovery from failed steps. We extend SwiftSage, a dual-process agent that combines a fast action p...

2026-09-16Researchresearchagent

Dreaming the Sound of Contact: Leveraging Video and Audio Generation for Zero-Shot Force-Aware Manipulation and Data Generation

Recent advances in video generation allow robots to learn manipulation trajectories from generated videos. However, these approaches produce purely kinematic trajectories that lack force information, causing failures in contact-rich tasks...

2026-09-16Researchresearch

Evidence-Grounded Agentic Formulation Development in an Autonomous Laboratory

Self-emulsifying drug delivery systems (SEDDS) can improve the oral bioavailability of poorly soluble drugs, but identifying high-performing formulations remains experimentally intensive. We present Andromeda 2, an agentic system that reas...

2026-09-16Researchresearchagent

Exponential Hardness of Off-Policy Evaluation under History-Dependent Logging

Can a logged dataset visit every hidden state frequently and still be exponentially uninformative about a target policy's value? We show that it can when the logger depends on history. For every horizon $H \ge 3$, we construct two POMDPs w...

2026-09-16Researchresearcheval

Flag Game: A Toy Model for Mechanistic Swarm Interpretability

Emergent coordinated behaviors of AI agents are starting to present critical safety risks. A key phenomenon driving these behaviors is the rapid formation and spread of beliefs about the world, and mechanistic understanding is crucial for...

2026-09-16Researchresearchagentmodel

How Model Growth, Recursion, and Boundary Operators Influence Scaling Exponents

Scaling laws predict how loss decreases with increases in computation. We show, contrary to conventional wisdom, that architectural interventions can modify scaling exponents in pre-training, leading to exponential improvements in performa...

2026-09-16Researchresearchmodel

Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations

As models scale, reward hacking becomes more frequent, more sophisticated, and more consequential. Does it leave a telltale signature in model representations? This work analyzes how reward hacking is represented internally in frontier ope...

2026-09-16Researchresearchevalllm

MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education

Large vision-language models have achieved remarkable progress in multi-modal understanding, yet their capabilities in educational settings remain insufficiently evaluated. In AI-assisted language learning, models must interpret artistic i...

2026-09-16Researchresearchbenchmarkeval

Objective vs. Search: Decomposing What Makes a Good Tokeniser

Two dominant tokenisation algorithms are used by modern language models: byte-pair encoding (BPE) and UnigramLM. These differ along two orthogonal axes: their optimisation objective (compression vs. log-likelihood) and their search procedu...

2026-09-16Researchresearchlanguagemodel

PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection

Intelligent systems that act in the world require image understanding that is both comprehensive and spatially grounded. Current vision-language models (VLMs) can generate fluent and detailed image captions, but reliably associating them w...

2026-09-16Researchresearchlanguagemodel

Playing log(N)-Questions over Wikipedia Abstracts: Communication Efficiency Between Paired Frontier Models

We evaluate six frontier language models on the two-agent $\log(N)$-Questions game. A questioner sees $N$ Wikipedia lead paragraphs and must identify a secretly chosen target using exactly $\log_2 N$ yes/no questions. An answerer sees only...

2026-09-16Researchresearchagenteval

Prepared Or Unprepared? Evaluating Healthcare Workforce Readiness for Clinical Adoption of Artificial Intelligence in Nigeria

Artificial intelligence (AI) is increasingly integrated into healthcare systems worldwide, yet its successful clinical adoption depends critically on workforce readiness, particularly in low- and middle-income countries (LMICs) where infra...

2026-09-16Researchresearchartificialeval

Probabilistic Linear Explanations

Formal explainability provides mathematically grounded justifications for individual predictions. However, abductive explanations often exceed human cognitive limits by involving too many features, while probabilistic relaxations have rema...

2026-09-16Researchresearch

Reporting Practice Matters: The Impact of Reference Choice on Chest X-ray Report Evaluation

Radiologists follow heterogeneous reporting practices. Two radiologists examining the same image and identifying the same clinical findings might nevertheless compose superficially distinct reports, varying in terminology, shorthand, forma...

2026-09-16Researchresearcheval

Resources

Resources

1 item

Claude Code relaunches Projects to manage multiple AI agents in the cloud

The revamped projects feature in Claude Code allows users to run multiple agents under the same roof, with a shared memory, goals, and library of files and artifacts. Similar to Grok Bot and other tools that manage groups of AI agents, eac...

2026-09-17Pressresourcepressagent