Realtime LLM news

Follow model releases, benchmark shifts, and research signals without losing the source.

EvalKit refreshes this feed from free public sources every 15 minutes and falls back to curated citations if a source is temporarily unavailable.

Today

Today

11 items

A Marc Benioff-backed startup thinks AI can solve the AI deployment problem

June emerged from stealth today with a $20 million pre-seed round to make AI adoption simpler.

2026-08-03Presstodaypress

After killer quarter, Palantir CEO Alex Karp calls AI industry ‘Marxist’

After a quarter that delivered $1 billion in profit, Palantir CEO Alex Karp on Monday once again warned that AI frontier labs are too untrustworthy for enterprises.

2026-08-03Presstodaypress

Apple finally fixed Siri. So why does it feel anticlimactic?

Apple’s long-awaited AI overhaul finally makes Siri the assistant it was always supposed to be. Yet it arrives at a moment when simply being a capable AI assistant no longer feels revolutionary.

2026-08-03Presstodaypress

China’s Alibaba takes another swipe at America’s AI supremacy

Chinese tech giant Alibaba released what it says is its largest and "most capable AI model to date," claiming performance rivaling the best systems from US frontier labs Anthropic and OpenAI, as well as domestic rivals like Moonshot AI's K...

2026-08-03Presstodaypressanthropic

Europe’s AI labeling and transparency rules are now in effect

The European Union has ushered in some additional rules that aim to make it easier for people to identify chatbots and AI deepfakes online. The new transparency obligations under the bloc's landmark AI Act came into effect on August 2nd, r...

2026-08-03Presstodaypress

Influencers draw backlash for attending OpenAI’s first luxury trip

OpenAI’s first-ever influencer brand trip is sparking online backlash as tensions over the use of AI continue.

2026-08-03Presstodaypressopenai

Fender’s CEO seems to think your bandmates are just analog AI

Fender CEO Edward "Bud" Cole gave an interview to T3 in May celebrating the 75th anniversary of the Telecaster with comments on AI and music that initially flew under the radar. But it has started making the rounds recently, pouring more f...

2026-08-02Presstodaypress

Sam Altman and AI’s decel debate

On the latest episode of Equity, we discuss why Sam Altman is calling on the industry to "pace the rate of AI development."

2026-08-02Presstodaypress

Is this Billboard Hot 100 hit AI slop?

Fenix Flexin is best known as a member of Shoreline Mafia, a rap duo from Los Angeles. But he's recently found solo success with the track "Rubberz," which has climbed to number 58 on the Billboard Hot 100. Almost immediately, though, ques...

2026-08-01Presstodaypress

YouTuber Hank Green says his AI usage is ‘not healthy’

Green offered a remarkable apology, saying that "the level of dopamine that I've been getting from interacting with LLMs ... is not healthy for me or good for the world."

2026-08-01Presstodaypressllm

Here’s the problem with putting an AI image generator in Google Earth

A text prompt was all it took to generate reality-warping images using Google Earth's satellite, aerial, and 3D imagery with a now-rolled back AI feature, like these images generated by Digital Digging's Henk van Ess that show "refugees ne...

2026-07-31Presstodaypress

New releases

New releases

6 items

Circles powers telco personalization with OpenAI technology

Circles uses the OpenAI API and Codex to power AI-native telco experiences, increasing ARPU by 22%, reducing churn by 9%, and improving development efficiency.

2026-08-03Officialreleaseofficialopenai

How we built a realtime system for responsive voice AI in six months

GPT-Live enables continuous voice interaction with AI, using a turnless speech model and low-latency architecture for faster, more natural conversations.

2026-08-03Officialreleaseofficialgpt

Ten advances in mathematics and theoretical computer science

OpenAI shares new results on long-standing open problems in mathematics and theoretical computer science, including advances in geometry, cryptography, and complexity.

2026-08-01Officialreleaseofficialopenai

Advancing responsible AI across Europe

OpenAI shares how its safety, security, transparency, and provenance practices support responsible AI governance in Europe. The work will continue as the EU AI Act advances.

2026-07-31Officialreleaseofficialopenai

Building abundant intelligence

A full-stack approach to making advanced AI more capable, more affordable, and more widely useful.

2026-07-31Officialreleaseofficial

Disrupting a Criminal Scam Operation

OpenAI disrupted a Cambodia-based scam operation using ChatGPT to support investment, romance, gambling, and impersonation schemes.

2026-07-31Officialreleaseofficialgpt

Research

Research

16 items

Characterizing Bias in Post-Bandit Inference under Index Algorithms

Bandit algorithms generate data for downstream inference, but adaptive sampling biases post-bandit sample means. We analyze this bias for stable index algorithms, including UCB1 and its generalizations, and derive sharp leading-order expre...

2026-08-02Researchresearchinference

Control Under Compression: Reliability Frontiers for Tool-Using Agents

Tool-using language-model agents are governed not only by task prompts but also by persistent system-side instructions that specify tools, arguments, policies, execution protocols, and recovery. Compressing these agent control contexts (AC...

2026-08-02Researchresearchagentmodel

Credit the Right Box: Marginal Contribution Assignment for Structured Visual Perception

Multimodal Large Language Models (MLLMs) are increasingly expected to solve structured perception tasks that require visual recognition, language-to-object binding, object cardinality preservation, and precisely localized grounding and seg...

2026-08-02Researchresearchlanguagelarge

DeBERTa-Sentinel: Toward Transparent and Trustworthy Detection of AI-Generated Text

The rapid spread of large language models (LLMs) across the web raises concerns about misinformation, academic integrity, automated content manipulation, and risks to vulnerable online communities. Existing transformer-based detectors, suc...

2026-08-02Researchresearchlanguagelarge

Decoy Images Amplify Caption-Mediated Defenses Against Encoded Jailbreaks

We report a counter-intuitive interaction between image inputs and existing black-box defenses on Vision--Language Models (VLMs): pairing an encoded jailbreak prompt with an unrelated decoy image can sharply lower attack success rate (ASR)...

2026-08-02Researchresearchlanguagemodel

Don't Offer What Can't Be Done: Deterministic Executability Gating for LLM Skill Selection at Scale

Production LLM agents that select from large skill libraries face a limitation that semantic relevance alone cannot resolve: a skill may match a user's topic yet be impossible to execute in the current account state. We present a deployed...

2026-08-02Researchresearchagentllm

FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds

World models have attracted significant attention for their ability to capture and predict the structure and dynamics of the physical world. In this emerging landscape, Joint Embedding Predictive Architectures (JEPA) offer a particularly c...

2026-08-02Researchresearchagentmodel

Is paying artists enough to convince them to embrace AI?

Illustrators have spent years sounding the alarm about generative artificial intelligence startups training their models on artists' work without permission. They've pointed out how the practice is tantamount to theft, and in response, man...

2026-08-02Pressresearchpressartificial

Logit-Origin Centering for Singleton Test-Time Adaptation

Tabular data is used extensively in many real-world use cases. Deep learning models have been developed to deal with tabular data, but generally perform poorly when the test data distribution differs from that of the training data. Researc...

2026-08-02Researchresearchmodel

On the Limits of Machine-Learned Ranking for Modern Microarchitectural Policies

Machine-learning predictors estimate processor performance far faster than cycle-level simulation. For design-space exploration, however, the valuable test is not merely reproducing the usual hardware ordering, but identifying how differen...

2026-08-02Researchresearch

Role-Decoupled Attention Residuals: Separating Matching and Content Retrieval Across Depth

Depth-routing residual architectures allow Transformer layers to retrieve earlier representations instead of inheriting only the immediately preceding state. Existing Block Attention Residuals, however, use a single content-dependent depth...

2026-08-02Researchresearchevaltransformer

Using Lower-Bound Representations for Trajectory Similarity Learning

Trajectory similarity learning is fundamental to efficient trajectory retrieval under complex distance measures. Existing learning-based methods typically rely on embeddings trained to approximate trajectory distances or rankings, but they...

2026-08-02Researchresearcheval

What Could the Agent See at 19:05? Generating Temporal Enterprise Scenarios from Real Research and Replaying Them to Evaluate Agents

Enterprise AI agents act across many apps whose data changes continuously, so an answer is correct only relative to what data existed and who could see it at the moment it was asked. Offline evaluation today grades against a single static...

2026-08-02Researchresearchagenteval

ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation

Standard AI-text detection benchmarks compare human-written text against text generated directly by large language models (LLMs). While prior work has shown that rewriting and paraphrasing can degrade detector performance, it remains uncle...

2026-07-31Researchresearchbenchmarkeval

Evidence-Type Competition: When Can Interventional Data Teach Language Models Causal Direction?

Interventional data is widely regarded as the gold standard for teaching models causal reasoning. We test this assumption in a fully controlled synthetic environment pitting observational correlation against causal effect, and find it fail...

2026-07-31Researchresearchlanguagemodel

FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models

Reading a social situation often depends on behavior, not words alone. We introduce FriendBench, a benchmark for inferring whether two people are already familiar or are meeting as strangers, from a 20-second clip of a dyadic ice-breaker c...

2026-07-31Researchresearchbenchmarkinference

Benchmarks

Benchmarks

1 item

Design Arena creators raise $7.9 million to bring taste to AI models

Design Arena is used by 5.3 million people around the world, providing critical human evaluations to frontier labs.

2026-08-03Pressbenchmarkpresseval

Resources

Resources

2 items

Congress’ favorite AI tool? ChatGPT

House spending records show OpenAI's ChatGPT dominates paid AI use on Capitol Hill, with congressional offices relying on the chatbot to draft memos, summarize legislation, and assist constituent communications.

2026-08-03Pressresourcepressgpt

Google Earth’s AI deepfake tool only lasted one day

Google has shut down Google Earth feature it launched Thursday that allowed users to edit satellite images with text prompts using AI. The tool essentially let users create AI deepfakes of the real world using text prompts; Digital Digging...

2026-07-31Pressresourcepress