LLM intelligence hub

Track the frontier of LLMs without losing the evidence.

EvalKit brings public leaderboards, model profiles, benchmark guides, and curated LLM news into one citation-first workspace.

Leaderboard rows464
Model profiles410
Public sources8
Verified claims0

Current leaders

Start with the models people are already comparing.

Open full explorer →

LLM news

Curated releases, research, and benchmark shifts.

Read news →

All the AI agents that can live in your text messages

We created a list of the most notable AI agents that can live in your text messages, from general assistants to agents designed for families, travel, and work.

TechCrunch AISource

An OpenAI safety employee has quit and is sounding the alarm

David Robinson used to write the safety reports that accompanied every major model release at OpenAI. This week, he resigned from his position and is now speaking out in an editorial in The Atlantic. It's understandable if you're feeling a...

The Verge AISource

Capcom is preparing for a ‘future where we create games together with AI’

Capcom's Pragmata might be all about the horrors of AI, but in practice the studio doesn't seem so down on the tech. During the Capcom Open Conference RE: 2026 programmer Satoshi Ishida gave a presentation with the mouthful of a title: "Th...

The Verge AISource

OpenAI safety employee resigns, claiming the company’s ‘culture is broken’

By his own admission, David Robinson is “something of a cliché”: an employee at a leading AI company who issues a dire warning while resigning from their job.

TechCrunch AISource

Benchmark guide

Read scores like a product decision, not a scoreboard.

Open guide →

Trust policy

No fake “tested by us” claims.

Public rows are labeled as replicated public-source data or editorial context. “Verified by EvalKit” stays at 0 until there is real run evidence attached.

Read citation policy