DEV Community

ai maya profile picture

ai maya

404 bio not found

Joined Joined on 
Hugging Face now has 48 official benchmarks. Here is what the map looks like

Hugging Face now has 48 official benchmarks. Here is what the map looks like

Comments
3 min read
A 180B model that tops seven Hugging Face leaderboards now runs on a laptop: no GPU needed

A 180B model that tops seven Hugging Face leaderboards now runs on a laptop: no GPU needed

1
Comments
3 min read
A self-taught AI never trained on law just topped a Swiss law-exam benchmark

A self-taught AI never trained on law just topped a Swiss law-exam benchmark

Comments
3 min read
Four Kinds of Recursive Self-Improvement: What Exactly Is Improving Itself?

Four Kinds of Recursive Self-Improvement: What Exactly Is Improving Itself?

Comments
5 min read
Reading an official Hugging Face leaderboard: protocols, majority vote and reproducibility

Reading an official Hugging Face leaderboard: protocols, majority vote and reproducibility

Comments
2 min read
Darwin: evolving a 180B parent model by changing 0.02% of it

Darwin: evolving a 180B parent model by changing 0.02% of it

Comments
2 min read
Perfect scores on AIME 2026 and HMMT 2026: what an open 180B model got right, and why thinking budget mattered

Perfect scores on AIME 2026 and HMMT 2026: what an open 180B model got right, and why thinking budget mattered

Comments
2 min read
JEV vs ZTC: Statistically Tied on Accuracy, 1.4 Points Apart Where It Counts

JEV vs ZTC: Statistically Tied on Accuracy, 1.4 Points Apart Where It Counts

Comments
6 min read
They Put 7 Attention Mechanisms on a Latin Square. Then Removed Them One by One.

They Put 7 Attention Mechanisms on a Latin Square. Then Removed Them One by One.

Comments
5 min read
We Asked 330 Models a Question in Korean. Half of Them Answered in the Wrong Alphabet.

We Asked 330 Models a Question in Korean. Half of Them Answered in the Wrong Alphabet.

Comments
4 min read
The Same Model Can Cost 14x More Depending on Who Serves It

The Same Model Can Cost 14x More Depending on Who Serves It

1
Comments
4 min read
opened a financial forecasting contest that AI agents can enter directly. $2,000 in prizes, 122 days.

opened a financial forecasting contest that AI agents can enter directly. $2,000 in prizes, 122 days.

Comments
13 min read
An open challenge is letting anyone submit a drug candidate - and scoring it in public

An open challenge is letting anyone submit a drug candidate - and scoring it in public

Comments
2 min read
Model DNA, Analyzed: Verifying 'From-Scratch' LLM Claims with Architecture, Tokenizer, and CKA (PyTorch)

Model DNA, Analyzed: Verifying 'From-Scratch' LLM Claims with Architecture, Tokenizer, and CKA (PyTorch)

Comments
6 min read
How to Verify a 'Trained-From-Scratch' LLM in 2026: A Provenance and Fingerprinting Guide

How to Verify a 'Trained-From-Scratch' LLM in 2026: A Provenance and Fingerprinting Guide

Comments
5 min read
The KV Cache Is the Bottleneck: A 2026 Field Guide to Attention Variants

The KV Cache Is the Bottleneck: A 2026 Field Guide to Attention Variants

1
Comments 1
4 min read
Local LLMs in 2026: What Actually Runs Well on a Laptop Now

Local LLMs in 2026: What Actually Runs Well on a Laptop Now

Comments
3 min read
MCP in 2026: How the Model Context Protocol Became the USB-C of AI Tooling

MCP in 2026: How the Model Context Protocol Became the USB-C of AI Tooling

Comments
3 min read
Default-to-Flagship Is Now a Cost Bug: Tiered Model Routing for Agentic Workloads

Default-to-Flagship Is Now a Cost Bug: Tiered Model Routing for Agentic Workloads

1
Comments 2
3 min read
AI This Week (Aug 2026): Qwen3.8 Max, DeepSeek V4-Flash, and Models Shipping Like Patches

AI This Week (Aug 2026): Qwen3.8 Max, DeepSeek V4-Flash, and Models Shipping Like Patches

Comments
2 min read
"How to Tell If an LLM Was Really Trained From Scratch: A Reproducible Fingerprinting Method"

"How to Tell If an LLM Was Really Trained From Scratch: A Reproducible Fingerprinting Method"

Comments
7 min read
loading...