DEV Community

Arun Kumar profile picture

Arun Kumar

I write primary-source explainers on transformers and LLM inference — reading the papers and checking what they actually measured. More at diffstudy.com

Joined Joined on 
Fine-tuning a 7B model needs 112 GB. The model is only 14 GB of it.

Fine-tuning a 7B model needs 112 GB. The model is only 14 GB of it.

Comments
3 min read

Want to connect with Arun Kumar?

Create an account to connect with Arun Kumar. You can also sign in below to proceed if you already have an account.

Already have an account? Sign in
If nothing was concatenated, it isn't prompt injection

If nothing was concatenated, it isn't prompt injection

Comments 1
3 min read
MCP doesn't replace function calling. Your MCP client still emits one.

MCP doesn't replace function calling. Your MCP client still emits one.

Comments 2
3 min read
GRPO doesn't remove the reward model. It removes the critic.

GRPO doesn't remove the reward model. It removes the critic.

2
Comments 1
3 min read
Muon doesn't replace AdamW. Every Muon run still has AdamW in it.

Muon doesn't replace AdamW. Every Muon run still has AdamW in it.

Comments
3 min read
"RoPE extrapolates to longer contexts" — someone finally measured it, and it doesn't

"RoPE extrapolates to longer contexts" — someone finally measured it, and it doesn't

Comments 1
3 min read
Speculative decoding won't change your model's distribution. It might still change your output.

Speculative decoding won't change your model's distribution. It might still change your output.

1
Comments 2
4 min read
loading...