DEV Community

#quantization

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Qwen3.8-27B on one RTX 3090, and the setting its README says chat clients should leave at default

Qwen3.8-27B on one RTX 3090, and the setting its README says chat clients should leave at default

Comments
4 min read
1.3 TB of RAM for 100M Vectors? Shrink Them Before They Hit OpenSearch

1.3 TB of RAM for 100M Vectors? Shrink Them Before They Hit OpenSearch

1
Comments
8 min read
Hot Take: Deploy Small Language Models to the Edge

Hot Take: Deploy Small Language Models to the Edge

1
Comments 1
4 min read
What Model Quantization Actually Does: From Float16 to 4-Bit Weights

What Model Quantization Actually Does: From Float16 to 4-Bit Weights

1
Comments 3
6 min read
Testing the 9x Smaller Local LLM Claim on a GPU-less VPS

Testing the 9x Smaller Local LLM Claim on a GPU-less VPS

Comments 1
6 min read
A Small Transformer Trained in 1.5 Hours Beat Many LLMs on ARC

A Small Transformer Trained in 1.5 Hours Beat Many LLMs on ARC

Comments
5 min read
A Better FP4 Gradient Quantizer That Training Couldn't Notice

A Better FP4 Gradient Quantizer That Training Couldn't Notice

Comments
7 min read
Why Does Your Local Model Crash at 32k Tokens?

Why Does Your Local Model Crash at 32k Tokens?

1
Comments 2
12 min read
Error Feedback, Gradient Compression, and Why Adam Breaks It

Error Feedback, Gradient Compression, and Why Adam Breaks It

5
Comments 1
8 min read
KV Cache INT4 Quantization for 1M+ Token Context Windows

KV Cache INT4 Quantization for 1M+ Token Context Windows

Comments
3 min read
GGUF Quantization: Which Level Should You Use?

GGUF Quantization: Which Level Should You Use?

Comments 1
5 min read
Comparing INT4 and NVFP4 Palettes on Real Gradient Tensors

Comparing INT4 and NVFP4 Palettes on Real Gradient Tensors

1
Comments 2
7 min read
Deploying a QAT Checkpoint Your Serving Stack Can't Load: Gemma 4 E2B in Pure JAX on One TPU

Deploying a QAT Checkpoint Your Serving Stack Can't Load: Gemma 4 E2B in Pure JAX on One TPU

8
Comments
10 min read
Unsloth Releases Qwen3.6-27B-NVFP4: Enhanced Throughput and Agentic Coding for Developers

Unsloth Releases Qwen3.6-27B-NVFP4: Enhanced Throughput and Agentic Coding for Developers

Comments
3 min read
Bonsai-27B: A 1-Bit LLM for On-Device Inference with Llama.cpp and MLX

Bonsai-27B: A 1-Bit LLM for On-Device Inference with Llama.cpp and MLX

Comments
3 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.