Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
quantization
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
Qwen3.8-27B on one RTX 3090, and the setting its README says chat clients should leave at default
Reno Lu
Reno Lu
Reno Lu
Follow
Oct 2
Qwen3.8-27B on one RTX 3090, and the setting its README says chat clients should leave at default
#
vllm
#
localllm
#
speculativedecoding
#
quantization
Comments
Add Comment
4 min read
1.3 TB of RAM for 100M Vectors? Shrink Them Before They Hit OpenSearch
Jon Handler
Jon Handler
Jon Handler
Follow
Oct 2
1.3 TB of RAM for 100M Vectors? Shrink Them Before They Hit OpenSearch
#
vectorsearch
#
airetrieval
#
opensearch
#
quantization
1
 reaction
Comments
Add Comment
8 min read
Hot Take: Deploy Small Language Models to the Edge
Nainik Mehta
Nainik Mehta
Nainik Mehta
Follow
Sep 23
Hot Take: Deploy Small Language Models to the Edge
#
edge
#
machinelearning
#
mlops
#
quantization
1
 reaction
Comments
1
 comment
4 min read
What Model Quantization Actually Does: From Float16 to 4-Bit Weights
Syed Anzar
Syed Anzar
Syed Anzar
Follow
Sep 18
What Model Quantization Actually Does: From Float16 to 4-Bit Weights
#
ai
#
llm
#
quantization
#
machinelearning
1
 reaction
Comments
3
 comments
6 min read
Testing the 9x Smaller Local LLM Claim on a GPU-less VPS
Qasim Parray
Qasim Parray
Qasim Parray
Follow
Sep 30
Testing the 9x Smaller Local LLM Claim on a GPU-less VPS
#
localllm
#
quantization
#
selfhosting
#
gguf
Comments
1
 comment
6 min read
A Small Transformer Trained in 1.5 Hours Beat Many LLMs on ARC
Anaz S. Aji
Anaz S. Aji
Anaz S. Aji
Follow
for
Codecora Dev
Sep 2
A Small Transformer Trained in 1.5 Hours Beat Many LLMs on ARC
#
machinelearning
#
vectorsearch
#
quantization
#
benchmark
Comments
Add Comment
5 min read
A Better FP4 Gradient Quantizer That Training Couldn't Notice
Seth Wheeler
Seth Wheeler
Seth Wheeler
Follow
Aug 25
A Better FP4 Gradient Quantizer That Training Couldn't Notice
#
llm
#
measurement
#
quantization
#
training
Comments
Add Comment
7 min read
Why Does Your Local Model Crash at 32k Tokens?
Eryk Kubiak
Eryk Kubiak
Eryk Kubiak
Follow
Sep 25
Why Does Your Local Model Crash at 32k Tokens?
#
llm
#
gpu
#
quantization
#
machinelearning
1
 reaction
Comments
2
 comments
12 min read
Error Feedback, Gradient Compression, and Why Adam Breaks It
Seth Wheeler
Seth Wheeler
Seth Wheeler
Follow
Aug 21
Error Feedback, Gradient Compression, and Why Adam Breaks It
#
llm
#
measurement
#
quantization
#
training
5
 reactions
Comments
1
 comment
8 min read
KV Cache INT4 Quantization for 1M+ Token Context Windows
Aomi Qaza
Aomi Qaza
Aomi Qaza
Follow
Aug 15
KV Cache INT4 Quantization for 1M+ Token Context Windows
#
aiengineering
#
quantization
Comments
Add Comment
3 min read
GGUF Quantization: Which Level Should You Use?
Mr Say Nothing
Mr Say Nothing
Mr Say Nothing
Follow
Sep 13
GGUF Quantization: Which Level Should You Use?
#
ai
#
gguf
#
quantization
#
localllm
Comments
1
 comment
5 min read
Comparing INT4 and NVFP4 Palettes on Real Gradient Tensors
Seth Wheeler
Seth Wheeler
Seth Wheeler
Follow
Aug 22
Comparing INT4 and NVFP4 Palettes on Real Gradient Tensors
#
llm
#
quantization
#
training
#
measurement
1
 reaction
Comments
2
 comments
7 min read
Deploying a QAT Checkpoint Your Serving Stack Can't Load: Gemma 4 E2B in Pure JAX on One TPU
xbill
xbill
xbill
Follow
for
Google Developer Experts
Aug 19
Deploying a QAT Checkpoint Your Serving Stack Can't Load: Gemma 4 E2B in Pure JAX on One TPU
#
tpu
#
jax
#
llm
#
quantization
8
 reactions
Comments
Add Comment
10 min read
Unsloth Releases Qwen3.6-27B-NVFP4: Enhanced Throughput and Agentic Coding for Developers
Pneumetron
Pneumetron
Pneumetron
Follow
Jul 17
Unsloth Releases Qwen3.6-27B-NVFP4: Enhanced Throughput and Agentic Coding for Developers
#
aiml
#
largelanguagemodels
#
quantization
#
qwen
Comments
Add Comment
3 min read
Bonsai-27B: A 1-Bit LLM for On-Device Inference with Llama.cpp and MLX
Pneumetron
Pneumetron
Pneumetron
Follow
Jul 15
Bonsai-27B: A 1-Bit LLM for On-Device Inference with Llama.cpp and MLX
#
llm
#
quantization
#
1bit
#
gguf
Comments
Add Comment
3 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account