DEV Community

#localllm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Run Qwen3.8-Flash-Next Locally with EXL3 and TabbyAPI on 40+ GB of RAM

Run Qwen3.8-Flash-Next Locally with EXL3 and TabbyAPI on 40+ GB of RAM

Comments
10 min read
Qwen3.8-27B on one RTX 3090, and the setting its README says chat clients should leave at default

Qwen3.8-27B on one RTX 3090, and the setting its README says chat clients should leave at default

Comments
4 min read
Still carrying the diagnosis

Still carrying the diagnosis

Comments
8 min read
Claude CLI 401 Unauthorized Refresh Token Issue

Claude CLI 401 Unauthorized Refresh Token Issue

Comments
3 min read
Discord bot not responding but still running: how to catch it

Discord bot not responding but still running: how to catch it

Comments
3 min read
How Much RAM Do You Need for Local LLMs on a Mac?

How Much RAM Do You Need for Local LLMs on a Mac?

Comments
3 min read
Claude Code MCP setup and usage rules

Claude Code MCP setup and usage rules

Comments
3 min read
Clone your own voice with a 17-second recording

Clone your own voice with a 17-second recording

Comments
2 min read
Prompt engineering: How to fix instructions that AI ignores

Prompt engineering: How to fix instructions that AI ignores

Comments
3 min read
Inside My llama.cpp Setup: Tuning Qwen 3.8 27B for 512K Context

Inside My llama.cpp Setup: Tuning Qwen 3.8 27B for 512K Context

Comments 2
7 min read
What Happens When You Ask an LLM a Question

What Happens When You Ask an LLM a Question

Comments
8 min read
Run vLLM on Kubernetes with Minikube, WSL2 and NVIDIA GPU

Run vLLM on Kubernetes with Minikube, WSL2 and NVIDIA GPU

Comments
13 min read
Testing the 9x Smaller Local LLM Claim on a GPU-less VPS

Testing the 9x Smaller Local LLM Claim on a GPU-less VPS

Comments 1
6 min read
I Ran DeepSeek V4 Flash Across Two DGX Sparks Over Ethernet

I Ran DeepSeek V4 Flash Across Two DGX Sparks Over Ethernet

Comments
11 min read
VRAM and RAM for local LLMs — honest planning bands, not a GPU tier list

VRAM and RAM for local LLMs — honest planning bands, not a GPU tier list

Comments
4 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.