DEV Community

Yuri Pocepaev profile picture

Yuri Pocepaev

Software engineer at Neuroprem. LLM quantization, efficient inference, retrieval and agent memory. Practical benchmarks and runtime fixes. Less hype. More engineering.

Work

Software engineer and lead developer at Neuroprem

When Offloading Stalls Inference: Queue Capacity, CUDA Event Ownership, and Misleading Health Checks

When Offloading Stalls Inference: Queue Capacity, CUDA Event Ownership, and Misleading Health Checks

Comments
9 min read
One Model, Many Roles: Specializing LLM Inference Without Training More Models

One Model, Many Roles: Specializing LLM Inference Without Training More Models

Comments
12 min read
Prevent Context Leakage in Multi-Tenant LLM Systems

Prevent Context Leakage in Multi-Tenant LLM Systems

Comments
6 min read
Smaller Context, Recoverable History: Inside an Agent Memory Handoff

Smaller Context, Recoverable History: Inside an Agent Memory Handoff

Comments 2
8 min read
How I Took GLiClass FP8 from 59 ms to 16 ms on an RTX 4050

How I Took GLiClass FP8 from 59 ms to 16 ms on an RTX 4050

Comments
8 min read
Building an NVFP4 KV Cache for a Hybrid Qwen Model

Building an NVFP4 KV Cache for a Hybrid Qwen Model

1
Comments
10 min read
Running Qwen Flash-Next NVFP4 in vLLM: PLE Loading, B12x Fixes, and Stable Inference

Running Qwen Flash-Next NVFP4 in vLLM: PLE Loading, B12x Fixes, and Stable Inference

Comments 2
13 min read
loading...