DEV Community

#pytorch

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Rebuilding a brain-to-text decoder: from 50% to 23.5% word errors

Rebuilding a brain-to-text decoder: from 50% to 23.5% word errors

1
Comments 1
5 min read
What Does a Neural Network Actually Receive When You Give It an Image?

What Does a Neural Network Actually Receive When You Give It an Image?

Comments
3 min read
Generating electrode stimulation patterns for a simulated cortical prosthesis

Generating electrode stimulation patterns for a simulated cortical prosthesis

5
Comments
5 min read
RoPE vs sinusoidal positional encoding, measured: 55 logits of drift against 0.0005

RoPE vs sinusoidal positional encoding, measured: 55 logits of drift against 0.0005

Comments
7 min read
Demystifying LLMs: Building a 124M-Parameter Decoder-Only Transformer in PyTorch

Demystifying LLMs: Building a 124M-Parameter Decoder-Only Transformer in PyTorch

Comments
24 min read
nn.Module Explained: The Same Model Built with Raw Tensors and with nn.Module

nn.Module Explained: The Same Model Built with Raw Tensors and with nn.Module

Comments
8 min read
25 LLM architecture blocks, side by side, in runnable PyTorch

25 LLM architecture blocks, side by side, in runnable PyTorch

Comments
8 min read
I Rebuilt Jev's Structure with Qwen (Not Its Capabilities)

I Rebuilt Jev's Structure with Qwen (Not Its Capabilities)

Comments
9 min read
Building a Large Language Model from Scratch: A Comprehensive Learning Guide

Building a Large Language Model from Scratch: A Comprehensive Learning Guide

Comments
10 min read
Discriminative Fine-Tuning: Why Your Backbone and Your Head Shouldn't Learn at the Same Speed

Discriminative Fine-Tuning: Why Your Backbone and Your Head Shouldn't Learn at the Same Speed

Comments
6 min read
DeepSpeed ZeRO vs PyTorch FSDP2, Measured: 1.8x on NVLink, Barely Finishing on PCIe (Plus How to Actually Get GPUs on GKE)

DeepSpeed ZeRO vs PyTorch FSDP2, Measured: 1.8x on NVLink, Barely Finishing on PCIe (Plus How to Actually Get GPUs on GKE)

Comments
10 min read
Undefined type Float8_e4m3fn on Apple Silicon: BF16 and GGUF Workarounds for FP8 Models

Undefined type Float8_e4m3fn on Apple Silicon: BF16 and GGUF Workarounds for FP8 Models

Comments 1
4 min read
OpenArch: PyTorch implementations of modern LLM architectures!

OpenArch: PyTorch implementations of modern LLM architectures!

Comments
5 min read
SSRF in PyTorch: How a Missing URL Validation in Dataset Loading Could Leak Cloud Credentials

SSRF in PyTorch: How a Missing URL Validation in Dataset Loading Could Leak Cloud Credentials

Comments
2 min read
Neural Networks: Weights, Activation, and Backpropagation

Neural Networks: Weights, Activation, and Backpropagation

Comments
5 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.