DEV Community

Cover image for Python 3.13 Free-Threading: What Actually Happens When the GIL Is Removed
Syed Anzar
Syed Anzar

Posted on

Python 3.13 Free-Threading: What Actually Happens When the GIL Is Removed

For over thirty years, every Python developer learned the same hard rule: CPU-bound code in standard Python cannot scale across multiple CPU cores using threading. The Global Interpreter Lock (GIL) forced every thread to take turns executing bytecode, serialization was inevitable, and multi-core speedups required multiprocessing with heavy IPC serialization overhead.

Starting in Python 3.13, CPython introduced an experimental free-threaded build (python3.13t via PEP 703) where the GIL is completely disabled.

Turning off the GIL was not just removing a single lock. The GIL served as the safety blanket protecting Python's memory layout, reference counters, and internal dicts from data races. Removing it required re-engineering the foundational memory architecture of CPython.

Here is what actually happens under the hood when the GIL is removed.


1. The Core Problem: Why the GIL Existed in the First Place

Every Python object contains a header with two fundamental fields: a pointer to its type object (ob_type) and a reference counter (ob_refcnt).

In standard GIL builds, whenever you pass a variable or read a list element, CPython increments or decrements ob_refcnt using simple non-atomic CPU instructions:

// Standard CPython non-atomic refcount macro
#define Py_INCREF(op) (((PyObject*)(op))->ob_refcnt++)
#define Py_DECREF(op) \
    if (--((PyObject*)(op))->ob_refcnt == 0) _Py_Dealloc((PyObject*)(op))
Enter fullscreen mode Exit fullscreen mode

Non-atomic instructions (++ and --) take multiple CPU cycles: load value into register, increment register, store back to memory. If two OS threads modify that counter simultaneously without synchronization, you hit race conditions, missed decrements, double frees, and memory corruption.

The GIL avoided this by guaranteeing that only one thread touched ob_refcnt at any given moment.

Simply replacing every refcount operation with atomic assembly instructions (LOCK XADD on x86) was tried in earlier experimental forks (like Larry Hastings' Gilectomy). The result: single-threaded Python code slowed down by 30% to 50% because atomic bus-locking instructions invalidate CPU L1/L2 caches and choke CPU pipelines.

To make free-threading viable, PEP 703 implemented Biased Reference Counting (BRC).


2. Biased Reference Counting (BRC): Fast-Path vs Shared-Path

Biased Reference Counting is built on a simple observation: in real-world programs, most objects are created and manipulated exclusively by the single thread that allocated them.

Instead of one global reference count, free-threaded CPython splits reference counting into two distinct paths:

  1. Local Refcount (ob_ref_local): Managed exclusively by the thread that owns (created) the object.
  2. Shared Refcount (ob_ref_shared): Managed atomically by any other thread accessing the object.
+-------------------------------------------------------------+
|                     Python Object Header                    |
+-------------------------------------------------------------+
|  ob_tid (Thread ID of Owner)                                |
|  ob_ref_local   (16-bit: Non-atomic fast updates by owner)  |
|  ob_ref_shared  (32-bit: Atomic atomic updates by peers)   |
|  ob_type        (Pointer to Type descriptor)                |
+-------------------------------------------------------------+
Enter fullscreen mode Exit fullscreen mode

The Fast Path (Owner Thread)

When the owning thread increments or decrements an object's reference counter, it checks if current_thread_id == object->ob_tid. If true, it performs a plain, non-atomic addition to ob_ref_local. This costs zero bus locks and keeps single-thread execution fast.

The Slow Path (Foreign Threads)

When another thread references that same object, it cannot touch ob_ref_local. Instead, it issues an atomic operation on ob_ref_shared.

Merging and Deallocation

When does the object actually get freed?
The owning thread periodically merges ob_ref_shared into ob_ref_local. When ob_ref_local + (ob_ref_shared >> 2) == 0, the owning thread triggers deallocation. If a non-owning thread drops the final reference while the owning thread is idle, an evaluation interrupt (eval_breaker) signals the owning thread to process the deferred cleanup safely.


3. Immortal Objects: Stopping Cache-Line Bouncing

Some objects are read constantly by every single running thread: None, True, False, small integers (0 to 256), built-in exceptions, type definitions (int, str, list), and top-level function code objects.

If 16 parallel threads constantly run Py_INCREF(Py_None) and Py_DECREF(Py_None), atomic cache-line invalidation would destroy multi-core performance. CPU cores would spend more time invalidating each other's L1 cache lines than executing code.

CPython solves this with Immortalization (PEP 683, adapted for PEP 703).

In free-threaded builds, immortal objects have their refcount bitfield set to a special constant (_Py_IMMORTAL_REFCNT). The refcount macros check this bitfield:

static inline void _Py_INCREF_STAT(PyObject *op) {
    if (_Py_IsImmortal(op)) {
        return; // No-op: Zero memory writes, zero cache invalidation
    }
    // Proceed to biased reference counting...
}
Enter fullscreen mode Exit fullscreen mode

Because immortal objects never modify memory during refcounting, thousands of threads can read None, interned strings, or module definitions concurrently with zero memory contention.


4. Goodbye pymalloc, Hello mimalloc

Standard CPython relies on pymalloc, an allocator optimized for small Python objects under 512 bytes. However, pymalloc is completely non-thread-safe and relies entirely on the GIL for synchronization.

Rather than bolting locks onto pymalloc (which would kill allocation speed), free-threaded Python 3.13 replaces it with Microsoft's mimalloc allocator.

Mimalloc brings critical architectural advantages:

  • Thread-Local Heaps: Each thread maintains its own allocation arenas. Allocating a new dictionary or string requires zero global locks.
  • Radix-Tree Page Lookups: Enables the Garbage Collector to find and traverse all Python objects directly through heap metadata, eliminating the need to maintain global doubly-linked lists across all objects.
  • QSBR (Quiescent-State Based Reclamation): Prevents use-after-free bugs when memory pages are returned to the operating system while other threads are finishing up reads.

5. Per-Object Critical Sections: Replacing the Global Lock

Without the GIL, what prevents two threads from mutating the same list or dict simultaneously and corrupting internal hash tables?

CPython 3.13 introduces Per-Object Critical Sections (Py_BEGIN_CRITICAL_SECTION).

Instead of locking the entire runtime, Python locks only the specific container being modified:

// Example: Modifying a list in free-threaded CPython
Py_BEGIN_CRITICAL_SECTION(list_obj);
// Exclusive lock on list_obj held here
list_append_internal(list_obj, item);
Py_END_CRITICAL_SECTION();
Enter fullscreen mode Exit fullscreen mode

Automatic Deadlock Avoidance

What if Thread A locks List 1 and wants List 2, while Thread B locks List 2 and wants List 1?

CPython critical sections maintain a thread-local stack of active locks. If a thread attempts to acquire an object lock that is currently held by another thread, the runtime automatically suspends all active critical sections held by the waiting thread, releases their locks, yields, and re-acquires them in a strict ordered hierarchy before continuing.


6. What This Means for Python Developers

Running Python 3.13 free-threaded (python3.13t) delivers real CPU-bound concurrency:

import sys
from concurrent.futures import ThreadPoolExecutor

print(f"Free threading enabled: {sys._is_gil_enabled() == False}")

def cpu_heavy_task(n):
    return sum(i * i for i in range(n))

# In Python 3.13t, this runs in true parallel across all 8 CPU cores
with ThreadPoolExecutor(max_workers=8) as executor:
    results = list(executor.map(cpu_heavy_task, [10_000_000] * 8))
Enter fullscreen mode Exit fullscreen mode

Key Takeaways

  1. True Parallel Threads: CPU-bound tasks in standard Python threads now scale linearly with CPU core counts without multiprocessing IPC overhead.
  2. Thread Safety Still Matters: Built-in containers are safe from memory crashes due to internal critical sections, but race conditions in application logic still require explicit locks (threading.Lock).
  3. C-Extensions Need Adaptation: C extensions must be compiled specifically for the free-threaded ABI (python3.13t) to verify they don't rely on the legacy GIL for internal struct safety.

Python's transition to free-threading is one of the most substantial engineering feats in modern language runtimes, unlocking decades of untapped multi-core performance while maintaining the syntax and ergonomics developers rely on.

Top comments (0)