Skip to main content

Command Palette

Search for a command to run...

The OOM crash that forced a model redesign

Updated
•3 min read•View as Markdown
D
Tech Lead–oriented Senior Full Stack Engineer with 7+ years of experience designing scalable, multi-tenant, and cloud-native systems across enterprise SaaS, e-learning, government, and enterprise platforms. Led architectural transformations spanning AI-driven systems, media streaming infrastructure, CI/CD modernization, multilingual content architecture, analytics, security, and distributed systems. Strong expertise in AWS-based cloud architecture, system design, microservices, API development, database architecture, performance optimization, and production reliability. Experienced in driving end-to-end feature delivery, mentoring engineers, collaborating with cross-functional stakeholders, and translating complex business requirements into production-grade technical solutions. Focused on driving architectural excellence and building scalable systems in high-impact engineering environments.

Early in this project, image generation was a single hardcoded call: one model, one code path. It worked, until I pointed it at a heavier anime-style checkpoint and the process died mid-run with an out-of-memory crash on Apple Silicon's MPS backend. That crash is the reason the pipeline now supports multiple interchangeable image engines instead of one.

The problem

The first image model was small enough to load comfortably. The moment I swapped in a larger, better-quality checkpoint for a more detailed anime style, generation crashed partway through a batch. It wasn't the first image, which made it worse to debug, but several images into a run, once accumulated memory pressure tipped it over.

The naive fix, "just use the smaller model", wasn't really a fix. It was giving up on quality to work around a memory ceiling. The real problem was that the code had no concept of "model that doesn't fit" as a condition to design around; it assumed one fixed model would always work.

The design decision

Two changes, in order:

  1. Move image generation to a bigger, more memory-efficient architecture for the higher-quality checkpoints, rather than fighting the smaller architecture's memory profile.

  2. Make the model a per-scene, per-episode choice instead of a hardcoded constant. Some scenes needed a specific model that rendered a particular subject (a certain character, or dog scenes with a model that handled animals noticeably better). Baking one model into the code meant every future "actually, this one scene needs a different model" request meant another crash-and-patch cycle.

That second decision is the one with lasting value. It turned a one-off bug fix into an engine-selection system: pick a default engine per episode, override per-scene when a specific model genuinely draws something better, and fall back gracefully across CUDA, Apple Silicon, and CPU with appropriate precision for each.

Before / after (illustrative)

Before: one model, hardcoded, no fallback

pipe = load_pipeline("anime-checkpoint-v1", dtype=torch.float16)

def generate_image(prompt):
    return pipe(prompt).images[0]

Great until that checkpoint doesn't fit in memory, or until scene 7 out of 12 in an episode actually needs a different model. Every such case meant editing this function directly.

After: device-aware loading with per-scene model resolution

def get_device_and_dtype():
    if torch.cuda.is_available():
        return "cuda", torch.float16
    if torch.backends.mps.is_available():
        return "mps", torch.float16
    return "cpu", torch.float32

_loaded_pipelines = {}

def get_pipe(style="anime", model_id=None):
    key = model_id or style
    if key not in _loaded_pipelines:
        device, dtype = get_device_and_dtype()
        target_model = model_id or ENGINE_MODELS[style]
        pipe = load_pipeline(target_model, dtype=dtype).to(device)
        _loaded_pipelines[key] = pipe
    return _loaded_pipelines[key]

def generate_image(scene, default_style):
    pipe = get_pipe(style=default_style, model_id=scene.get("image_model"))
    return pipe(scene["image_prompt"]).images[0]

Now a crash on one model doesn't mean rewriting the calling code. It means adding an entry to a model map and, if needed, an image_model override on the one scene that needs it.

How a scene picks its model

Why it mattered

The crash forced a question I should have asked from the start: what happens when the "one model" assumption breaks? Once image selection became data (a field on the episode or scene) instead of code, every subsequent model experiment, whether a more realistic engine or a model better suited to a specific character or subject, was a config change, not a redesign. The OOM crash cost an afternoon; the fix it produced saved every model swap after it.

A

Making image_model scene data removes the hardcoded-engine assumption. The illustrated _loaded_pipelines dictionary introduces a separate memory risk, though: every newly selected checkpoint remains resident, so an episode using several overrides can recreate the original OOM even when each model fits on its own.

I would give the cache an explicit residency budget and define when an inactive pipeline can be evicted. An active generation needs a lease so its pipeline cannot disappear halfway through a scene. A sequence alternating three checkpoints would be a useful regression case, recording resident memory after each switch rather than testing only a fresh process with one model.