The OOM crash that forced a model redesign
Early in this project, image generation was a single hardcoded call: one model, one code path. It worked, until I pointed it at a heavier anime-style checkpoint and the process died mid-run with an out-of-memory crash on Apple Silicon's MPS backend. That crash is the reason the pipeline now supports multiple interchangeable image engines instead of one.
The problem
The first image model was small enough to load comfortably. The moment I swapped in a larger, better-quality checkpoint for a more detailed anime style, generation crashed partway through a batch. It wasn't the first image, which made it worse to debug, but several images into a run, once accumulated memory pressure tipped it over.
The naive fix, "just use the smaller model", wasn't really a fix. It was giving up on quality to work around a memory ceiling. The real problem was that the code had no concept of "model that doesn't fit" as a condition to design around; it assumed one fixed model would always work.
The design decision
Two changes, in order:
Move image generation to a bigger, more memory-efficient architecture for the higher-quality checkpoints, rather than fighting the smaller architecture's memory profile.
Make the model a per-scene, per-episode choice instead of a hardcoded constant. Some scenes needed a specific model that rendered a particular subject (a certain character, or dog scenes with a model that handled animals noticeably better). Baking one model into the code meant every future "actually, this one scene needs a different model" request meant another crash-and-patch cycle.
That second decision is the one with lasting value. It turned a one-off bug fix into an engine-selection system: pick a default engine per episode, override per-scene when a specific model genuinely draws something better, and fall back gracefully across CUDA, Apple Silicon, and CPU with appropriate precision for each.
Before / after (illustrative)
Before: one model, hardcoded, no fallback
pipe = load_pipeline("anime-checkpoint-v1", dtype=torch.float16)
def generate_image(prompt):
return pipe(prompt).images[0]
Great until that checkpoint doesn't fit in memory, or until scene 7 out of 12 in an episode actually needs a different model. Every such case meant editing this function directly.
After: device-aware loading with per-scene model resolution
def get_device_and_dtype():
if torch.cuda.is_available():
return "cuda", torch.float16
if torch.backends.mps.is_available():
return "mps", torch.float16
return "cpu", torch.float32
_loaded_pipelines = {}
def get_pipe(style="anime", model_id=None):
key = model_id or style
if key not in _loaded_pipelines:
device, dtype = get_device_and_dtype()
target_model = model_id or ENGINE_MODELS[style]
pipe = load_pipeline(target_model, dtype=dtype).to(device)
_loaded_pipelines[key] = pipe
return _loaded_pipelines[key]
def generate_image(scene, default_style):
pipe = get_pipe(style=default_style, model_id=scene.get("image_model"))
return pipe(scene["image_prompt"]).images[0]
Now a crash on one model doesn't mean rewriting the calling code. It means adding an entry to a model map and, if needed, an image_model override on the one scene that needs it.
How a scene picks its model
Why it mattered
The crash forced a question I should have asked from the start: what happens when the "one model" assumption breaks? Once image selection became data (a field on the episode or scene) instead of code, every subsequent model experiment, whether a more realistic engine or a model better suited to a specific character or subject, was a config change, not a redesign. The OOM crash cost an afternoon; the fix it produced saved every model swap after it.

