Skip to main content

Command Palette

Search for a command to run...

Designing a Data Model for Translatable Content Without Breaking Every Existing Record

Updated
•5 min read•View as Markdown
Designing a Data Model for Translatable Content Without Breaking Every Existing Record
D
Tech Lead–oriented Senior Full Stack Engineer with 7+ years of experience designing scalable, multi-tenant, and cloud-native systems across enterprise SaaS, e-learning, government, and enterprise platforms. Led architectural transformations spanning AI-driven systems, media streaming infrastructure, CI/CD modernization, multilingual content architecture, analytics, security, and distributed systems. Strong expertise in AWS-based cloud architecture, system design, microservices, API development, database architecture, performance optimization, and production reliability. Experienced in driving end-to-end feature delivery, mentoring engineers, collaborating with cross-functional stakeholders, and translating complex business requirements into production-grade technical solutions. Focused on driving architectural excellence and building scalable systems in high-impact engineering environments.

Adding multi-language support to a content platform sounds like a straightforward data modeling problem until you remember one thing: you already have production data in the old shape, and you can't just migrate it all at once without downtime, data loss risk, or both. Here's how we designed a translatable-content model that could hold both the old shape and the new shape at once, and the change-detection bug that came with it.

The type that let old and new data coexist

The starting point was simple: every text field on a piece of content was just a string. Adding translations meant a field needed to potentially hold content in multiple languages at once. The tempting-but-wrong move is to migrate every field to a language-keyed object in one big migration. The safer move — the one we took — was a union type:

type TranslatableField = string | LanguageObject;

interface LanguageObject {
  en: string;
  es?: string;
  fr?: string;
  // ...other language codes
  _meta?: Record<string, { translatedAt: string; source: 'human' | 'machine' }>;
}

Every existing record with a plain string stays valid, untouched, forever if needed. Every new piece of translated content gets the richer object shape. No big-bang migration, no downtime, no risk of a half-migrated dataset.

The cost of that flexibility is that you can't read these fields directly anymore — every call site that used to do block.title now has to handle both shapes:

function getLanguageContent(field: TranslatableField, lang: string, fallback = 'en'): string {
  if (typeof field === 'string') return field;         // legacy record, no translations yet
  return field[lang] ?? field[fallback] ?? '';
}

We centralized that logic into a small set of helper functions and made it a hard rule: nothing reads a translatable field directly, everything goes through the helper. That rule is the entire reason the union type is livable — without it, every component and every API response handler would need to duplicate the "is this a string or an object" check, and someone would eventually forget.

Here's the shape of that type, and how a legacy record and a new record both satisfy it:

The "base language" concept, and why editing needed a lock

Once content can exist in multiple languages, you run into a subtler problem: what happens when someone edits a piece of content that has already been translated? If they edit the English version, does the Spanish translation silently go stale? Does the UI let them edit the Spanish version directly and diverge from what a translation pipeline would produce?

We settled on a "base language" model: one language is the source of truth for a given piece of content, and non-base-language editing is intentionally restricted — direct structural edits (reordering answer options, changing what's "correct," restructuring content) are locked on non-base languages, because those changes need to flow from the base language outward through translation, not be made independently in each language and drift apart. Only wording, not structure, is safe to hand-edit per language.

This is as much a UX decision as a data model one, but it only works because the data model can represent "language X is derived from language Y" cleanly — you can't build sane rules around a mush of independent strings.

The bug: detecting "does this need re-translation" by hashing content

Here's where it got interesting. Once content can be translated, you need to know when a translation is stale — i.e., the base-language content changed since the last translation ran. The first approach was straightforward: hash the base content, store the hash, and compare on every check.

// Before: hash comparison to detect staleness
function needsRetranslation(block) {
  return block.contentHash !== hash(block.content);
}

This is a classic "looks reasonable, breaks in practice" pattern. A few ways it went wrong:

  • False positives: any structural change to the content object — even ones that didn't affect translatable text, like reordering internal metadata fields — changed the hash and triggered an unnecessary re-translation.

  • False negatives: certain update paths mutated content without going through the code path that recalculated the hash, so a genuine change could go undetected.

  • Opacity: when a translation did get marked stale, there was no record of why — was it a real content edit, or a hash quirk? Debugging meant diffing raw content blobs.

The fix was to stop inferring staleness from a hash and instead track it explicitly as a status flag, set at the moment a base-language edit actually happens:

// After: explicit status, set at the moment of a real content edit
function onBaseContentEdit(block) {
  block.translationStatus = 'stale';
  markDependentLanguagesStale(block);
}

function needsRetranslation(block) {
  return block.translationStatus === 'stale';
}

This moved the "did the content meaningfully change" decision to the one place that actually knows the answer — the edit handler itself — instead of trying to reverse-engineer it later from a hash diff. It also meant we could log why something went stale, because the state transition was explicit and traceable instead of implicit and inferred.

The propagation is straightforward once it's explicit instead of inferred:

The takeaway

The general lesson generalizes past translation: whenever you catch yourself computing a derived signal (a hash, a checksum, a diff) to infer whether "something meaningful changed," ask whether you can instead set an explicit flag at the one place where that meaningful change actually happens. Implicit state inferred after the fact is fragile in ways that are hard to predict up front and annoying to debug after the fact; explicit state set at the source is boring, traceable, and almost always worth the small amount of extra plumbing.