<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Dharmesh Joshi]]></title><description><![CDATA[Dharmesh Joshi]]></description><link>https://dharmeshjoshi.hashnode.dev</link><image><url>https://cdn.hashnode.com/res/hashnode/image/upload/v1593680282896/kNC7E8IR4.png</url><title>Dharmesh Joshi</title><link>https://dharmeshjoshi.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Mon, 05 Oct 2026 11:45:20 GMT</lastBuildDate><atom:link href="https://dharmeshjoshi.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[Designing a Data Model for Translatable Content Without Breaking Every Existing Record]]></title><description><![CDATA[Adding multi-language support to a content platform sounds like a straightforward data modeling problem until you remember one thing: you already have production data in the old shape, and you can't j]]></description><link>https://dharmeshjoshi.hashnode.dev/designing-a-data-model-for-translatable-content-without-breaking-every-existing-record</link><guid isPermaLink="true">https://dharmeshjoshi.hashnode.dev/designing-a-data-model-for-translatable-content-without-breaking-every-existing-record</guid><category><![CDATA[TypeScript]]></category><category><![CDATA[i18n]]></category><category><![CDATA[Software Engineering]]></category><category><![CDATA[data-modeling]]></category><category><![CDATA[System Design]]></category><dc:creator><![CDATA[Dharmesh Joshi]]></dc:creator><pubDate>Fri, 02 Oct 2026 04:36:56 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6aa951c2af6c71a9e9a987b4/611e89de-f940-42b1-9930-f8b88ef1ebbc.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Adding multi-language support to a content platform sounds like a straightforward data modeling problem until you remember one thing: you already have production data in the old shape, and you can't just migrate it all at once without downtime, data loss risk, or both. Here's how we designed a translatable-content model that could hold both the old shape and the new shape at once, and the change-detection bug that came with it.</p>
<h3>The type that let old and new data coexist</h3>
<p>The starting point was simple: every text field on a piece of content was just a <code>string</code>. Adding translations meant a field needed to potentially hold content in multiple languages at once. The tempting-but-wrong move is to migrate every field to a language-keyed object in one big migration. The safer move — the one we took — was a union type:</p>
<pre><code class="language-ts">type TranslatableField = string | LanguageObject;

interface LanguageObject {
  en: string;
  es?: string;
  fr?: string;
  // ...other language codes
  _meta?: Record&lt;string, { translatedAt: string; source: 'human' | 'machine' }&gt;;
}
</code></pre>
<p>Every existing record with a plain string stays valid, untouched, forever if needed. Every <em>new</em> piece of translated content gets the richer object shape. No big-bang migration, no downtime, no risk of a half-migrated dataset.</p>
<p>The cost of that flexibility is that you can't read these fields directly anymore — every call site that used to do <code>block.title</code> now has to handle both shapes:</p>
<pre><code class="language-ts">function getLanguageContent(field: TranslatableField, lang: string, fallback = 'en'): string {
  if (typeof field === 'string') return field;         // legacy record, no translations yet
  return field[lang] ?? field[fallback] ?? '';
}
</code></pre>
<p>We centralized that logic into a small set of helper functions and made it a hard rule: nothing reads a translatable field directly, everything goes through the helper. That rule is the entire reason the union type is livable — without it, every component and every API response handler would need to duplicate the "is this a string or an object" check, and someone would eventually forget.</p>
<p>Here's the shape of that type, and how a legacy record and a new record both satisfy it:</p>
<img src="https://cdn.hashnode.com/uploads/covers/6aa951c2af6c71a9e9a987b4/624ca345-4299-4060-a5ab-1a341060f955.png" alt="" style="display:block;margin:0 auto" />

<h3>The "base language" concept, and why editing needed a lock</h3>
<p>Once content can exist in multiple languages, you run into a subtler problem: what happens when someone edits a piece of content that has <em>already been translated</em>? If they edit the English version, does the Spanish translation silently go stale? Does the UI let them edit the Spanish version directly and diverge from what a translation pipeline would produce?</p>
<p>We settled on a "base language" model: one language is the source of truth for a given piece of content, and non-base-language editing is intentionally restricted — direct structural edits (reordering answer options, changing what's "correct," restructuring content) are locked on non-base languages, because those changes need to flow from the base language outward through translation, not be made independently in each language and drift apart. Only wording, not structure, is safe to hand-edit per language.</p>
<p>This is as much a UX decision as a data model one, but it only works <em>because</em> the data model can represent "language X is derived from language Y" cleanly — you can't build sane rules around a mush of independent strings.</p>
<h3>The bug: detecting "does this need re-translation" by hashing content</h3>
<p>Here's where it got interesting. Once content can be translated, you need to know when a translation is stale — i.e., the base-language content changed since the last translation ran. The first approach was straightforward: hash the base content, store the hash, and compare on every check.</p>
<pre><code class="language-js">// Before: hash comparison to detect staleness
function needsRetranslation(block) {
  return block.contentHash !== hash(block.content);
}
</code></pre>
<p>This is a classic "looks reasonable, breaks in practice" pattern. A few ways it went wrong:</p>
<ul>
<li><p><strong>False positives</strong>: any structural change to the content object — even ones that didn't affect translatable text, like reordering internal metadata fields — changed the hash and triggered an unnecessary re-translation.</p>
</li>
<li><p><strong>False negatives</strong>: certain update paths mutated content without going through the code path that recalculated the hash, so a genuine change could go undetected.</p>
</li>
<li><p><strong>Opacity</strong>: when a translation <em>did</em> get marked stale, there was no record of <em>why</em> — was it a real content edit, or a hash quirk? Debugging meant diffing raw content blobs.</p>
</li>
</ul>
<p>The fix was to stop inferring staleness from a hash and instead track it explicitly as a status flag, set at the moment a base-language edit actually happens:</p>
<pre><code class="language-js">// After: explicit status, set at the moment of a real content edit
function onBaseContentEdit(block) {
  block.translationStatus = 'stale';
  markDependentLanguagesStale(block);
}

function needsRetranslation(block) {
  return block.translationStatus === 'stale';
}
</code></pre>
<p>This moved the "did the content meaningfully change" decision to the one place that actually knows the answer — the edit handler itself — instead of trying to reverse-engineer it later from a hash diff. It also meant we could log <em>why</em> something went stale, because the state transition was explicit and traceable instead of implicit and inferred.</p>
<p>The propagation is straightforward once it's explicit instead of inferred:</p>
<img src="https://cdn.hashnode.com/uploads/covers/6aa951c2af6c71a9e9a987b4/3604b28d-472f-4117-a0c5-3dfd75c10d8d.png" alt="" style="display:block;margin:0 auto" />

<h3>The takeaway</h3>
<p>The general lesson generalizes past translation: whenever you catch yourself computing a derived signal (a hash, a checksum, a diff) to infer whether "something meaningful changed," ask whether you can instead set an explicit flag at the one place where that meaningful change actually happens. Implicit state inferred after the fact is fragile in ways that are hard to predict up front and annoying to debug after the fact; explicit state set at the source is boring, traceable, and almost always worth the small amount of extra plumbing.</p>
]]></content:encoded></item><item><title><![CDATA[Streaming Video Behind a Signed-Cookie CDN: Lessons from a Real Migration]]></title><description><![CDATA[We had a video player in an eLearning platform that was serving video directly from origin, with a signed URL generated per request. It worked, but it didn't scale well and it wasn't resilient to some]]></description><link>https://dharmeshjoshi.hashnode.dev/streaming-video-behind-a-signed-cookie-cdn-lessons-from-a-real-migration</link><guid isPermaLink="true">https://dharmeshjoshi.hashnode.dev/streaming-video-behind-a-signed-cookie-cdn-lessons-from-a-real-migration</guid><category><![CDATA[System Design]]></category><category><![CDATA[webdev]]></category><category><![CDATA[JavaScript]]></category><category><![CDATA[CDN]]></category><category><![CDATA[streaming]]></category><dc:creator><![CDATA[Dharmesh Joshi]]></dc:creator><pubDate>Fri, 25 Sep 2026 06:51:18 GMT</pubDate><content:encoded><![CDATA[<p>We had a video player in an eLearning platform that was serving video directly from origin, with a signed URL generated per request. It worked, but it didn't scale well and it wasn't resilient to some real-world browser behavior we kept running into. So we migrated it to CDN-based HLS streaming with signed cookies. Here's what that migration actually involved, and the bugs that taught me the most.</p>
<h3>Why signed URLs on origin weren't enough</h3>
<p>A signed-URL-per-video approach means every video request round-trips to your origin to mint a URL, then the actual bytes still have to come from <em>somewhere</em> — and if that somewhere is your own servers, you're paying for bandwidth and scaling problems a CDN exists to solve. Worse, HLS doesn't fetch "a video" — it fetches a manifest, then a stream of segment files, potentially dozens of small requests per minute of video. Signing each of those individually is a non-starter.</p>
<p>The standard answer is CDN-level signed cookies: authenticate once, get a cookie that scopes CDN access to a specific resource path and time window, and let every subsequent manifest/segment request ride on that cookie without a fresh signature each time.</p>
<pre><code class="language-js">// Before: a fresh signed URL minted per video request, hitting origin logic every time
function getVideoUrl(videoId, user) {
  return `${ORIGIN}/videos/${videoId}?token=${signToken(videoId, user)}`;
}
</code></pre>
<pre><code class="language-js">// After: signed cookies scoped to a whole course's worth of video, set once per session
function setStreamingCookies(res, { courseId }) {
  const policy = buildCdnPolicy({
    resource: `${CDN_DOMAIN}/videos/${courseId}/*`,
    expiresAt: Date.now() + SESSION_TTL,
  });
  const { signature, keyPairId } = signPolicy(policy, CDN_PRIVATE_KEY);

  res.cookie('CDN-Policy', policy, cookieOpts);
  res.cookie('CDN-Signature', signature, cookieOpts);
  res.cookie('CDN-Key-Pair-Id', keyPairId, cookieOpts);
}
</code></pre>
<p>This cut a huge number of origin round-trips down to one cookie-set per session, and let the CDN do what it's good at.</p>
<p>Here's the shape of the request flow, cookie path vs. fallback path:</p>
<img src="https://cdn.hashnode.com/uploads/covers/6aa951c2af6c71a9e9a987b4/efcd19fe-b1a4-46c2-9086-7b76b439c17a.png" alt="Request flow: signed-cookie happy path vs. signed-token fallback when third-party cookies are blocked" style="display:block;margin:0 auto" />

<h3>The fallback nobody thinks about until it breaks in production</h3>
<p>Signed cookies are, well, cookies — and modern browsers have gotten aggressive about blocking third-party cookies by default. If your CDN domain differs from your app's domain (it usually does), a meaningful slice of your users will silently fail to authenticate against the CDN, with no obvious error beyond "video won't play."</p>
<p>The fix was a fallback path: detect when cookie-based auth isn't landing, and fall back to appending a signed token directly on the manifest/segment request as a query parameter instead of relying on the cookie jar. It's less efficient (you're back to per-request signing for that user), but it only kicks in for the subset of sessions that need it, and it turns a hard failure into a degraded-but-working experience.</p>
<p>The real lesson: don't design an auth strategy for a CDN with only the happy path in mind. Assume a meaningful percentage of your users have some cookie-blocking behavior on by default, and have a tested fallback — not a TODO — before you ship.</p>
<h3>The race condition that taught me the most: switching audio tracks</h3>
<p>Once the streaming layer was solid, the trickiest bugs moved into the player itself. The one I learned the most from: switching the active audio track (for multi-language audio) while a previous switch was still in flight.</p>
<p>The naive implementation just swapped the active track index directly:</p>
<pre><code class="language-js">// Before: no guard against overlapping switches
function switchAudioTrack(player, trackId) {
  player.audioTrack = trackId;
}
</code></pre>
<p>Under normal use this is fine. But if a user clicked through two languages quickly, or if an automatic "match audio to caption language" switch overlapped with a manual switch, you'd get two async track-loading operations racing — and whichever one <em>finished</em> last won, regardless of which one the user actually wanted last. The symptom was intermittent: audio would occasionally play in the wrong language, or briefly glitch/mute during a switch, and it was maddening to reproduce because it depended on network timing.</p>
<p>The fix was to make sure only the most recent switch is allowed to apply. Every switch gets an ID, and when its track finishes loading, it checks whether a newer switch has started in the meantime. If one has, it throws its own result away:</p>
<pre><code class="language-js">// After: every switch gets an ID, and only the most recent one is allowed to apply
let latestRequest = 0;

async function switchAudioTrack(player, trackId) {
  const requestId = ++latestRequest;
  await loadAudioTrack(player, trackId); // fetch and prepare the track, but don't activate it yet
  if (requestId !== latestRequest) return; // a newer switch started while this one loaded, so drop it
  player.audioTrack = trackId;
}
</code></pre>
<p>It doesn't matter how many times the user clicks or which load finishes first: the track that plays is always the one they picked last.</p>
<p>This one bug taught me to be much more suspicious of any UI action that triggers an async operation and mutates shared state on completion — "the last click wins" only holds if you guarantee ordering, and async code doesn't guarantee ordering for free.</p>
<h3>Wrapping up</h3>
<p>The overall shape of this migration was: move the expensive path (bytes) to infrastructure built for it (CDN), be honest about the auth mechanism's failure modes (cookies get blocked — build the fallback), and once the transport layer is solid, expect the interesting bugs to move into concurrent state management in the player itself.</p>
]]></content:encoded></item></channel></rss>