<?xml version="1.0" encoding="utf-8"?>
<rss xmlns:atom="http://www.w3.org/2005/Atom" xmlns:webfeeds="http://webfeeds.org/rss/1.0" version="2.0">
  <channel>
    <atom:link href="http://pubsubhubbub.appspot.com/" rel="hub"/>
    <atom:link href="https://f43.me/netflix-tech-blog.xml" rel="self" type="application/rss+xml"/>
    <title>Netflix Tech Blog</title>
    <description>Learn about Netflix’s world class engineering efforts, company culture, product developments and more. - Medium</description>
    <link>http://medium.com</link>
    <webfeeds:icon>https://s2.googleusercontent.com/s2/favicons?alt=feed&amp;domain=medium.com</webfeeds:icon>
    <generator>f43.me</generator>
    <lastBuildDate>Sat, 05 Sep 2026 03:45:42 +0200</lastBuildDate>
    <item>
      <title><![CDATA[MAPS: Netflix’s Multimodal Asset Personalization at Scale]]></title>
      <description><![CDATA[Attention Required! | Cloudflare
  <div id="cf-wrapper"><p>Please enable cookies.</p><div id="cf-error-details" class="cf-error-details-wrapper"><div class="cf-section cf-wrapper"><div class="cf-columns two"><div class="cf-column"><h2 data-translate="blocked_why_headline">Why have I been blocked?</h2><p data-translate="blocked_why_detail">This website is using a security service to protect itself from online attacks. The action you just performed triggered the security solution. There are several actions that could trigger this block including submitting a certain word or phrase, a SQL command or malformed data.</p></div><div class="cf-column"><h2 data-translate="blocked_resolve_headline">What can I do to resolve this?</h2><p data-translate="blocked_resolve_detail">You can email the site owner to let them know you were blocked. Please include what you were doing when this page came up and the Cloudflare Ray ID found at the bottom of this page.</p></div></div></div></div></div>]]></description>
      <link>https://netflixtechblog.com/maps-netflixs-multimodal-asset-personalization-at-scale-32f96320785e</link>
      <guid>https://netflixtechblog.com/maps-netflixs-multimodal-asset-personalization-at-scale-32f96320785e</guid>
      <pubDate>Fri, 28 Aug 2026 18:01:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[A Tale of Two Flink Autoscalers]]></title>
      <description><![CDATA[Attention Required! | Cloudflare
  <div id="cf-wrapper"><p>Please enable cookies.</p><div id="cf-error-details" class="cf-error-details-wrapper"><div class="cf-section cf-wrapper"><div class="cf-columns two"><div class="cf-column"><h2 data-translate="blocked_why_headline">Why have I been blocked?</h2><p data-translate="blocked_why_detail">This website is using a security service to protect itself from online attacks. The action you just performed triggered the security solution. There are several actions that could trigger this block including submitting a certain word or phrase, a SQL command or malformed data.</p></div><div class="cf-column"><h2 data-translate="blocked_resolve_headline">What can I do to resolve this?</h2><p data-translate="blocked_resolve_detail">You can email the site owner to let them know you were blocked. Please include what you were doing when this page came up and the Cloudflare Ray ID found at the bottom of this page.</p></div></div></div></div></div>]]></description>
      <link>https://netflixtechblog.com/a-tale-of-two-flink-autoscalers-e9f6a1b1492b</link>
      <guid>https://netflixtechblog.com/a-tale-of-two-flink-autoscalers-e9f6a1b1492b</guid>
      <pubDate>Fri, 21 Aug 2026 18:01:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[How and Why Netflix Built a Real-Time Distributed Graph: Part 3 — Querying the graph with gRPC…]]></title>
      <description><![CDATA[Attention Required! | Cloudflare
  <div id="cf-wrapper"><p>Please enable cookies.</p><div id="cf-error-details" class="cf-error-details-wrapper"><div class="cf-section cf-wrapper"><div class="cf-columns two"><div class="cf-column"><h2 data-translate="blocked_why_headline">Why have I been blocked?</h2><p data-translate="blocked_why_detail">This website is using a security service to protect itself from online attacks. The action you just performed triggered the security solution. There are several actions that could trigger this block including submitting a certain word or phrase, a SQL command or malformed data.</p></div><div class="cf-column"><h2 data-translate="blocked_resolve_headline">What can I do to resolve this?</h2><p data-translate="blocked_resolve_detail">You can email the site owner to let them know you were blocked. Please include what you were doing when this page came up and the Cloudflare Ray ID found at the bottom of this page.</p></div></div></div></div></div>]]></description>
      <link>https://netflixtechblog.com/how-and-why-netflix-built-a-real-time-distributed-graph-part-3-querying-the-graph-with-grpc-0f3468349607</link>
      <guid>https://netflixtechblog.com/how-and-why-netflix-built-a-real-time-distributed-graph-part-3-querying-the-graph-with-grpc-0f3468349607</guid>
      <pubDate>Fri, 07 Aug 2026 18:01:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Modeling Device Capabilities for Analytics]]></title>
      <description><![CDATA[<article><div class="e"><div class="e"><section><div><div><div class="am ew"><div class="fd cs hf hg hh hi"></div></div><div class="gz ib ic id ie"><div class="am ew"><div class="fd cs hf hg hh hi"><div><div></div><p id="661a" class="pw-post-body-paragraph nm nn ih no b np nq nr ns nt nu nv nw fo nx ny nz fr oa ob oc fu od oe of og gz cv">by <a class="bx oh" href="https://www.linkedin.com/in/aarti-laddha-70666557/" rel="noopener ugc nofollow" target="_blank">Aarti Laddha</a>, <a class="bx oh" href="https://www.linkedin.com/in/richardjcool/" rel="noopener ugc nofollow" target="_blank">Richard Diaz-Cool</a>, <a class="bx oh" href="https://www.linkedin.com/in/rishikaidnani/" rel="noopener ugc nofollow" target="_blank">Rishika Idnani</a>, <a class="bx oh" href="https://www.linkedin.com/in/venkatesh-selvaraj-88824137/" rel="noopener ugc nofollow" target="_blank">Venkatesh Selveraj</a></p><p id="7433" class="pw-post-body-paragraph nm nn ih no b np nq nr ns nt nu nv nw fo nx ny nz fr oa ob oc fu od oe of og gz cv">Netflix supports a vast and evolving set of features and content types, ranging from 4K streaming and immersive audio to live streaming and cloud gaming, across a diverse ecosystem of devices. However, not all devices are created equal. Hardware limitations such as available RAM, CPU cores, display capabilities, or platform support mean that some features cannot be supported on certain device models. To ensure the best possible user experience, we rely on a deep understanding of device capabilities. We have invested in building a comprehensive device capability data model and integrating feature flags from internal systems, paving the way for smarter, more granular feature management across our global device landscape. This approach helps us identify bottlenecks in feature penetration and accelerates the pace of innovation.</p><div class="oi et oj ok am j gt"><div class="ol om on oo op e bv"><h2 class="oq b or os ot ou ov">Get Netflix Technology Blog’s stories in your inbox</h2></div><div class="ow e bv"><p class="z b cr u w">Join Medium for free to get updates from this writer.</p></div><div class="pa pb pc pd pe pf am"><div class="pg am gt"><div class="am pg br bq hl ph pi ei mf pj pk pl pm pn po pp"><input class="pq ca by cb pr ps dx ls cw pt cs" placeholder="Enter your email" type="text" value="" /></div></div><div class="ck cz pu ej"><div class="ox oy oz"><button class="z b cr u pv pw px py pz qa qb bk bl qc qd qe bp bq br bs bt bu bv">Subscribe</button></div></div><div class="cs da r s"><button class="z b cr u pv pw px py pz qa qb bk bl qc qd qe bp cs bq br bs bt bu bv">Subscribe</button></div></div><div class="qf e"><label class="am j"><input class="qg qh qi dr gg qj qk ql qm qn qo qp" type="checkbox" checked="checked" /></label></div><div class="e"><p class="z b cr u cv">Remember me for faster sign in</p></div></div><div class="qx e"></div><p id="d3a7" class="pw-post-body-paragraph nm nn ih no b np nr ns nt nv nw fo ny nz fr ob oc fu oe of ok og gz cv">We have designed our data storage and modeling strategies to efficiently support analytics at scale. We use a cumulative table to process information about the device’s capabilities. This table is structured to efficiently capture the latest state of each device and its associated capabilities (like Screen resolutions, Video Profiles Supported, Surround Sound, RAM size etc) making it ideal for analytics and reporting use cases.</p><pre class="qy qz ra rb rc rd re rf hl rg cn cv">{<br />"Screen Height": ["720"],<br />"Screen Width": ["1280"],<br />"Video Profiles": <br />[<br />"playready",<br />"hevc",<br />],<br />}</pre><p id="cd26" class="pw-post-body-paragraph nm nn ih no b np nq nr ns nt nu nv nw fo nx ny nz fr oa ob oc fu od oe of og gz cv">For aggregate analytics, we leverage a histogram table that captures active device counts over the past 28 days, broken down by device model and software version. This table also records the number of devices supporting specific capabilities, enabling detailed distribution analysis. One use case for this histogram data is to analyze the distribution of external display capabilities attached to streaming sticks. For example, the histogram below shows that out of total X number of devices, all supported the HD profile (playready), while only 20% devices supported the UHD profile (hevc).</p><pre class="qy qz ra rb rc rd re rf hl rg cn cv">{<br />"Video Profiles": {<br />      "playready": 100%, # HD profile<br />      "hevc": 20% # UHD profile<br />}<br />}</pre><p id="3d72" class="pw-post-body-paragraph nm nn ih no b np nq nr ns nt nu nv nw fo nx ny nz fr oa ob oc fu od oe of og gz cv">We have built analytical products that leverage these datasets to provide a comprehensive view of feature reach such as 4K Ultra HD, Netflix Spatial Audio, Cloud Gaming and the latest UI. By relying on data-driven insights, we can make informed decisions about which features to enable on specific devices, ensuring both performance and reliability.</p></div></div></div></div></div></div></section></div></div></article><article aria-label="In-House LLM Serving at Netflix" class="al" data-testid="post-preview"><div class="al rt e"><div class="cs"><div class="al e"><div class="dm al we wf wg wh wi wj wk wl wm wn wo wp wq"><div class="wr"><div aria-label="In-House LLM Serving at Netflix"><div class="wt wu wv ww ek"><img alt="In-House LLM Serving at Netflix" class="cs wx wy wz xa" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*GKGOrp0xddZwHomMhiSeCA.png" /></div></div></div><div class="ws am ew gt"><div class="am gt qs cs xb xc xd xe"><div class="xf xg xh xi xj fd am j"><div class="fd am j xk"><div class="xl e ej"><div><div class="e"><div tabindex="-1" class="cq"><a href="https://netflixtechblog.com" rel="noopener follow"><div class="dm"><img alt="Netflix TechBlog" class="ei gv e xn xm" src="https://miro.medium.com/v2/resize:fill:40:40/1*ty4NvNrGg4ReETxqU2N3Og.png" width="20" height="20" /></div></a></div></div></div></div><div class="td e ej dd"><p class="z b y u w">In</p></div><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai jz am j" href="https://netflixtechblog.com" rel="noopener follow"><p class="z b y u dv xo es xp xq xr xs xt cv">Netflix TechBlog</p></a></div></div></div></div></div><div class="xu e ej dd"><p class="z b y u w">by</p></div><div class="fd am j xk"><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai jz am j" href="https://netflixtechblog.medium.com" rel="noopener follow"><p class="z b y u dv xo es xp xq xr xs xt cv">Netflix Technology Blog</p></a></div></div></div></div></div><div class="mk am j ej db dd"><p class="z b y u w">·</p><p class="z b y u w">Jul 17</p></div></div><div class="xv e xw xx xy xz ya gz"><div class="yb yc yd ye yf yg yh yi"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" rel="noopener follow" href="https://netflixtechblog.com/in-house-llm-serving-at-netflix-a5a8e799ea2c" data-discover="true"><div title=""><h2 class="z ii tn to yj yk yl ym yn yo fl tp tq yp yq yr ys yt yu fn fo fp yv yw yx yy yz za fq fr fs zb zc zd ze zf zg ft fu fv zh zi zj zk zl zm fw xt cv">In-House LLM Serving at Netflix</h2></div><div class="zn e"><h3 class="z b fz u dv zo es xp zp xr xt w">By AI Platform’s Model Runtime team and Inference team</h3></div></a></div></div><div class="su am je ce"><div class="am j jp"><div class="ck cz da r s"><div class="dm zq zr am j"><div class="dr mf zs am j jp"><a class="dr gg zs am j jp" tabindex="-1" rel="noopener follow" href="https://netflixtechblog.com/in-house-llm-serving-at-netflix-a5a8e799ea2c" data-discover="true"><div><div class="am"><div tabindex="-1" class="cq"><div class="am j db">A response icon10</div></div></div></div></a></div></div></div><div class="am j zu zv"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fmodeling-device-capabilities-for-analytics-e7607acebde8" rel="noopener follow"><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></a><div class="ck cz da r s"><div><div class="bt"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2Fa5a8e799ea2c&amp;operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fin-house-llm-serving-at-netflix-a5a8e799ea2c" rel="noopener follow"></a></div></div></div></div></div></div><div class="da r s"></div></div></div></div></div></div></div></div></article><article aria-label="Building Service Topology at Scale: Architecture, Challenges, and Lessons Learned" class="al" data-testid="post-preview"><div class="al rt e"><div class="cs"><div class="al e"><div class="dm al we wf wg wh wi wj wk wl wm wn wo wp wq"><div class="wr"><div aria-label="Diagram showing two separate ingestion pipelines — a flow log pipeline and an IPC pipeline, each fed by data enrichment — writing to their own graph store, with a shared API serving UI and backend clients"><div class="wt wu wv ww ek"><img alt="Diagram showing two separate ingestion pipelines — a flow log pipeline and an IPC pipeline, each fed by data enrichment — writing to their own graph store, with a shared API serving UI and backend clients" class="cs wx wy wz xa" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*5_09VEnyPtFR0tbO1zzzMQ.png" /></div></div></div><div class="ws am ew gt"><div class="am gt qs cs xb xc xd xe"><div class="xf xg xh xi xj fd am j"><div class="fd am j xk"><div class="xl e ej"><div><div class="e"><div tabindex="-1" class="cq"><a href="https://netflixtechblog.com" rel="noopener follow"><div class="dm"><img alt="Netflix TechBlog" class="ei gv e xn xm" src="https://miro.medium.com/v2/resize:fill:40:40/1*ty4NvNrGg4ReETxqU2N3Og.png" width="20" height="20" /></div></a></div></div></div></div><div class="td e ej dd"><p class="z b y u w">In</p></div><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai jz am j" href="https://netflixtechblog.com" rel="noopener follow"><p class="z b y u dv xo es xp xq xr xs xt cv">Netflix TechBlog</p></a></div></div></div></div></div><div class="xu e ej dd"><p class="z b y u w">by</p></div><div class="fd am j xk"><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai jz am j" href="https://netflixtechblog.medium.com" rel="noopener follow"><p class="z b y u dv xo es xp xq xr xs xt cv">Netflix Technology Blog</p></a></div></div></div></div></div><div class="mk am j ej db dd"><p class="z b y u w">·</p><p class="z b y u w">Jul 13</p></div></div><div class="xv e xw xx xy xz ya gz"><div class="yb yc yd ye yf yg yh yi"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" rel="noopener follow" href="https://netflixtechblog.com/building-service-topology-at-scale-architecture-challenges-and-lessons-learned-f4b792f3f0d8" data-discover="true"><div title="Building Service Topology at Scale: Architecture, Challenges, and Lessons Learned"><h2 class="z ii tn to yj yk yl ym yn yo fl tp tq yp yq yr ys yt yu fn fo fp yv yw yx yy yz za fq fr fs zb zc zd ze zf zg ft fu fv zh zi zj zk zl zm fw xt cv">Building Service Topology at Scale: Architecture, Challenges, and Lessons Learned</h2></div><div class="zn e"><h3 class="z b fz u dv zo es xp zp xr xt w">By Parth Jain, Rakesh Sukumar, Yingwu Zhao, Renzo Sanchez-Silva &amp; Nathan Fisher. A deep dive into the engineering challenges of building a…</h3></div></a></div></div><div class="su am je ce"><div class="am j jp"><div class="ck cz da r s"><div class="dm zq zr am j"><div class="dr mf zs am j jp"><a class="dr gg zs am j jp" tabindex="-1" rel="noopener follow" href="https://netflixtechblog.com/building-service-topology-at-scale-architecture-challenges-and-lessons-learned-f4b792f3f0d8" data-discover="true"><div><div class="am"><div tabindex="-1" class="cq"><div class="am j db">A response icon8</div></div></div></div></a></div></div></div><div class="am j zu zv"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fmodeling-device-capabilities-for-analytics-e7607acebde8" rel="noopener follow"><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></a><div class="ck cz da r s"><div><div class="bt"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2Ff4b792f3f0d8&amp;operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fbuilding-service-topology-at-scale-architecture-challenges-and-lessons-learned-f4b792f3f0d8" rel="noopener follow"></a></div></div></div></div></div></div><div class="da r s"></div></div></div></div></div></div></div></div></article><article aria-label="GenPage: Towards End-to-End Generative Homepage Construction at Netflix" class="al" data-testid="post-preview"><div class="al rt e"><div class="cs"><div class="al e"><div class="dm al we wf wg wh wi wj wk wl wm wn wo wp wq"><div class="wr"><div aria-label="GenPage: Towards End-to-End Generative Homepage Construction at Netflix"><div class="wt wu wv ww ek"><img alt="GenPage: Towards End-to-End Generative Homepage Construction at Netflix" class="cs wx wy wz xa" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*NurrizMgC7_QbsGuuW42Dg.gif" /></div></div></div><div class="ws am ew gt"><div class="am gt qs cs xb xc xd xe"><div class="xf xg xh xi xj fd am j"><div class="fd am j xk"><div class="xl e ej"><div><div class="e"><div tabindex="-1" class="cq"><a href="https://netflixtechblog.com" rel="noopener follow"><div class="dm"><img alt="Netflix TechBlog" class="ei gv e xn xm" src="https://miro.medium.com/v2/resize:fill:40:40/1*ty4NvNrGg4ReETxqU2N3Og.png" width="20" height="20" /></div></a></div></div></div></div><div class="td e ej dd"><p class="z b y u w">In</p></div><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai jz am j" href="https://netflixtechblog.com" rel="noopener follow"><p class="z b y u dv xo es xp xq xr xs xt cv">Netflix TechBlog</p></a></div></div></div></div></div><div class="xu e ej dd"><p class="z b y u w">by</p></div><div class="fd am j xk"><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai jz am j" href="https://netflixtechblog.medium.com" rel="noopener follow"><p class="z b y u dv xo es xp xq xr xs xt cv">Netflix Technology Blog</p></a></div></div></div></div></div><div class="mk am j ej db dd"><p class="z b y u w">·</p><p class="z b y u w">Jun 29</p></div></div><div class="xv e xw xx xy xz ya gz"><div class="yb yc yd ye yf yg yh yi"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" rel="noopener follow" href="https://netflixtechblog.com/genpage-towards-end-to-end-generative-homepage-construction-at-netflix-77146fba8a08" data-discover="true"><div title="GenPage: Towards End-to-End Generative Homepage Construction at Netflix"><h2 class="z ii tn to yj yk yl ym yn yo fl tp tq yp yq yr ys yt yu fn fo fp yv yw yx yy yz za fq fr fs zb zc zd ze zf zg ft fu fv zh zi zj zk zl zm fw xt cv">GenPage: Towards End-to-End Generative Homepage Construction at Netflix</h2></div><div class="zn e"><h3 class="z b fz u dv zo es xp zp xr xt w">Authors: Lequn Wang, Jiangwei Pan, and Linas Baltrunas</h3></div></a></div></div><div class="su am je ce"><div class="am j jp"><div class="ck cz da r s"><div class="dm zq zr am j"><div class="dr mf zs am j jp"><a class="dr gg zs am j jp" tabindex="-1" rel="noopener follow" href="https://netflixtechblog.com/genpage-towards-end-to-end-generative-homepage-construction-at-netflix-77146fba8a08" data-discover="true"><div><div class="am"><div tabindex="-1" class="cq"><div class="am j db">A response icon12</div></div></div></div></a></div></div></div><div class="am j zu zv"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fmodeling-device-capabilities-for-analytics-e7607acebde8" rel="noopener follow"><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></a><div class="ck cz da r s"><div><div class="bt"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2F77146fba8a08&amp;operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fgenpage-towards-end-to-end-generative-homepage-construction-at-netflix-77146fba8a08" rel="noopener follow"></a></div></div></div></div></div></div><div class="da r s"></div></div></div></div></div></div></div></div></article><article aria-label="Measuring the Impact of Personalized Recommendations" class="al" data-testid="post-preview"><div class="al rt e"><div class="cs"><div class="al e"><div class="dm al we wf wg wh wi wj wk wl wm wn wo wp wq"><div class="wr"><div aria-label="Measuring the Impact of Personalized Recommendations"><div class="wt wu wv ww ek"><img alt="Measuring the Impact of Personalized Recommendations" class="cs wx wy wz xa" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*m_HbeqimgsdkEvyOGZE97A.png" /></div></div></div><div class="ws am ew gt"><div class="am gt qs cs xb xc xd xe"><div class="xf xg xh xi xj fd am j"><div class="fd am j xk"><div class="xl e ej"><div><div class="e"><div tabindex="-1" class="cq"><a tabindex="-1" href="https://netflixtechblog.medium.com" rel="noopener follow"><div class="e dm"><img alt="Netflix Technology Blog" class="e bs ee xm xn ei" src="https://miro.medium.com/v2/resize:fill:40:40/1*BJWRqfSMf9Da9vsXG9EBRQ.jpeg" width="20" height="20" /></div></a></div></div></div></div><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai jz am j" href="https://netflixtechblog.medium.com" rel="noopener follow"><p class="z b y u dv xo es xp xq xr xs xt cv">Netflix Technology Blog</p></a></div></div></div></div></div><div class="mk am j ej db dd"><p class="z b y u w">·</p><p class="z b y u w">Jul 10</p></div></div><div class="xv e xw xx xy xz ya gz"><div class="yb yc yd ye yf yg yh yi"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://netflixtechblog.medium.com/measuring-the-impact-of-personalized-recommendations-4c26be3a4d96" rel="noopener follow"><div title=""><h2 class="z ii tn to yj yk yl ym yn yo fl tp tq yp yq yr ys yt yu fn fo fp yv yw yx yy yz za fq fr fs zb zc zd ze zf zg ft fu fv zh zi zj zk zl zm fw xt cv">Measuring the Impact of Personalized Recommendations</h2></div><div class="zn e"><h3 class="z b fz u dv zo es xp zp xr xt w">By Kevin Zielnicki, Guy Aridor, Aurélien Bibaut, Allen Tran, Winston Chou, and Nathan Kallus</h3></div></a></div></div><div class="su am je ce"><div class="am j jp"><div class="ck cz da r s"><div class="dm zq zr am j"><div class="dr mf zs am j jp"><a class="dr gg zs am j jp" tabindex="-1" href="https://netflixtechblog.medium.com/measuring-the-impact-of-personalized-recommendations-4c26be3a4d96" rel="noopener follow"><div><div class="am"><div tabindex="-1" class="cq"><div class="am j db">A response icon9</div></div></div></div></a></div></div></div><div class="am j zu zv"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fmodeling-device-capabilities-for-analytics-e7607acebde8" rel="noopener follow"><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></a><div class="ck cz da r s"><div><div class="bt"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2F4c26be3a4d96&amp;operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.medium.com%2Fmeasuring-the-impact-of-personalized-recommendations-4c26be3a4d96" rel="noopener follow"></a></div></div></div></div></div></div></div></div></div></div></div></div></div></article><article aria-label="The Medallion Architecture Reconsidered: What It Solved, and Where It’s Cracking" class="al" data-testid="post-preview"><div class="al rt e"><div class="cs"><div class="al e"><div class="dm al we wf wg wh wi wj wk wl wm wn wo wp wq"><div class="wr"><div aria-label="The Medallion Architecture Reconsidered: What It Solved, and Where It’s Cracking"><div class="wt wu wv ww ek"><img alt="The Medallion Architecture Reconsidered: What It Solved, and Where It’s Cracking" class="cs wx wy wz xa" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*5zGpctNCnFEA7A-Mmn9SIQ.png" /></div></div></div><div class="ws am ew gt"><div class="am gt qs cs xb xc xd xe"><div class="xf xg xh xi xj fd am j"><div class="fd am j xk"><div class="xl e ej"><div><div class="e"><div tabindex="-1" class="cq"><a href="https://levelup.gitconnected.com" rel="noopener follow"><div class="dm"><img alt="Level Up Coding" class="ei gv e xn xm" src="https://miro.medium.com/v2/resize:fill:40:40/1*5D9oYBd58pyjMkV_5-zXXQ.jpeg" width="20" height="20" /></div></a></div></div></div></div><div class="td e ej dd"><p class="z b y u w">In</p></div><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai jz am j" href="https://levelup.gitconnected.com" rel="noopener follow"><p class="z b y u dv xo es xp xq xr xs xt cv">Level Up Coding</p></a></div></div></div></div></div><div class="xu e ej dd"><p class="z b y u w">by</p></div><div class="fd am j xk"><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai jz am j" href="https://blog.santoshshinde.com" rel="noopener follow"><p class="z b y u dv xo es xp xq xr xs xt cv">Santosh Shinde</p></a></div></div></div></div></div><div class="mk am j ej db dd"><p class="z b y u w">·</p><p class="z b y u w">Jul 4</p></div></div><div class="xv e xw xx xy xz ya gz"><div class="yb yc yd ye yf yg yh yi"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://levelup.gitconnected.com/the-medallion-architecture-reconsidered-what-it-solved-and-where-its-cracking-d6073aa1b0fd" rel="noopener follow"><div title="The Medallion Architecture Reconsidered: What It Solved, and Where It’s Cracking"><h2 class="z ii tn to yj yk yl ym yn yo fl tp tq yp yq yr ys yt yu fn fo fp yv yw yx yy yz za fq fr fs zb zc zd ze zf zg ft fu fv zh zi zj zk zl zm fw xt cv">The Medallion Architecture Reconsidered: What It Solved, and Where It’s Cracking</h2></div><div class="zn e"><h3 class="z b fz u dv zo es xp zp xr xt w">Rethinking Bronze, Silver, and Gold for the AI era</h3></div></a></div></div><div class="su am je ce"><div class="am j jp"><div class="rt am"><button class="e cf af ac" aria-label="Member-only story"><div class=""><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></div></button></div><div class="ck cz da r s"><div class="dm zq zr am j"><div class="dr mf zs am j jp"><a class="dr gg zs am j jp" tabindex="-1" href="https://levelup.gitconnected.com/the-medallion-architecture-reconsidered-what-it-solved-and-where-its-cracking-d6073aa1b0fd" rel="noopener follow"><div><div class="am"><div tabindex="-1" class="cq"><div class="am j db">A response icon4</div></div></div></div></a></div></div></div><div class="am j zu zv"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fmodeling-device-capabilities-for-analytics-e7607acebde8" rel="noopener follow"><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></a><div class="ck cz da r s"><div><div class="bt"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2Fd6073aa1b0fd&amp;operation=register&amp;redirect=https%3A%2F%2Flevelup.gitconnected.com%2Fthe-medallion-architecture-reconsidered-what-it-solved-and-where-its-cracking-d6073aa1b0fd" rel="noopener follow"></a></div></div></div></div></div></div><div class="da r s"></div></div></div></div></div></div></div></div></article><article aria-label="In-House LLM Serving at Netflix" class="al" data-testid="post-preview"><div class="al rt e"><div class="cs"><div class="al e"><div class="dm al we wf wg wh wi wj wk wl wm wn wo wp wq"><div class="wr"><div aria-label="In-House LLM Serving at Netflix"><div class="wt wu wv ww ek"><img alt="In-House LLM Serving at Netflix" class="cs wx wy wz xa" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*GKGOrp0xddZwHomMhiSeCA.png" /></div></div></div><div class="ws am ew gt"><div class="am gt qs cs xb xc xd xe"><div class="xf xg xh xi xj fd am j"><div class="fd am j xk"><div class="xl e ej"><div><div class="e"><div tabindex="-1" class="cq"><a href="https://netflixtechblog.com" rel="noopener follow"><div class="dm"><img alt="Netflix TechBlog" class="ei gv e xn xm" src="https://miro.medium.com/v2/resize:fill:40:40/1*ty4NvNrGg4ReETxqU2N3Og.png" width="20" height="20" /></div></a></div></div></div></div><div class="td e ej dd"><p class="z b y u w">In</p></div><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai jz am j" href="https://netflixtechblog.com" rel="noopener follow"><p class="z b y u dv xo es xp xq xr xs xt cv">Netflix TechBlog</p></a></div></div></div></div></div><div class="xu e ej dd"><p class="z b y u w">by</p></div><div class="fd am j xk"><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai jz am j" href="https://netflixtechblog.medium.com" rel="noopener follow"><p class="z b y u dv xo es xp xq xr xs xt cv">Netflix Technology Blog</p></a></div></div></div></div></div><div class="mk am j ej db dd"><p class="z b y u w">·</p><p class="z b y u w">Jul 17</p></div></div><div class="xv e xw xx xy xz ya gz"><div class="yb yc yd ye yf yg yh yi"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" rel="noopener follow" href="https://netflixtechblog.com/in-house-llm-serving-at-netflix-a5a8e799ea2c" data-discover="true"><div title=""><h2 class="z ii tn to yj yk yl ym yn yo fl tp tq yp yq yr ys yt yu fn fo fp yv yw yx yy yz za fq fr fs zb zc zd ze zf zg ft fu fv zh zi zj zk zl zm fw xt cv">In-House LLM Serving at Netflix</h2></div><div class="zn e"><h3 class="z b fz u dv zo es xp zp xr xt w">By AI Platform’s Model Runtime team and Inference team</h3></div></a></div></div><div class="su am je ce"><div class="am j jp"><div class="ck cz da r s"><div class="dm zq zr am j"><div class="dr mf zs am j jp"><a class="dr gg zs am j jp" tabindex="-1" rel="noopener follow" href="https://netflixtechblog.com/in-house-llm-serving-at-netflix-a5a8e799ea2c" data-discover="true"><div><div class="am"><div tabindex="-1" class="cq"><div class="am j db">A response icon10</div></div></div></div></a></div></div></div><div class="am j zu zv"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fmodeling-device-capabilities-for-analytics-e7607acebde8" rel="noopener follow"><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></a><div class="ck cz da r s"><div><div class="bt"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2Fa5a8e799ea2c&amp;operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fin-house-llm-serving-at-netflix-a5a8e799ea2c" rel="noopener follow"></a></div></div></div></div></div></div></div></div></div></div></div></div></div></article><article aria-label="8 Pillars of Modern Data Engineering" class="al" data-testid="post-preview"><div class="al rt e"><div class="cs"><div class="al e"><div class="dm al we wf wg wh wi wj wk wl wm wn wo wp wq"><div class="wr"><div aria-label="8 Pillars of Modern Data Engineering"><div class="wt wu wv ww ek"><img alt="8 Pillars of Modern Data Engineering" class="cs wx wy wz xa" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*NfEnDY7I02i6clqjM2kQSQ.png" /></div></div></div><div class="ws am ew gt"><div class="am gt qs cs xb xc xd xe"><div class="xf xg xh xi xj fd am j"><div class="fd am j xk"><div class="xl e ej"><div><div class="e"><div tabindex="-1" class="cq"><a tabindex="-1" href="https://medium.com/@bvsarathc06" rel="noopener follow"><div class="e dm"><img alt="B V Sarath Chandra" class="e bs ee xm xn ei" src="https://miro.medium.com/v2/resize:fill:40:40/1*8WdpRZsuzsiKXa8IWvf22A@2x.jpeg" width="20" height="20" /></div></a></div></div></div></div><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai jz am j" href="https://medium.com/@bvsarathc06" rel="noopener follow"><p class="z b y u dv xo es xp xq xr xs xt cv">B V Sarath Chandra</p></a></div></div></div></div></div><div class="mk am j ej db dd"><p class="z b y u w">·</p><p class="z b y u w">4d ago</p></div></div><div class="xv e xw xx xy xz ya gz"><div class="yb yc yd ye yf yg yh yi"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/@bvsarathc06/8-pillars-of-modern-data-engineering-5b36870c729f" rel="noopener follow"><div title=""><h2 class="z ii tn to yj yk yl ym yn yo fl tp tq yp yq yr ys yt yu fn fo fp yv yw yx yy yz za fq fr fs zb zc zd ze zf zg ft fu fv zh zi zj zk zl zm fw xt cv">8 Pillars of Modern Data Engineering</h2></div><div class="zn e"><h3 class="z b fz u dv zo es xp zp xr xt w">Tools come and go. Architectural principles stay forever.</h3></div></a></div></div><div class="su am je ce"><div class="am j jp"><div class="rt am"><button class="e cf af ac" aria-label="Member-only story"><div class=""><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></div></button></div><div class="ck cz da r s"><div class="dm zq zr am j"><div class="dr mf zs am j jp"><a class="dr gg zs am j jp" tabindex="-1" href="https://medium.com/@bvsarathc06/8-pillars-of-modern-data-engineering-5b36870c729f" rel="noopener follow"><div><div class="am"><div tabindex="-1" class="cq"><div class="am j db">A response icon3</div></div></div></div></a></div></div></div><div class="am j zu zv"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fmodeling-device-capabilities-for-analytics-e7607acebde8" rel="noopener follow"><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></a><div class="ck cz da r s"><div><div class="bt"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2F5b36870c729f&amp;operation=register&amp;redirect=https%3A%2F%2Fmedium.com%2F%40bvsarathc06%2F8-pillars-of-modern-data-engineering-5b36870c729f" rel="noopener follow"></a></div></div></div></div></div></div><div class="da r s"></div></div></div></div></div></div></div></div></article><article aria-label="Self-Service Analytics, According to Anthropic" class="al" data-testid="post-preview"><div class="al rt e"><div class="cs"><div class="al e"><div class="dm al we wf wg wh wi wj wk wl wm wn wo wp wq"><div class="wr"><div aria-label="Self-Service Analytics, According to Anthropic"><div class="wt wu wv ww ek"><img alt="Self-Service Analytics, According to Anthropic" class="cs wx wy acg xa" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*OkIQbvesGUYpHyXVnmIpxw.png" /></div></div></div><div class="ws am ew gt"><div class="am gt qs cs xb xc xd xe"><div class="xf xg xh xi xj fd am j"><div class="fd am j xk"><div class="xl e ej"><div><div class="e"><div tabindex="-1" class="cq"><a tabindex="-1" href="https://annujackson.medium.com" rel="noopener follow"><div class="e dm"><img alt="Ann Jackson" class="e bs ee xm xn ei" src="https://miro.medium.com/v2/resize:fill:40:40/1*Mzv4x0Nkwi5gS-7iezL42Q.png" width="20" height="20" /></div></a></div></div></div></div><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai jz am j" href="https://annujackson.medium.com" rel="noopener follow"><p class="z b y u dv xo es xp xq xr xs xt cv">Ann Jackson</p><div class="ach aci e"><div class="am acj"><div class="am"></div></div></div></a></div></div></div></div></div><div class="mk am j ej db dd"><p class="z b y u w">·</p><p class="z b y u w">Jun 19</p></div></div><div class="xv e xw xx xy xz ya gz"><div class="yb yc yd ye yf yg yh yi"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://annujackson.medium.com/self-service-analytics-according-to-anthropic-722d8fbceab9" rel="noopener follow"><div title=""><h2 class="z ii tn to yj yk yl ym yn yo fl tp tq yp yq yr ys yt yu fn fo fp yv yw yx yy yz za fq fr fs zb zc zd ze zf zg ft fu fv zh zi zj zk zl zm fw xt cv">Self-Service Analytics, According to Anthropic</h2></div><div class="zn e"><h3 class="z b fz u dv zo es xp zp xr xt w">An AI company mapped the path to agentic self-service analytics; 79% of it is human judgment</h3></div></a></div></div><div class="su am je ce"><div class="am j jp"><div class="rt am"><button class="e cf af ac" aria-label="Member-only story"><div class=""><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></div></button></div><div class="ck cz da r s"><div class="dm zq zr am j"><div class="dr mf zs am j jp"><a class="dr gg zs am j jp" tabindex="-1" href="https://annujackson.medium.com/self-service-analytics-according-to-anthropic-722d8fbceab9" rel="noopener follow"><div><div class="am"><div tabindex="-1" class="cq"><div class="am j db">A response icon6</div></div></div></div></a></div></div></div><div class="am j zu zv"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fmodeling-device-capabilities-for-analytics-e7607acebde8" rel="noopener follow"><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></a><div class="ck cz da r s"><div><div class="bt"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2F722d8fbceab9&amp;operation=register&amp;redirect=https%3A%2F%2Fannujackson.medium.com%2Fself-service-analytics-according-to-anthropic-722d8fbceab9" rel="noopener follow"></a></div></div></div></div></div></div><div class="da r s"></div></div></div></div></div></div></div></div></article><article aria-label="30 Core Agentic Engineering Concepts Every Developer Should Know" class="al" data-testid="post-preview"><div class="al rt e"><div class="cs"><div class="al e"><div class="dm al we wf wg wh wi wj wk wl wm wn wo wp wq"><div class="wr"><div aria-label="30 Core Agentic Engineering Concepts Every Developer Should Know"><div class="wt wu wv ww ek"><img alt="30 Core Agentic Engineering Concepts Every Developer Should Know" class="cs wx wy wz xa" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*GbRct1qcyZBLBX1XCdTl-Q.png" /></div></div></div><div class="ws am ew gt"><div class="am gt qs cs xb xc xd xe"><div class="xf xg xh xi xj fd am j"><div class="fd am j xk"><div class="xl e ej"><div><div class="e"><div tabindex="-1" class="cq"><a href="https://medium.com/lets-code-future" rel="noopener follow"><div class="dm"><img alt="Let’s Code Future" class="ei gv e xn xm" src="https://miro.medium.com/v2/resize:fill:40:40/1*QXfeVFVbIzUGnlwXoOZvyQ.png" width="20" height="20" /></div></a></div></div></div></div><div class="td e ej dd"><p class="z b y u w">In</p></div><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai jz am j" href="https://medium.com/lets-code-future" rel="noopener follow"><p class="z b y u dv xo es xp xq xr xs xt cv">Let’s Code Future</p></a></div></div></div></div></div><div class="xu e ej dd"><p class="z b y u w">by</p></div><div class="fd am j xk"><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai jz am j" href="https://medium.com/@Deep-concept" rel="noopener follow"><p class="z b y u dv xo es xp xq xr xs xt cv">Deep concept</p></a></div></div></div></div></div><div class="mk am j ej db dd"><p class="z b y u w">·</p><p class="z b y u w">Jun 21</p></div></div><div class="xv e xw xx xy xz ya gz"><div class="yb yc yd ye yf yg yh yi"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/lets-code-future/30-core-agentic-engineering-concepts-every-developer-should-know-5066b3117f69" rel="noopener follow"><div title=""><h2 class="z ii tn to yj yk yl ym yn yo fl tp tq yp yq yr ys yt yu fn fo fp yv yw yx yy yz za fq fr fs zb zc zd ze zf zg ft fu fv zh zi zj zk zl zm fw xt cv">30 Core Agentic Engineering Concepts Every Developer Should Know</h2></div><div class="zn e"><h3 class="z b fz u dv zo es xp zp xr xt w">A simple guide to AI agents, tools, memory, multi-agent systems, and how to build them safely</h3></div></a></div></div><div class="su am je ce"><div class="am j jp"><div class="rt am"><button class="e cf af ac" aria-label="Member-only story"><div class=""><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></div></button></div><div class="ck cz da r s"><div class="dm zq zr am j"><div class="dr mf zs am j jp"><a class="dr gg zs am j jp" tabindex="-1" href="https://medium.com/lets-code-future/30-core-agentic-engineering-concepts-every-developer-should-know-5066b3117f69" rel="noopener follow"><div><div class="am"><div tabindex="-1" class="cq"><div class="am j db">A response icon40</div></div></div></div></a></div></div></div><div class="am j zu zv"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fmodeling-device-capabilities-for-analytics-e7607acebde8" rel="noopener follow"><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></a><div class="ck cz da r s"><div><div class="bt"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2F5066b3117f69&amp;operation=register&amp;redirect=https%3A%2F%2Fmedium.com%2Flets-code-future%2F30-core-agentic-engineering-concepts-every-developer-should-know-5066b3117f69" rel="noopener follow"></a></div></div></div></div></div></div><div class="da r s"></div></div></div></div></div></div></div></div></article><article aria-label="11 Charts the AI Industry Doesn’t Want You to See" class="al" data-testid="post-preview"><div class="al rt e"><div class="cs"><div class="al e"><div class="dm al we wf wg wh wi wj wk wl wm wn wo wp wq"><div class="wr"><div aria-label="11 Charts the AI Industry Doesn’t Want You to See"><div class="wt wu wv ww ek"><img alt="11 Charts the AI Industry Doesn’t Want You to See" class="cs wx wy wz xa" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/0*CR0xu69pl2ve2--9.png" /></div></div></div><div class="ws am ew gt"><div class="am gt qs cs xb xc xd xe"><div class="xf xg xh xi xj fd am j"><div class="fd am j xk"><div class="xl e ej"><div><div class="e"><div tabindex="-1" class="cq"><a tabindex="-1" href="https://albertoromgar.medium.com" rel="noopener follow"><div class="e dm"><img alt="Alberto Romero" class="e bs ee xm xn ei" src="https://miro.medium.com/v2/resize:fill:40:40/1*_BorRAHo8o40sBLbZ7VE3Q.jpeg" width="20" height="20" /></div></a></div></div></div></div><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai jz am j" href="https://albertoromgar.medium.com" rel="noopener follow"><p class="z b y u dv xo es xp xq xr xs xt cv">Alberto Romero</p></a></div></div></div></div></div><div class="mk am j ej db dd"><p class="z b y u w">·</p><p class="z b y u w">Jul 14</p></div></div><div class="xv e xw xx xy xz ya gz"><div class="yb yc yd ye yf yg yh yi"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://albertoromgar.medium.com/11-charts-the-ai-industry-doesnt-want-you-to-see-55a73805cf5b" rel="noopener follow"><div title=""><h2 class="z ii tn to yj yk yl ym yn yo fl tp tq yp yq yr ys yt yu fn fo fp yv yw yx yy yz za fq fr fs zb zc zd ze zf zg ft fu fv zh zi zj zk zl zm fw xt cv">11 Charts the AI Industry Doesn’t Want You to See</h2></div><div class="zn e"><h3 class="z b fz u dv zo es xp zp xr xt w">This is what the AI story looks like</h3></div></a></div></div><div class="su am je ce"><div class="am j jp"><div class="rt am"><button class="e cf af ac" aria-label="Member-only story"><div class=""><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></div></button></div><div class="ck cz da r s"><div class="dm zq zr am j"><div class="dr mf zs am j jp"><a class="dr gg zs am j jp" tabindex="-1" href="https://albertoromgar.medium.com/11-charts-the-ai-industry-doesnt-want-you-to-see-55a73805cf5b" rel="noopener follow"><div><div class="am"><div tabindex="-1" class="cq"><div class="am j db">A response icon141</div></div></div></div></a></div></div></div><div class="am j zu zv"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fmodeling-device-capabilities-for-analytics-e7607acebde8" rel="noopener follow"><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></a><div class="ck cz da r s"><div><div class="bt"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2F55a73805cf5b&amp;operation=register&amp;redirect=https%3A%2F%2Falbertoromgar.medium.com%2F11-charts-the-ai-industry-doesnt-want-you-to-see-55a73805cf5b" rel="noopener follow"></a></div></div></div></div></div></div></div></div></div></div></div></div></div></article>]]></description>
      <link>https://netflixtechblog.com/modeling-device-capabilities-for-analytics-e7607acebde8</link>
      <guid>https://netflixtechblog.com/modeling-device-capabilities-for-analytics-e7607acebde8</guid>
      <pubDate>Fri, 31 Jul 2026 18:01:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[GenRec: Towards LLM-Native Recommendation at Netflix]]></title>
      <description><![CDATA[<article><div class="e"><div class="e"><section><div><div><div class="am ew"><div class="fd cs iz ja jb jc"></div></div><div class="it ju jv jw jx"><div class="am ew"><div class="fd cs iz ja jb jc"><div><div></div><p id="8e03" class="pw-post-body-paragraph pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px it cv">Authors: <a class="bx py" href="https://www.linkedin.com/in/ying-li-94016850/" rel="noopener ugc nofollow" target="_blank">Ying Li</a>, <a class="bx py" href="https://www.linkedin.com/in/arjunra0/" rel="noopener ugc nofollow" target="_blank">Arjun Rao</a>, <a class="bx py" href="https://www.linkedin.com/in/shradha-sehgal/" rel="noopener ugc nofollow" target="_blank">Shradha Sehgal</a></p><h2 id="41c3" class="pz qa ka z qb qc qd qe fl qf qg qh fn qi qj qk ql qm qn qo qp qq qr qs qt qu cv">Introduction</h2><p id="feec" class="pw-post-body-paragraph pd pe ka pf b pg qv pi pj pk qw pm pn fo qx pp pq fr qy ps pt fu qz pv pw px it cv">Recommendations sit at the heart of the Netflix experience. Our current production models rely on thousands of hand‑crafted features over users, items, and interactions, along with specialized architectures for sequence modeling, feature interactions, and multi‑task objectives. This stack has evolved over many years to support diverse content types (movies, series, games, live, podcasts) and product surfaces, but its complexity makes it costly to onboard new use cases: adding a content type or surface can require significant feature engineering, architecture change, infrastructure work, and experimentation.</p><p id="770f" class="pw-post-body-paragraph pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px it cv">At the same time, large language models (LLMs) are changing how we think about recommendation, as shown by recent work such as <a class="bx py" href="https://arxiv.org/abs/2510.07784" rel="noopener ugc nofollow" target="_blank">PLUM</a>, <a class="bx py" href="https://arxiv.org/abs/2603.17540" rel="noopener ugc nofollow" target="_blank">GLIDE</a>, and <a class="bx py" href="https://arxiv.org/abs/2510.11639" rel="noopener ugc nofollow" target="_blank">OneRec-Think</a>. Their broad world knowledge and strong language understanding make it possible to represent user histories and item metadata directly as text, capture rich relationships in a shared semantic space, and steer recommendations via natural‑language prompts. However, off‑the‑shelf LLMs are still far from production‑ready recommenders: they often over‑recommend globally popular content, hallucinate out‑of‑catalog items, ignore business constraints, and provide only limited personalization.</p><p id="f5e2" class="pw-post-body-paragraph pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px it cv">To address this, we built <strong class="pf kb">GenRec</strong>, an LLM‑backed recommendation ranker that post‑trains an internal foundation LLM on Netflix‑specific data and objectives. GenRec shows that an LLM‑based ranker can match or exceed a mature production system while relying on far fewer labeled examples and input signals.</p><figure class="rd re rf rg rh ri ra rb paragraph-image"><div role="button" tabindex="0" class="rj rk dm rl cs rm">Press enter or click to view image in full size<div class="ra rb rc"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*7Jb6-3knD_VCB_wWN5M3Mw.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*7Jb6-3knD_VCB_wWN5M3Mw.png" /><img alt="" class="cs er rn c" width="700" height="269" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="ro bv rp ra rb rq rr z b cr u w">Figure 1: GenRec pipeline. Raw logs of user history, item metadata, and context are transformed via context engineering into natural-language prompts and fed into the GenRec, which runs on vLLM in prefill-only mode and outputs scores for each catalog item, yielding a recommendation ranking.</figcaption></figure><p id="3d00" class="pw-post-body-paragraph pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px it cv">At a high level, GenRec:</p><ul class=""><li id="03fd" class="pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px rs rt ru cv">Verbalizes user histories, item metadata, and context as text.</li><li id="2e78" class="pd pe ka pf b pg rv pi pj pk rw pm pn fo rx pp pq fr ry ps pt fu rz pv pw px rs rt ru cv">Post‑trains a Netflix‑adapted foundation LLM for ranking.</li><li id="374c" class="pd pe ka pf b pg rv pi pj pk rw pm pn fo rx pp pq fr ry ps pt fu rz pv pw px rs rt ru cv">Adds a catalog‑aware scoring head over Netflix titles.</li><li id="e0d5" class="pd pe ka pf b pg rv pi pj pk rw pm pn fo rx pp pq fr ry ps pt fu rz pv pw px rs rt ru cv">Uses reward signals to align with long‑term member value and business goals.</li><li id="6287" class="pd pe ka pf b pg rv pi pj pk rw pm pn fo rx pp pq fr ry ps pt fu rz pv pw px rs rt ru cv">Runs in prefill‑only mode on Netflix’s LLM serving stack for cost efficiency.</li></ul><p id="cc36" class="pw-post-body-paragraph pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px it cv">In a large‑scale A/B test against a well‑tuned production ranker, GenRec achieves statistically significant improvements in both short‑term and long‑term online metrics, while using only a small fraction of the Phase‑2 labeled data and input signals. It reduces our reliance on hand‑engineered features and shifts the focus from <strong class="pf kb">feature engineering</strong> to <strong class="pf kb">context engineering</strong>. In this blog post, we will describe how GenRec works, how it performs, and why we believe it points toward a more LLM‑centric future for recommendation at Netflix.</p><h2 id="493b" class="pz qa ka z qb qc qd qe fl qf qg qh fn qi qj qk ql qm qn qo qp qq qr qs qt qu cv">Problem Setting</h2><p id="8a3d" class="pw-post-body-paragraph pd pe ka pf b pg qv pi pj pk qw pm pn fo qx pp pq fr qy ps pt fu qz pv pw px it cv">We focus on a <strong class="pf kb">full‑catalog ranking</strong> <strong class="pf kb">task</strong> (or top‑<em class="sa">K</em> ranking when a candidate set is provided).</p><p id="4776" class="pw-post-body-paragraph pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px it cv">Given a user 푢, their interaction history 퐻, and the current context 휏 (device, surface, locale, time, etc.), GenRec scores each item and produces a personalized ranking that can directly power recommendations or serve as input for downstream personalization systems.</p><p id="576c" class="pw-post-body-paragraph pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px it cv">Formally, we map a request (<em class="sa">u</em>,<em class="sa">τ</em>,<em class="sa">t</em>,<em class="sa">H</em>) — user, context, time, and history — to a ranking 휋 over the catalog <em class="sa">C</em>, where <em class="sa">π(i)</em> is the position assigned to item <em class="sa">i</em>. We optimize <em class="sa">π</em> for <strong class="pf kb">expected long‑term member utility</strong> (a proxy for satisfaction and retention), not just short‑term engagements.</p><h2 id="fd49" class="pz qa ka z qb qc qd qe fl qf qg qh fn qi qj qk ql qm qn qo qp qq qr qs qt qu cv">From Foundation LLM to Recommendation Ranker</h2><p id="f003" class="pw-post-body-paragraph pd pe ka pf b pg qv pi pj pk qw pm pn fo qx pp pq fr qy ps pt fu qz pv pw px it cv">GenRec follows a <strong class="pf kb">two‑phase training framework</strong> (Figure 2):</p><figure class="rd re rf rg rh ri ra rb paragraph-image"><div role="button" tabindex="0" class="rj rk dm rl cs rm">Press enter or click to view image in full size<div class="ra rb sb"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*2Gpp9wQ1MBAqeAZ0ojd1ew.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*2Gpp9wQ1MBAqeAZ0ojd1ew.png" /><img alt="" class="cs er rn c" width="700" height="156" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="ro bv rp ra rb rq rr z b cr u w">Figure 2: Two Phase Framework. Phase 1 trains a foundational LLM on Netflix data for user and content understanding, and Phase 2 post-trains on ranking-specific data and objectives.</figcaption></figure><h3 id="8c24" class="sc qa ka z qb fk sd ao fl fm se aq fn fo sf fp fq fr sg fs ft fu sh fv fw si cv"><strong class="cb">Phase 1 — Netflix-Adapted</strong> <strong class="cb">Foundation LLM</strong>.</h3><p id="8ae6" class="pw-post-body-paragraph pd pe ka pf b pg qv pi pj pk qw pm pn fo qx pp pq fr qy ps pt fu qz pv pw px it cv">We start from an open‑source LLM and adapt it on proprietary Netflix corpora, so it learns foundational capabilities such as</p><ul class=""><li id="4ca7" class="pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px rs rt ru cv">Netflix content understanding</li><li id="933f" class="pd pe ka pf b pg rv pi pj pk rw pm pn fo rx pp pq fr ry ps pt fu rz pv pw px rs rt ru cv">Member behavior and preference patterns</li><li id="478e" class="pd pe ka pf b pg rv pi pj pk rw pm pn fo rx pp pq fr ry ps pt fu rz pv pw px rs rt ru cv">General language understanding and generation.</li></ul><p id="1339" class="pw-post-body-paragraph pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px it cv">Phase 1 is updated relatively infrequently and serves as a shared, Netflix‑aware backbone for many applications.</p><h3 id="dd0f" class="sc qa ka z qb fk sd ao fl fm se aq fn fo sf fp fq fr sg fs ft fu sh fv fw si cv"><strong class="cb">Phase 2</strong> — <strong class="cb">GenRec</strong>.</h3><p id="ec9b" class="pw-post-body-paragraph pd pe ka pf b pg qv pi pj pk qw pm pn fo qx pp pq fr qy ps pt fu qz pv pw px it cv">We then turn this foundation model into a high‑quality ranking model by post‑training on ranking‑specific data and objectives. Phase 2:</p><ul class=""><li id="f1ee" class="pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px rs rt ru cv">Focuses on ranking quality and steering</li><li id="a9e0" class="pd pe ka pf b pg rv pi pj pk rw pm pn fo rx pp pq fr ry ps pt fu rz pv pw px rs rt ru cv">Incorporates multiple reward signals via reward‑weighted losses</li><li id="a663" class="pd pe ka pf b pg rv pi pj pk rw pm pn fo rx pp pq fr ry ps pt fu rz pv pw px rs rt ru cv">Is refreshed more frequently to track new content and evolving tastes</li><li id="820b" class="pd pe ka pf b pg rv pi pj pk rw pm pn fo rx pp pq fr ry ps pt fu rz pv pw px rs rt ru cv">Is explicitly optimized under serving cost constraints.</li></ul><h2 id="d199" class="pz qa ka z qb qc qd qe fl qf qg qh fn qi qj qk ql qm qn qo qp qq qr qs qt qu cv">Training Data as Conversations</h2><p id="249e" class="pw-post-body-paragraph pd pe ka pf b pg qv pi pj pk qw pm pn fo qx pp pq fr qy ps pt fu qz pv pw px it cv">Netflix members generate <strong class="pf kb">hundreds of billions </strong>of interaction events spanning many surfaces (views, plays, durations, thumbs up/down, add to list, abandons, etc.). We convert this log data into <strong class="pf kb">single‑turn or multi‑turn “conversations”</strong> between a user and a recommender. Each turn contains:</p><ul class=""><li id="4d62" class="pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px rs rt ru cv"><strong class="pf kb">User message</strong>: verbalized context, profile, history, item metadata, and task (e.g., recommend what the user will watch or thumb next).</li><li id="52db" class="pd pe ka pf b pg rv pi pj pk rw pm pn fo rx pp pq fr ry ps pt fu rz pv pw px rs rt ru cv"><strong class="pf kb">Assistant message</strong>: the member’s actual engagement (e.g., which titles were played, for how long, what feedback they provided).</li></ul><p id="d4c6" class="pw-post-body-paragraph pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px it cv">During Phase‑2 training, the LLM learns how assistant messages depend on user messages. This allows us to express rich recommendation signals as text, jointly supporting both the language-modeling (LM) and ranking objectives.</p><p id="c9da" class="pw-post-body-paragraph pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px it cv">At inference time, we feed in the verbalized context and apply a catalog‑aware scoring head to rank items; we do not decode assistant messages. The conversational format is primarily used during training to support the LM objective and preserve strong language understanding over the verbalized text.</p><h2 id="0a55" class="pz qa ka z qb qc qd qe fl qf qg qh fn qi qj qk ql qm qn qo qp qq qr qs qt qu cv">Verbalization and Context Engineering</h2><p id="6c90" class="pw-post-body-paragraph pd pe ka pf b pg qv pi pj pk qw pm pn fo qx pp pq fr qy ps pt fu qz pv pw px it cv">Traditional recommenders operate on dense features and embeddings. GenRec takes a different approach: it <strong class="pf kb">verbalizes</strong> rich user histories and context as natural language, encoding raw interaction signals directly in the LLM’s semantic space. In doing so, it relies on the model to discover higher‑level patterns — such as item relationships and evolving user interests — rather than on manual feature engineering.</p><p id="41aa" class="pw-post-body-paragraph pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px it cv">Naively verbalizing every interaction in a user’s history can quickly exceed the token budget and be too expensive at Netflix scale. The context window becomes our new “feature budget”, so we apply <a class="bx py" href="https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents" rel="noopener ugc nofollow" target="_blank"><strong class="pf kb">context</strong> <strong class="pf kb">engineering</strong></a><strong class="pf kb">:</strong></p><ul class=""><li id="1396" class="pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px rs rt ru cv"><strong class="pf kb">Retain in full:</strong> high‑signal engagements (e.g., long plays, thumbs‑up) with richer details</li><li id="f36f" class="pd pe ka pf b pg rv pi pj pk rw pm pn fo rx pp pq fr ry ps pt fu rz pv pw px rs rt ru cv"><strong class="pf kb">Omit: </strong>low‑signal events (e.g., very short plays or quick hovers)</li><li id="8b85" class="pd pe ka pf b pg rv pi pj pk rw pm pn fo rx pp pq fr ry ps pt fu rz pv pw px rs rt ru cv"><strong class="pf kb">Summarize or compress: </strong>repetitive behaviors (e.g., binge‑watching )</li><li id="6a7e" class="pd pe ka pf b pg rv pi pj pk rw pm pn fo rx pp pq fr ry ps pt fu rz pv pw px rs rt ru cv"><strong class="pf kb">Elaborate selectively:</strong> important or cold‑start items (e.g., new releases)</li></ul><p id="06de" class="pw-post-body-paragraph pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px it cv">Within a fixed token budget, we prioritize recent, high‑signal history and compress or drop older history. We also structure the prompt to maximize shared prefixes for better prefix caching. The goal is a <strong class="pf kb">compact, high‑information prompt</strong> that preserves ranking quality without prohibitive costs.</p><h2 id="d433" class="pz qa ka z qb qc qd qe fl qf qg qh fn qi qj qk ql qm qn qo qp qq qr qs qt qu cv">Objectives: Ranking, Language, and Rewards</h2><p id="ff57" class="pw-post-body-paragraph pd pe ka pf b pg qv pi pj pk qw pm pn fo qx pp pq fr qy ps pt fu qz pv pw px it cv">The overall GenRec model is trained with a <strong class="pf kb">multi‑objective loss </strong>that combines a recommendation ranking objective, language modeling objectives, and alignment via reward‑weighted training.</p><h3 id="cf9b" class="sc qa ka z qb fk sd ao fl fm se aq fn fo sf fp fq fr sg fs ft fu sh fv fw si cv">1. Catalog‑Aware Ranking Objective</h3><p id="bb14" class="pw-post-body-paragraph pd pe ka pf b pg qv pi pj pk qw pm pn fo qx pp pq fr qy ps pt fu qz pv pw px it cv">The primary task is a <strong class="pf kb">ranking objective</strong> that teaches the model to score items by engagement quality. We label positives using high‑value engagements (e.g., sufficiently long plays, strong explicit feedback), with thresholds and denoising logic, and train the model — via a cross‑entropy loss over the catalog or candidate set — to assign higher scores to these positives given a verbalized context.</p><h3 id="6ec2" class="sc qa ka z qb fk sd ao fl fm se aq fn fo sf fp fq fr sg fs ft fu sh fv fw si cv">2. Language Modeling Objective</h3><p id="529b" class="pw-post-body-paragraph pd pe ka pf b pg qv pi pj pk qw pm pn fo qx pp pq fr qy ps pt fu qz pv pw px it cv">We also retain a <strong class="pf kb">language modeling (LM)</strong> objective over the verbalized inputs and outputs. This helps preserve the model’s general language understanding, improves its ability to interpret rich natural‑language histories and item metadata, and keeps the door open for text‑generation use cases such as recommendation explanations.</p><h3 id="b983" class="sc qa ka z qb fk sd ao fl fm se aq fn fo sf fp fq fr sg fs ft fu sh fv fw si cv">3. Reward‑Weighted Loss for Alignment</h3><p id="8c37" class="pw-post-body-paragraph pd pe ka pf b pg qv pi pj pk qw pm pn fo qx pp pq fr qy ps pt fu qz pv pw px it cv">Beyond raw ranking accuracy, GenRec must (1) respect business requirements — for example, balancing movies, series, games, live, and podcasts — and (2) optimize long‑term member satisfaction rather than just immediate clicks or plays.</p><div class="sj et sk sl am j gh"><div class="sm sn so sp sq e bv"><h2 class="sr b ss st su sv sw">Get Netflix Technology Blog’s stories in your inbox</h2></div><div class="sx e bv"><p class="z b cr u w">Join Medium for free to get updates from this writer.</p></div><div class="tb tc td te tf tg am"><div class="th am gh"><div class="am th br bq jf ti tj ei go tk tl tm tn to tp tq"><input class="tr ca by cb ts tt dx nl cw hp cs" placeholder="Enter your email" type="text" value="" /></div></div><div class="ck cz tu ej"><div class="sy sz ta"><button class="z b cr u tv tw tx ty tz ua ub bk bl uc ud ue bp bq br bs bt bu bv">Subscribe</button></div></div><div class="cs da r s"><button class="z b cr u tv tw tx ty tz ua ub bk bl uc ud ue bp cs bq br bs bt bu bv">Subscribe</button></div></div><div class="uf e"><label class="am j"><input class="ug uh ui dr hd uj uk ul um un uo up" type="checkbox" checked="checked" /></label></div><div class="e"><p class="z b cr u cv">Remember me for faster sign in</p></div></div><div class="ux e"></div><p id="9ad3" class="pw-post-body-paragraph pd pe ka pf b pg pi pj pk pm pn fo pp pq fr ps pt fu pv pw sl px it cv">Training only on raw interaction sequences can lead to undesirable behaviors, such as over‑favoring binge‑watching or over‑focusing on a single content type. To address this, we <strong class="pf kb">weight the ranking loss </strong>usingsignalsfrom separate <a class="bx py" href="https://dl.acm.org/doi/abs/10.1145/3604915.3608873" rel="noopener ugc nofollow" target="_blank">reward models</a>. Each training example receives a scalar weight derived from two types of signals:</p><ul class=""><li id="c449" class="pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px rs rt ru cv"><strong class="pf kb">Long‑term satisfaction proxies:</strong> estimate how much a short‑term engagement contributes to long‑term outcomes, such as return behavior, catalog exploration, or sustained engagement.</li><li id="1129" class="pd pe ka pf b pg rv pi pj pk rw pm pn fo rx pp pq fr ry ps pt fu rz pv pw px rs rt ru cv"><strong class="pf kb">Behavior rebalancing: </strong>adjust behaviors across content types and launch stages (for example, games vs. movies, new releases vs. evergreen titles) to better align with business goals.</li></ul><p id="85b4" class="pw-post-body-paragraph pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px it cv">The example’s ranking loss is then scaled by this weight: high‑value engagements receive larger weights, and low‑value ones are down‑weighted. This <strong class="pf kb">reward‑weighted approach</strong> is simpler and more cost-efficient than full reinforcement learning, yet provides effective alignment in practice. We have seen additional gains from RL‑style methods (e.g., GRPO), but leave them to future work due to their higher cost.</p><h2 id="9f95" class="pz qa ka z qb qc qd qe fl qf qg qh fn qi qj qk ql qm qn qo qp qq qr qs qt qu cv">Model Architecture and Serving</h2><h3 id="a5bc" class="sc qa ka z qb fk sd ao fl fm se aq fn fo sf fp fq fr sg fs ft fu sh fv fw si cv">Backbone and Scoring Head</h3><p id="0e00" class="pw-post-body-paragraph pd pe ka pf b pg qv pi pj pk qw pm pn fo qx pp pq fr qy ps pt fu qz pv pw px it cv">GenRec’s architecture closely follows our foundational LLM: a <strong class="pf kb">decoder‑only Transformer</strong> trained with next‑token‑prediction style objectives, augmented with a <strong class="pf kb">catalog‑aware ranking head</strong> that scores only Netflix in-catalog items. The scoring pipeline works as follows:</p><ol class=""><li id="d72c" class="pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px uy rt ru cv"><strong class="pf kb">Verbalization:</strong> A verbalizer <em class="sa">V</em> serializes user history <em class="sa">H</em>, context <em class="sa">휏</em> , and relevant item metadata into a single text sequence <em class="sa">x</em>.</li><li id="de4b" class="pd pe ka pf b pg rv pi pj pk rw pm pn fo rx pp pq fr ry ps pt fu rz pv pw px uy rt ru cv"><strong class="pf kb">Pooled representation:</strong> The LLM processes <em class="sa">x,</em> and we extract a pooled hidden state <em class="sa">h</em> that summarizes the user’s current preferences and context.</li><li id="6c38" class="pd pe ka pf b pg rv pi pj pk rw pm pn fo rx pp pq fr ry ps pt fu rz pv pw px uy rt ru cv"><strong class="pf kb">Catalog‑aware scoring:</strong> Each catalog item <em class="sa">i </em>has a learned embedding <em class="sa">eᵢ</em>. A scoring head<em class="sa"> ϕ</em> combines <em class="sa">h</em> and <em class="sa">eᵢ </em>(e.g., via dot product or small MLP) to produce a score s<em class="sa">ᵢ</em>. Applying a softmax over scores yields a probability distribution which we convert into a ranking<em class="sa"> π</em>.</li></ol><p id="87fb" class="pw-post-body-paragraph pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px it cv">All parameters — the backbone, scoring head, and item embeddings — are trained jointly. For very large catalogs, we can use sampled softmax or candidate sets for efficient training and inference. This architecture constrains recommendations to the Netflix catalog while supporting efficient scoring over large candidate sets.</p><h3 id="80cb" class="sc qa ka z qb fk sd ao fl fm se aq fn fo sf fp fq fr sg fs ft fu sh fv fw si cv">Serving and Cost Optimization</h3><p id="b877" class="pw-post-body-paragraph pd pe ka pf b pg qv pi pj pk qw pm pn fo qx pp pq fr qy ps pt fu qz pv pw px it cv">GenRec is served on <a class="bx py" rel="noopener ugc nofollow" href="https://netflixtechblog.com/in-house-llm-serving-at-netflix-a5a8e799ea2c" data-discover="true">Netflix’s internal LLM stack</a> using vLLM. At Netflix scale, serving cost is driven primarily by 1) Model size; 2) Context length; 3) Inference mode (prefill vs. autoregressive decoding). We control cost through three strategies:</p><ul class=""><li id="00d1" class="pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px rs rt ru cv"><strong class="pf kb">Smaller / distilled models:</strong> We train GenRec on smaller or distilled foundation models, often with larger or more targeted datasets, to capture most of the quality of larger models at lower serving cost.</li><li id="aec0" class="pd pe ka pf b pg rv pi pj pk rw pm pn fo rx pp pq fr ry ps pt fu rz pv pw px rs rt ru cv"><strong class="pf kb">Aggressive context compaction:</strong> Using the context engineering described earlier, we minimize tokens while preserving ranking quality.</li><li id="91e7" class="pd pe ka pf b pg rv pi pj pk rw pm pn fo rx pp pq fr ry ps pt fu rz pv pw px rs rt ru cv"><strong class="pf kb">Prefill‑only inference: </strong>Autoregressive decoding over large candidate sets would be prohibitively expensive. Instead, we run in prefill‑only mode: the model consumes the prompt once and scores the entire candidate set in a single forward pass, with no token‑by‑token decoding.</li></ul><p id="3329" class="pw-post-body-paragraph pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px it cv">Together, these choices make it feasible to serve GenRec on high‑volume workloads within compute budgets.</p><h2 id="5351" class="pz qa ka z qb qc qd qe fl qf qg qh fn qi qj qk ql qm qn qo qp qq qr qs qt qu cv">Offline and Online Experiments</h2><p id="99b7" class="pw-post-body-paragraph pd pe ka pf b pg qv pi pj pk qw pm pn fo qx pp pq fr qy ps pt fu qz pv pw px it cv">We evaluated GenRec against a <strong class="pf kb">mature production ranker </strong>that has been tuned over many years. The baseline model relies on thousands of engineered dense and embedding features, as well as custom architectures for modeling feature interactions and sequences. We assessed performance using both offline evaluation metrics and a large‑scale online A/B test.</p><h3 id="f03c" class="sc qa ka z qb fk sd ao fl fm se aq fn fo sf fp fq fr sg fs ft fu sh fv fw si cv">GenRec vs Production Baseline</h3><p id="fd04" class="pw-post-body-paragraph pd pe ka pf b pg qv pi pj pk qw pm pn fo qx pp pq fr qy ps pt fu qz pv pw px it cv"><strong class="pf kb">Offline</strong>, GenRec outperformed the production ranker on ranking metrics despite using far fewer input signals and labeled examples. With roughly <strong class="pf kb">40× fewer Phase‑2 labeled training examples</strong>, GenRec achieved about <strong class="pf kb">+1.6%</strong> <strong class="pf kb">improvement</strong> in Mean Reciprocal Rank (MRR). As we increased Phase‑2 training data and enriched the input signals, GenRec’s offline metrics continued to improve.</p><p id="297a" class="pw-post-body-paragraph pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px it cv"><strong class="pf kb">Online</strong>, we ran a large A/B test on batch‑compute recommendation surfaces, covering <strong class="pf kb">~10% of Netflix traffic</strong> over ~4 weeks. In this low‑data, low‑signal configuration, GenRec delivered <strong class="pf kb">statistically significant gains</strong> over the production baseline on both short‑term and long‑term online metrics (Figure 3).</p><p id="623b" class="pw-post-body-paragraph pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px it cv">These results indicate that a properly post‑trained and aligned LLM‑backed ranker can be a strong alternative to traditional recommendation models, with substantial headroom as we further scale data and input signals.</p><figure class="rd re rf rg rh ri ra rb paragraph-image"><div role="button" tabindex="0" class="rj rk dm rl cs rm">Press enter or click to view image in full size<div class="ra rb uz"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*Z8IJ48SEZuHf3rViZXKZsA.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*Z8IJ48SEZuHf3rViZXKZsA.png" /><img alt="" class="cs er rn c" width="700" height="385" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="ro bv rp ra rb rq rr z b cr u w">Figure 3: Online metrics of GenRec vs. production model. GenRec achieves statistically significant improvements on both short-term and long-term online metrics.</figcaption></figure><h3 id="f65c" class="sc qa ka z qb fk sd ao fl fm se aq fn fo sf fp fq fr sg fs ft fu sh fv fw si cv">Data, Model, and Phase Contributions</h3><p id="6f08" class="pw-post-body-paragraph pd pe ka pf b pg qv pi pj pk qw pm pn fo qx pp pq fr qy ps pt fu qz pv pw px it cv">We ran ablations to understand where GenRec’s gains come from.</p><p id="ec94" class="pw-post-body-paragraph pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px it cv"><strong class="pf kb">Data and Model Scaling</strong></p><ul class=""><li id="8f42" class="pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px rs rt ru cv"><strong class="pf kb">Data scaling: </strong>For both ~1B and ~10B parameter backbones, offline MRR improves as we increase Phase‑2 post‑training data. Larger models reach higher absolute MRR but follow a similar scaling curve (see Figure 4).</li><li id="8332" class="pd pe ka pf b pg rv pi pj pk rw pm pn fo rx pp pq fr ry ps pt fu rz pv pw px rs rt ru cv"><strong class="pf kb">Model scaling: </strong>Under a fixed training budget, we post‑trained GenRec variants from ~1B to ~10B parameters. Within this budget, larger backbones consistently achieved higher offline MRR than smaller ones.</li></ul><figure class="rd re rf rg rh ri ra rb paragraph-image"><div role="button" tabindex="0" class="rj rk dm rl cs rm">Press enter or click to view image in full size<div class="ra rb va"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*1EcVs9E3h7XFfypn7nBX7g.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*1EcVs9E3h7XFfypn7nBX7g.png" /><img alt="" class="cs er rn c" width="700" height="376" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="ro bv rp ra rb rq rr z b cr u w">Figure 4: GenRec Phase-2 data scaling for the∼10B model.</figcaption></figure><p id="38db" class="pw-post-body-paragraph pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px it cv"><strong class="pf kb">Phase-1 vs. OSS, Phase-2 vs. Phase-1</strong></p><ul class=""><li id="240c" class="pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px rs rt ru cv"><strong class="pf kb">Phase-1 vs. OSS:</strong> Using the Phase‑1 Netflix‑adapted foundation LLM as the base model improves offline ranking metrics by roughly <strong class="pf kb">10–20%</strong> compared to starting directly from an off‑the‑shelf LLM.</li><li id="0f68" class="pd pe ka pf b pg rv pi pj pk rw pm pn fo rx pp pq fr ry ps pt fu rz pv pw px rs rt ru cv"><strong class="pf kb">Phase-2 vs. Phase-1:</strong> Phase‑2 post‑training adds another <strong class="pf kb">35–50%</strong> gain in offline ranking metrics when evaluated near the Phase‑1 training cutoff (i.e. when Phase‑1 model is the freshest). As time passes and Phase‑1 becomes stale with new content and shifting tastes, the relative benefit of Phase 2 grows to about 80% after 2 weeks.</li></ul><figure class="rd re rf rg rh ri ra rb paragraph-image"><div role="button" tabindex="0" class="rj rk dm rl cs rm">Press enter or click to view image in full size<div class="ra rb vb"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*HEMkanM7bkzEMji36XkMoQ.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*HEMkanM7bkzEMji36XkMoQ.png" /><img alt="" class="cs er rn c" width="700" height="179" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="68ce" class="pw-post-body-paragraph pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px it cv"><strong class="pf kb">Data efficiency vs. production ranker</strong></p><ul class=""><li id="8c91" class="pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px rs rt ru cv">Starting from a strong Phase‑1 model, GenRec matches or exceeds the production ranker using <strong class="pf kb">10–40× fewer</strong> <strong class="pf kb">Phase‑2 labeled examples</strong>, depending on configuration. This marginal data efficiency is especially valuable because Phase 2 is refreshed far more frequently than Phase 1.</li></ul><h3 id="2228" class="sc qa ka z qb fk sd ao fl fm se aq fn fo sf fp fq fr sg fs ft fu sh fv fw si cv">Context Length Optimization</h3><p id="8c64" class="pw-post-body-paragraph pd pe ka pf b pg qv pi pj pk qw pm pn fo qx pp pq fr qy ps pt fu qz pv pw px it cv">Context length drives both <strong class="pf kb">quality</strong> and <strong class="pf kb">cost</strong>: longer verbalizations expose more behavior and context but increase training and serving cost. To study this trade‑off, we varied context length and verbosity and optimized them in three steps:</p><ol class=""><li id="2a10" class="pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px uy rt ru cv"><strong class="pf kb">Clean and compress events:</strong> drop low‑signal engagements and compress repetitive behavior to form a cleaned sequence of events.</li><li id="c388" class="pd pe ka pf b pg rv pi pj pk rw pm pn fo rx pp pq fr ry ps pt fu rz pv pw px uy rt ru cv"><strong class="pf kb">Find the “elbow point”:</strong> vary how many historical events we include and plot MRR vs. number of events to identify an elbow beyond which additional context yields diminishing returns (see Figure 5).</li><li id="9e89" class="pd pe ka pf b pg rv pi pj pk rw pm pn fo rx pp pq fr ry ps pt fu rz pv pw px uy rt ru cv"><strong class="pf kb">Optimize verbosity:</strong> for the retained events, test different levels of details and simplified wordings, measuring MRR each time.</li></ol><p id="b43e" class="pw-post-body-paragraph pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px it cv">In our experiments, we can reduce the context tokens to roughly <strong class="pf kb">one-third</strong> of the original budget with <strong class="pf kb">negligible degradation</strong> in offline ranking metrics. Since serving cost is approximately proportional to context length, we observed a similar reduction in serving cost.</p><figure class="rd re rf rg rh ri ra rb paragraph-image"><div role="button" tabindex="0" class="rj rk dm rl cs rm">Press enter or click to view image in full size<div class="ra rb va"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*W_u0qbjvFHQ7z9deBnyPsA.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*W_u0qbjvFHQ7z9deBnyPsA.png" /><img alt="" class="cs er rn c" width="700" height="403" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="ro bv rp ra rb rq rr z b cr u w">Figure 5: Offline ranking metric (MRR) vs. number of user engagement events included in the prompt. The dashed line marks the elbow point: increasing the number of events beyond this yields diminishing returns.</figcaption></figure><h2 id="e08f" class="pz qa ka z qb qc qd qe fl qf qg qh fn qi qj qk ql qm qn qo qp qq qr qs qt qu cv">Towards LLM‑Native Recommendation</h2><p id="09e0" class="pw-post-body-paragraph pd pe ka pf b pg qv pi pj pk qw pm pn fo qx pp pq fr qy ps pt fu qz pv pw px it cv">GenRec is more than “swapping in a Transformer” for an existing ranker. It hints at a broader shift toward <strong class="pf kb">LLM‑native recommendation</strong> at Netflix. A few notable changes:</p><h3 id="31e1" class="sc qa ka z qb fk sd ao fl fm se aq fn fo sf fp fq fr sg fs ft fu sh fv fw si cv">From Feature Engineering to Context Engineering</h3><p id="f429" class="pw-post-body-paragraph pd pe ka pf b pg qv pi pj pk qw pm pn fo qx pp pq fr qy ps pt fu qz pv pw px it cv">Traditional RecSys stacks revolve around large feature sets and heavy feature infrastructure. LLM‑centric systems instead revolve around <strong class="pf kb">constructing rich textual contexts</strong> from raw logs, metadata, and tools. The “prompt” becomes the new feature vector.</p><p id="30c9" class="pw-post-body-paragraph pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px it cv">Modeling effort shifts from designing features to deciding <strong class="pf kb">which signals to include, how far back in time to go, how to compress or summarize history within a token budget</strong>. Our experiments on verbalization compaction illustrate this shift: careful context design can preserve quality while dramatically reducing serving cost.</p><h3 id="1132" class="sc qa ka z qb fk sd ao fl fm se aq fn fo sf fp fq fr sg fs ft fu sh fv fw si cv">From Customized Architectures to Foundation Backbones</h3><p id="c949" class="pw-post-body-paragraph pd pe ka pf b pg qv pi pj pk qw pm pn fo qx pp pq fr qy ps pt fu qz pv pw px it cv">Historically, each recommendation task often had its own custom architecture (two‑tower models, DLRM‑style networks, bespoke attention blocks). In an LLM‑centric world, many tasks share a <strong class="pf kb">common foundation backbone</strong>, with differentiation coming from data and verbalization strategies, post‑training objectives and rewards, and inference optimization.</p><p id="28cb" class="pw-post-body-paragraph pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px it cv">GenRec leverages the same backbone as our foundation LLM rather than introducing a new architecture built from scratch. This makes it easier to share learnings across applications, and opens the door to <strong class="pf kb">natural‑language steering</strong> for future experiences.</p><h3 id="0a97" class="sc qa ka z qb fk sd ao fl fm se aq fn fo sf fp fq fr sg fs ft fu sh fv fw si cv">Scaling Laws as Design Guides</h3><p id="5f57" class="pw-post-body-paragraph pd pe ka pf b pg qv pi pj pk qw pm pn fo qx pp pq fr qy ps pt fu qz pv pw px it cv">Traditional RecSys can hit diminishing returns due to sparse IDs, heavy engineering objectives, and task‑specific architectures. With an LLM‑backed backbone, recommendation inherits clearer <strong class="pf kb">data and model scaling behavior</strong>: within cost limits, more data and larger models consistently improve quality. This brings RecSys design closer to the broader LLM paradigm, where <strong class="pf kb">scaling laws</strong> help guide model and data investment.</p><h3 id="7f97" class="sc qa ka z qb fk sd ao fl fm se aq fn fo sf fp fq fr sg fs ft fu sh fv fw si cv">From RecSys Infra to LLM Infra</h3><p id="b08e" class="pw-post-body-paragraph pd pe ka pf b pg qv pi pj pk qw pm pn fo qx pp pq fr qy ps pt fu qz pv pw px it cv">LLM‑backed recommenders push us toward <strong class="pf kb">LLM‑style infrastructure</strong>: GPU‑accelerated, vLLM/Triton‑based, with careful batching and caching. Over time, recommendation serving infra starts to look more like general LLM infra than classic RecSys stacks built around MLPs or factorization models.</p><h2 id="78eb" class="pz qa ka z qb qc qd qe fl qf qg qh fn qi qj qk ql qm qn qo qp qq qr qs qt qu cv">Conclusions</h2><p id="badd" class="pw-post-body-paragraph pd pe ka pf b pg qv pi pj pk qw pm pn fo qx pp pq fr qy ps pt fu qz pv pw px it cv">We have presented <strong class="pf kb">GenRec</strong>, an LLM‑backed recommendation ranker at Netflix that adapts an internal foundation LLM for large‑scale personalization. By verbalizing user histories, context, and item metadata, adding a catalog‑aware ranking head, using reward‑weighted objectives aligned to long‑term satisfaction and business goals, and serving efficiently on our LLM infrastructure, we obtain a model that <strong class="pf kb">improves on a strong production ranker while using far fewer Phase‑2 labels and input signals</strong>.</p><p id="cd4d" class="pw-post-body-paragraph pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px it cv">GenRec is an early but promising step toward a more <strong class="pf kb">LLM‑centric recommendation stack </strong>at Netflix. Our results suggest that, with careful attention to cost, infrastructure, and alignment, LLM‑backed recommenders can play a central role in large‑scale personalization.</p><h2 id="c8a7" class="pz qa ka z qb qc qd qe fl qf qg qh fn qi qj qk ql qm qn qo qp qq qr qs qt qu cv">Acknowledgments</h2><p id="3a98" class="pw-post-body-paragraph pd pe ka pf b pg qv pi pj pk qw pm pn fo qx pp pq fr qy ps pt fu qz pv pw px it cv">GenRec is the result of close collaboration among multiple teams and organizations across Netflix. The contributors to this work (in alphabetical order):</p><p id="7813" class="pw-post-body-paragraph pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px it cv"><strong class="pf kb">AI for members: </strong>Arjun Rao, Ashish Rastogi, Baolin Li, Fernando Amat Gil, Grace Huang, Justin Basilico, Kamelia Aryafar, Linas Baltrunas, Moumita Bhattacharya, Ogheneovo Dibie, Rein Houthooft, Shradha Sehgal, Sejoon Oh, Sergi Perez, Sourabh Medapati, Thea Wang, Yaochen Zhu, Yesu Feng, Ying Li, Yun Li, Yucheng Shi, Yunan Hu</p><p id="20f1" class="pw-post-body-paragraph pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px it cv"><strong class="pf kb">AI platform and serving: </strong>Abhishek Agrawal, Adam Singer, Binh Tang, Daneo Zhang, Derek Olejnik, Ed Maddox, Erik Osheim, Lingyi Liu, Liping Peng, Meghana Chilukuri, Nicolas Hortiguera, Shaojing Li, ZQ Zhang</p><p id="2864" class="pw-post-body-paragraph pd pe ka pf b pg ph pi pj pk pl pm pn fo po pp pq fr pr ps pt fu pu pv pw px it cv"><strong class="pf kb">Product: </strong>Ilke Kaya, Michelle Kislak, Scarlet Chen, Si Cheng</p></div></div></div></div></div></div></section></div></div></article><article aria-label="In-House LLM Serving at Netflix" class="al" data-testid="post-preview"><div class="al vi e"><div class="cs"><div class="al e"><div class="dm al abl abm abn abo abp abq abr abs abt abu abv abw abx"><div class="aby"><div aria-label="In-House LLM Serving at Netflix"><div class="aca acb acc acd ek"><img alt="In-House LLM Serving at Netflix" class="cs ace acf acg ach" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*GKGOrp0xddZwHomMhiSeCA.png" /></div></div></div><div class="abz am ew gh"><div class="am gh us cs aci acj ack acl"><div class="acm acn aco acp acq fd am j"><div class="fd am j acr"><div class="acs e ej"><div><div class="e"><div tabindex="-1" class="cq"><a href="https://netflixtechblog.com" rel="noopener follow"><div class="dm"><img alt="Netflix TechBlog" class="ei ip e acu act" src="https://miro.medium.com/v2/resize:fill:40:40/1*ty4NvNrGg4ReETxqU2N3Og.png" width="20" height="20" /></div></a></div></div></div></div><div class="wq e ej dd"><p class="z b y u w">In</p></div><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai ls am j" href="https://netflixtechblog.com" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">Netflix TechBlog</p></a></div></div></div></div></div><div class="acv e ej dd"><p class="z b y u w">by</p></div><div class="fd am j acr"><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai ls am j" href="https://netflixtechblog.medium.com" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">Netflix Technology Blog</p></a></div></div></div></div></div><div class="ny am j ej db dd"><p class="z b y u w">·</p><p class="z b y u w">Jul 17</p></div></div><div class="acw e acx acy acz ada adb it"><div class="adc add ade adf adg adh adi adj"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" rel="noopener follow" href="https://netflixtechblog.com/in-house-llm-serving-at-netflix-a5a8e799ea2c" data-discover="true"><div title=""><h2 class="z kb qc qe adk adl adm adn ado adp fl qf qh adq adr ads adt adu adv fn fo fp adw adx ady adz aea aeb fq fr fs aec aed aee aef aeg aeh ft fu fv aei aej aek ael aem aen fw hv cv">In-House LLM Serving at Netflix</h2></div><div class="aeo e"><h3 class="z b fz u dv aep es hr aeq ht hv w">By AI Platform’s Model Runtime team and Inference team</h3></div></a></div></div><div class="wh am kx ce"><div class="am j li"><div class="ck cz da r s"><div class="dm aer xr am j"><div class="dr go aes am j li"><a class="dr hd aes am j li" tabindex="-1" rel="noopener follow" href="https://netflixtechblog.com/in-house-llm-serving-at-netflix-a5a8e799ea2c" data-discover="true"><div><div class="am"><div tabindex="-1" class="cq"><div class="am j db">A response icon10</div></div></div></div></a></div></div></div><div class="am j aeu aev"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fgenrec-towards-llm-native-recommendation-at-netflix-f20be6f643e3" rel="noopener follow"><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></a><div class="ck cz da r s"><div><div class="bt"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2Fa5a8e799ea2c&amp;operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fin-house-llm-serving-at-netflix-a5a8e799ea2c" rel="noopener follow"></a></div></div></div></div></div></div><div class="da r s"></div></div></div></div></div></div></div></div></article><article aria-label="GenPage: Towards End-to-End Generative Homepage Construction at Netflix" class="al" data-testid="post-preview"><div class="al vi e"><div class="cs"><div class="al e"><div class="dm al abl abm abn abo abp abq abr abs abt abu abv abw abx"><div class="aby"><div aria-label="GenPage: Towards End-to-End Generative Homepage Construction at Netflix"><div class="aca acb acc acd ek"><img alt="GenPage: Towards End-to-End Generative Homepage Construction at Netflix" class="cs ace acf acg ach" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*NurrizMgC7_QbsGuuW42Dg.gif" /></div></div></div><div class="abz am ew gh"><div class="am gh us cs aci acj ack acl"><div class="acm acn aco acp acq fd am j"><div class="fd am j acr"><div class="acs e ej"><div><div class="e"><div tabindex="-1" class="cq"><a href="https://netflixtechblog.com" rel="noopener follow"><div class="dm"><img alt="Netflix TechBlog" class="ei ip e acu act" src="https://miro.medium.com/v2/resize:fill:40:40/1*ty4NvNrGg4ReETxqU2N3Og.png" width="20" height="20" /></div></a></div></div></div></div><div class="wq e ej dd"><p class="z b y u w">In</p></div><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai ls am j" href="https://netflixtechblog.com" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">Netflix TechBlog</p></a></div></div></div></div></div><div class="acv e ej dd"><p class="z b y u w">by</p></div><div class="fd am j acr"><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai ls am j" href="https://netflixtechblog.medium.com" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">Netflix Technology Blog</p></a></div></div></div></div></div><div class="ny am j ej db dd"><p class="z b y u w">·</p><p class="z b y u w">Jun 29</p></div></div><div class="acw e acx acy acz ada adb it"><div class="adc add ade adf adg adh adi adj"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" rel="noopener follow" href="https://netflixtechblog.com/genpage-towards-end-to-end-generative-homepage-construction-at-netflix-77146fba8a08" data-discover="true"><div title="GenPage: Towards End-to-End Generative Homepage Construction at Netflix"><h2 class="z kb qc qe adk adl adm adn ado adp fl qf qh adq adr ads adt adu adv fn fo fp adw adx ady adz aea aeb fq fr fs aec aed aee aef aeg aeh ft fu fv aei aej aek ael aem aen fw hv cv">GenPage: Towards End-to-End Generative Homepage Construction at Netflix</h2></div><div class="aeo e"><h3 class="z b fz u dv aep es hr aeq ht hv w">Authors: Lequn Wang, Jiangwei Pan, and Linas Baltrunas</h3></div></a></div></div><div class="wh am kx ce"><div class="am j li"><div class="ck cz da r s"><div class="dm aer xr am j"><div class="dr go aes am j li"><a class="dr hd aes am j li" tabindex="-1" rel="noopener follow" href="https://netflixtechblog.com/genpage-towards-end-to-end-generative-homepage-construction-at-netflix-77146fba8a08" data-discover="true"><div><div class="am"><div tabindex="-1" class="cq"><div class="am j db">A response icon12</div></div></div></div></a></div></div></div><div class="am j aeu aev"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fgenrec-towards-llm-native-recommendation-at-netflix-f20be6f643e3" rel="noopener follow"><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></a><div class="ck cz da r s"><div><div class="bt"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2F77146fba8a08&amp;operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fgenpage-towards-end-to-end-generative-homepage-construction-at-netflix-77146fba8a08" rel="noopener follow"></a></div></div></div></div></div></div><div class="da r s"></div></div></div></div></div></div></div></div></article><article aria-label="Building Service Topology at Scale: Architecture, Challenges, and Lessons Learned" class="al" data-testid="post-preview"><div class="al vi e"><div class="cs"><div class="al e"><div class="dm al abl abm abn abo abp abq abr abs abt abu abv abw abx"><div class="aby"><div aria-label="Diagram showing two separate ingestion pipelines — a flow log pipeline and an IPC pipeline, each fed by data enrichment — writing to their own graph store, with a shared API serving UI and backend clients"><div class="aca acb acc acd ek"><img alt="Diagram showing two separate ingestion pipelines — a flow log pipeline and an IPC pipeline, each fed by data enrichment — writing to their own graph store, with a shared API serving UI and backend clients" class="cs ace acf acg ach" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*5_09VEnyPtFR0tbO1zzzMQ.png" /></div></div></div><div class="abz am ew gh"><div class="am gh us cs aci acj ack acl"><div class="acm acn aco acp acq fd am j"><div class="fd am j acr"><div class="acs e ej"><div><div class="e"><div tabindex="-1" class="cq"><a href="https://netflixtechblog.com" rel="noopener follow"><div class="dm"><img alt="Netflix TechBlog" class="ei ip e acu act" src="https://miro.medium.com/v2/resize:fill:40:40/1*ty4NvNrGg4ReETxqU2N3Og.png" width="20" height="20" /></div></a></div></div></div></div><div class="wq e ej dd"><p class="z b y u w">In</p></div><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai ls am j" href="https://netflixtechblog.com" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">Netflix TechBlog</p></a></div></div></div></div></div><div class="acv e ej dd"><p class="z b y u w">by</p></div><div class="fd am j acr"><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai ls am j" href="https://netflixtechblog.medium.com" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">Netflix Technology Blog</p></a></div></div></div></div></div><div class="ny am j ej db dd"><p class="z b y u w">·</p><p class="z b y u w">Jul 13</p></div></div><div class="acw e acx acy acz ada adb it"><div class="adc add ade adf adg adh adi adj"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" rel="noopener follow" href="https://netflixtechblog.com/building-service-topology-at-scale-architecture-challenges-and-lessons-learned-f4b792f3f0d8" data-discover="true"><div title="Building Service Topology at Scale: Architecture, Challenges, and Lessons Learned"><h2 class="z kb qc qe adk adl adm adn ado adp fl qf qh adq adr ads adt adu adv fn fo fp adw adx ady adz aea aeb fq fr fs aec aed aee aef aeg aeh ft fu fv aei aej aek ael aem aen fw hv cv">Building Service Topology at Scale: Architecture, Challenges, and Lessons Learned</h2></div><div class="aeo e"><h3 class="z b fz u dv aep es hr aeq ht hv w">By Parth Jain, Rakesh Sukumar, Yingwu Zhao, Renzo Sanchez-Silva &amp; Nathan Fisher. A deep dive into the engineering challenges of building a…</h3></div></a></div></div><div class="wh am kx ce"><div class="am j li"><div class="ck cz da r s"><div class="dm aer xr am j"><div class="dr go aes am j li"><a class="dr hd aes am j li" tabindex="-1" rel="noopener follow" href="https://netflixtechblog.com/building-service-topology-at-scale-architecture-challenges-and-lessons-learned-f4b792f3f0d8" data-discover="true"><div><div class="am"><div tabindex="-1" class="cq"><div class="am j db">A response icon7</div></div></div></div></a></div></div></div><div class="am j aeu aev"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fgenrec-towards-llm-native-recommendation-at-netflix-f20be6f643e3" rel="noopener follow"><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></a><div class="ck cz da r s"><div><div class="bt"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2Ff4b792f3f0d8&amp;operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fbuilding-service-topology-at-scale-architecture-challenges-and-lessons-learned-f4b792f3f0d8" rel="noopener follow"></a></div></div></div></div></div></div><div class="da r s"></div></div></div></div></div></div></div></div></article><article aria-label="How Netflix Simplified Batch Compute with Kueue" class="al" data-testid="post-preview"><div class="al vi e"><div class="cs"><div class="al e"><div class="dm al abl abm abn abo abp abq abr abs abt abu abv abw abx"><div class="aby"><div aria-label="How Netflix Simplified Batch Compute with Kueue"><div class="aca acb acc acd ek"><img alt="How Netflix Simplified Batch Compute with Kueue" class="cs ace acf afa ach" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*GL_sNsqautjh4lnu5XmRww.png" /></div></div></div><div class="abz am ew gh"><div class="am gh us cs aci acj ack acl"><div class="acm acn aco acp acq fd am j"><div class="fd am j acr"><div class="acs e ej"><div><div class="e"><div tabindex="-1" class="cq"><a href="https://netflixtechblog.com" rel="noopener follow"><div class="dm"><img alt="Netflix TechBlog" class="ei ip e acu act" src="https://miro.medium.com/v2/resize:fill:40:40/1*ty4NvNrGg4ReETxqU2N3Og.png" width="20" height="20" /></div></a></div></div></div></div><div class="wq e ej dd"><p class="z b y u w">In</p></div><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai ls am j" href="https://netflixtechblog.com" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">Netflix TechBlog</p></a></div></div></div></div></div><div class="acv e ej dd"><p class="z b y u w">by</p></div><div class="fd am j acr"><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai ls am j" href="https://netflixtechblog.medium.com" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">Netflix Technology Blog</p></a></div></div></div></div></div><div class="ny am j ej db dd"><p class="z b y u w">·</p><p class="z b y u w">Jun 22</p></div></div><div class="acw e acx acy acz ada adb it"><div class="adc add ade adf adg adh adi adj"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" rel="noopener follow" href="https://netflixtechblog.com/how-netflix-simplified-batch-compute-with-kueue-87860682629c" data-discover="true"><div title=""><h2 class="z kb qc qe adk adl adm adn ado adp fl qf qh adq adr ads adt adu adv fn fo fp adw adx ady adz aea aeb fq fr fs aec aed aee aef aeg aeh ft fu fv aei aej aek ael aem aen fw hv cv">How Netflix Simplified Batch Compute with Kueue</h2></div><div class="aeo e"><h3 class="z b fz u dv aep es hr aeq ht hv w">By Alvin Bao, Alex Petrov, Jennifer Lai, Aidan Sherr, and Samartha Chandrashekar</h3></div></a></div></div><div class="wh am kx ce"><div class="am j li"><div class="ck cz da r s"><div class="dm aer xr am j"><div class="dr go aes am j li"><a class="dr hd aes am j li" tabindex="-1" rel="noopener follow" href="https://netflixtechblog.com/how-netflix-simplified-batch-compute-with-kueue-87860682629c" data-discover="true"><div><div class="am"><div tabindex="-1" class="cq"><div class="am j db">A response icon1</div></div></div></div></a></div></div></div><div class="am j aeu aev"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fgenrec-towards-llm-native-recommendation-at-netflix-f20be6f643e3" rel="noopener follow"><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></a><div class="ck cz da r s"><div><div class="bt"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2F87860682629c&amp;operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fhow-netflix-simplified-batch-compute-with-kueue-87860682629c" rel="noopener follow"></a></div></div></div></div></div></div></div></div></div></div></div></div></div></article><article aria-label="In-House LLM Serving at Netflix" class="al" data-testid="post-preview"><div class="al vi e"><div class="cs"><div class="al e"><div class="dm al abl abm abn abo abp abq abr abs abt abu abv abw abx"><div class="aby"><div aria-label="In-House LLM Serving at Netflix"><div class="aca acb acc acd ek"><img alt="In-House LLM Serving at Netflix" class="cs ace acf acg ach" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*GKGOrp0xddZwHomMhiSeCA.png" /></div></div></div><div class="abz am ew gh"><div class="am gh us cs aci acj ack acl"><div class="acm acn aco acp acq fd am j"><div class="fd am j acr"><div class="acs e ej"><div><div class="e"><div tabindex="-1" class="cq"><a href="https://netflixtechblog.com" rel="noopener follow"><div class="dm"><img alt="Netflix TechBlog" class="ei ip e acu act" src="https://miro.medium.com/v2/resize:fill:40:40/1*ty4NvNrGg4ReETxqU2N3Og.png" width="20" height="20" /></div></a></div></div></div></div><div class="wq e ej dd"><p class="z b y u w">In</p></div><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai ls am j" href="https://netflixtechblog.com" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">Netflix TechBlog</p></a></div></div></div></div></div><div class="acv e ej dd"><p class="z b y u w">by</p></div><div class="fd am j acr"><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai ls am j" href="https://netflixtechblog.medium.com" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">Netflix Technology Blog</p></a></div></div></div></div></div><div class="ny am j ej db dd"><p class="z b y u w">·</p><p class="z b y u w">Jul 17</p></div></div><div class="acw e acx acy acz ada adb it"><div class="adc add ade adf adg adh adi adj"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" rel="noopener follow" href="https://netflixtechblog.com/in-house-llm-serving-at-netflix-a5a8e799ea2c" data-discover="true"><div title=""><h2 class="z kb qc qe adk adl adm adn ado adp fl qf qh adq adr ads adt adu adv fn fo fp adw adx ady adz aea aeb fq fr fs aec aed aee aef aeg aeh ft fu fv aei aej aek ael aem aen fw hv cv">In-House LLM Serving at Netflix</h2></div><div class="aeo e"><h3 class="z b fz u dv aep es hr aeq ht hv w">By AI Platform’s Model Runtime team and Inference team</h3></div></a></div></div><div class="wh am kx ce"><div class="am j li"><div class="ck cz da r s"><div class="dm aer xr am j"><div class="dr go aes am j li"><a class="dr hd aes am j li" tabindex="-1" rel="noopener follow" href="https://netflixtechblog.com/in-house-llm-serving-at-netflix-a5a8e799ea2c" data-discover="true"><div><div class="am"><div tabindex="-1" class="cq"><div class="am j db">A response icon10</div></div></div></div></a></div></div></div><div class="am j aeu aev"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fgenrec-towards-llm-native-recommendation-at-netflix-f20be6f643e3" rel="noopener follow"><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></a><div class="ck cz da r s"><div><div class="bt"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2Fa5a8e799ea2c&amp;operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fin-house-llm-serving-at-netflix-a5a8e799ea2c" rel="noopener follow"></a></div></div></div></div></div></div><div class="da r s"></div></div></div></div></div></div></div></div></article><article aria-label="We Got Women Into Tech, Then Failed To Keep Them" class="al" data-testid="post-preview"><div class="al vi e"><div class="cs"><div class="al e"><div class="dm al abl abm abn abo abp abq abr abs abt abu abv abw abx"><div class="aby"><div aria-label="We Got Women Into Tech, Then Failed To Keep Them"><div class="aca acb acc acd ek"><img alt="We Got Women Into Tech, Then Failed To Keep Them" class="cs ace acf acg ach" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/0*XGqJp9guR2FlNVVO" /></div></div></div><div class="abz am ew gh"><div class="am gh us cs aci acj ack acl"><div class="acm acn aco acp acq fd am j"><div class="fd am j acr"><div class="acs e ej"><div><div class="e"><div tabindex="-1" class="cq"><a href="https://levelup.gitconnected.com" rel="noopener follow"><div class="dm"><img alt="Level Up Coding" class="ei ip e acu act" src="https://miro.medium.com/v2/resize:fill:40:40/1*5D9oYBd58pyjMkV_5-zXXQ.jpeg" width="20" height="20" /></div></a></div></div></div></div><div class="wq e ej dd"><p class="z b y u w">In</p></div><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai ls am j" href="https://levelup.gitconnected.com" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">Level Up Coding</p></a></div></div></div></div></div><div class="acv e ej dd"><p class="z b y u w">by</p></div><div class="fd am j acr"><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai ls am j" href="https://attilavago.medium.com" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">Attila Vágó</p><div class="agi agj e"><div class="am agk"><div class="am"></div></div></div></a></div></div></div></div></div><div class="ny am j ej db dd"><p class="z b y u w">·</p><p class="z b y u w">5d ago</p></div></div><div class="acw e acx acy acz ada adb it"><div class="adc add ade adf adg adh adi adj"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://levelup.gitconnected.com/we-got-women-into-tech-then-failed-to-keep-them-4035564979ea" rel="noopener follow"><div title=""><h2 class="z kb qc qe adk adl adm adn ado adp fl qf qh adq adr ads adt adu adv fn fo fp adw adx ady adz aea aeb fq fr fs aec aed aee aef aeg aeh ft fu fv aei aej aek ael aem aen fw hv cv">We Got Women Into Tech, Then Failed To Keep Them</h2></div><div class="aeo e"><h3 class="z b fz u dv aep es hr aeq ht hv w">Come join us, they said… It’s gonna be great, they said… Until it wasn’t. An engineer mourning the loss of great talent and diversity in…</h3></div></a></div></div><div class="wh am kx ce"><div class="am j li"><div class="vi am"><button class="e cf af ac" aria-label="Member-only story"><div class=""><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></div></button></div><div class="ck cz da r s"><div class="dm aer xr am j"><div class="dr go aes am j li"><a class="dr hd aes am j li" tabindex="-1" href="https://levelup.gitconnected.com/we-got-women-into-tech-then-failed-to-keep-them-4035564979ea" rel="noopener follow"><div><div class="am"><div tabindex="-1" class="cq"><div class="am j db">A response icon24</div></div></div></div></a></div></div></div><div class="am j aeu aev"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fgenrec-towards-llm-native-recommendation-at-netflix-f20be6f643e3" rel="noopener follow"><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></a><div class="ck cz da r s"><div><div class="bt"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2F4035564979ea&amp;operation=register&amp;redirect=https%3A%2F%2Flevelup.gitconnected.com%2Fwe-got-women-into-tech-then-failed-to-keep-them-4035564979ea" rel="noopener follow"></a></div></div></div></div></div></div></div></div></div></div></div></div></div></article><article aria-label="Beyond RAG: How Google’s Open Knowledge Format (OKF) is Replacing the Vector Database" class="al" data-testid="post-preview"><div class="al vi e"><div class="cs"><div class="al e"><div class="dm al abl abm abn abo abp abq abr abs abt abu abv abw abx"><div class="aby"><div aria-label="Beyond RAG: How Google’s Open Knowledge Format (OKF) is Replacing the Vector Database"><div class="aca acb acc acd ek"><img alt="Beyond RAG: How Google’s Open Knowledge Format (OKF) is Replacing the Vector Database" class="cs ace acf acg ach" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*56aqjvj8yWXomR1Jp1Ekag.png" /></div></div></div><div class="abz am ew gh"><div class="am gh us cs aci acj ack acl"><div class="acm acn aco acp acq fd am j"><div class="fd am j acr"><div class="acs e ej"><div><div class="e"><div tabindex="-1" class="cq"><a href="https://medium.com/the-code-frontier" rel="noopener follow"><div class="dm"><img alt="The Code Frontier" class="ei ip e acu act" src="https://miro.medium.com/v2/resize:fill:40:40/1*ufoOMT9lBLP9K5OvnknIbw.png" width="20" height="20" /></div></a></div></div></div></div><div class="wq e ej dd"><p class="z b y u w">In</p></div><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai ls am j" href="https://medium.com/the-code-frontier" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">The Code Frontier</p></a></div></div></div></div></div><div class="acv e ej dd"><p class="z b y u w">by</p></div><div class="fd am j acr"><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai ls am j" href="https://secret-dev.medium.com" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">Secret Dev</p></a></div></div></div></div></div><div class="ny am j ej db dd"><p class="z b y u w">·</p><p class="z b y u w">Jul 3</p></div></div><div class="acw e acx acy acz ada adb it"><div class="adc add ade adf adg adh adi adj"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/the-code-frontier/beyond-rag-how-googles-open-knowledge-format-okf-is-replacing-the-vector-database-2ffb5bc2f8eb" rel="noopener follow"><div title="Beyond RAG: How Google’s Open Knowledge Format (OKF) is Replacing the Vector Database"><h2 class="z kb qc qe adk adl adm adn ado adp fl qf qh adq adr ads adt adu adv fn fo fp adw adx ady adz aea aeb fq fr fs aec aed aee aef aeg aeh ft fu fv aei aej aek ael aem aen fw hv cv">Beyond RAG: How Google’s Open Knowledge Format (OKF) is Replacing the Vector Database</h2></div><div class="aeo e"><h3 class="z b fz u dv aep es hr aeq ht hv w">For the last three years, the default engineering response to any enterprise AI context problem was automated: “Just build a RAG pipeline.”</h3></div></a></div></div><div class="wh am kx ce"><div class="am j li"><div class="vi am"><button class="e cf af ac" aria-label="Member-only story"><div class=""><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></div></button></div><div class="ck cz da r s"><div class="dm aer xr am j"><div class="dr go aes am j li"><a class="dr hd aes am j li" tabindex="-1" href="https://medium.com/the-code-frontier/beyond-rag-how-googles-open-knowledge-format-okf-is-replacing-the-vector-database-2ffb5bc2f8eb" rel="noopener follow"><div><div class="am"><div tabindex="-1" class="cq"><div class="am j db">A response icon35</div></div></div></div></a></div></div></div><div class="am j aeu aev"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fgenrec-towards-llm-native-recommendation-at-netflix-f20be6f643e3" rel="noopener follow"><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></a><div class="ck cz da r s"><div><div class="bt"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2F2ffb5bc2f8eb&amp;operation=register&amp;redirect=https%3A%2F%2Fmedium.com%2Fthe-code-frontier%2Fbeyond-rag-how-googles-open-knowledge-format-okf-is-replacing-the-vector-database-2ffb5bc2f8eb" rel="noopener follow"></a></div></div></div></div></div></div><div class="da r s"></div></div></div></div></div></div></div></div></article><article aria-label="You only have weeks left to vibe code" class="al" data-testid="post-preview"><div class="al vi e"><div class="cs"><div class="al e"><div class="dm al abl abm abn abo abp abq abr abs abt abu abv abw abx"><div class="aby"><div aria-label="Vibe coding is really over this time, claude terminal window telling you to stop"><div class="aca acb acc acd ek"><img alt="Vibe coding is really over this time, claude terminal window telling you to stop" class="cs ace acf acg ach" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*zKR3rUwFqBImPUfH9OdFAw.png" /></div></div></div><div class="abz am ew gh"><div class="am gh us cs aci acj ack acl"><div class="acm acn aco acp acq fd am j"><div class="fd am j acr"><div class="acs e ej"><div><div class="e"><div tabindex="-1" class="cq"><a tabindex="-1" href="https://michalmalewicz.medium.com" rel="noopener follow"><div class="e dm"><img alt="Michal Malewicz" class="e bs ee act acu ei" src="https://miro.medium.com/v2/resize:fill:40:40/1*149zXrb2FXvS_mctL4NKSg.png" width="20" height="20" /></div></a></div></div></div></div><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai ls am j" href="https://michalmalewicz.medium.com" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">Michal Malewicz</p><div class="agi agj e"><div class="am agk"><div class="am"></div></div></div></a></div></div></div></div></div><div class="ny am j ej db dd"><p class="z b y u w">·</p><p class="z b y u w">Jun 29</p></div></div><div class="acw e acx acy acz ada adb it"><div class="adc add ade adf adg adh adi adj"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://michalmalewicz.medium.com/you-only-have-weeks-left-to-vibe-code-09a89c3d3f9b" rel="noopener follow"><div title=""><h2 class="z kb qc qe adk adl adm adn ado adp fl qf qh adq adr ads adt adu adv fn fo fp adw adx ady adz aea aeb fq fr fs aec aed aee aef aeg aeh ft fu fv aei aej aek ael aem aen fw hv cv">You only have weeks left to vibe code</h2></div><div class="aeo e"><h3 class="z b fz u dv aep es hr aeq ht hv w">Then it’s over. You better hurry up!</h3></div></a></div></div><div class="wh am kx ce"><div class="am j li"><div class="vi am"><button class="e cf af ac" aria-label="Member-only story"><div class=""><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></div></button></div><div class="ck cz da r s"><div class="dm aer xr am j"><div class="dr go aes am j li"><a class="dr hd aes am j li" tabindex="-1" href="https://michalmalewicz.medium.com/you-only-have-weeks-left-to-vibe-code-09a89c3d3f9b" rel="noopener follow"><div><div class="am"><div tabindex="-1" class="cq"><div class="am j db">A response icon219</div></div></div></div></a></div></div></div><div class="am j aeu aev"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fgenrec-towards-llm-native-recommendation-at-netflix-f20be6f643e3" rel="noopener follow"><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></a><div class="ck cz da r s"><div><div class="bt"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2F09a89c3d3f9b&amp;operation=register&amp;redirect=https%3A%2F%2Fmichalmalewicz.medium.com%2Fyou-only-have-weeks-left-to-vibe-code-09a89c3d3f9b" rel="noopener follow"></a></div></div></div></div></div></div><div class="da r s"></div></div></div></div></div></div></div></div></article><article aria-label="Stop Wasting LLM Tokens: Building a Self-Updating Codebase Knowledge Graph with OKF" class="al" data-testid="post-preview"><div class="al vi e"><div class="cs"><div class="al e"><div class="dm al abl abm abn abo abp abq abr abs abt abu abv abw abx"><div class="aby"><div aria-label="Stop Wasting LLM Tokens: Building a Self-Updating Codebase Knowledge Graph with OKF"><div class="aca acb acc acd ek"><img alt="Stop Wasting LLM Tokens: Building a Self-Updating Codebase Knowledge Graph with OKF" class="cs ace acf acg ach" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*6DDfG1myntoUimy5wZN1Lg.png" /></div></div></div><div class="abz am ew gh"><div class="am gh us cs aci acj ack acl"><div class="acm acn aco acp acq fd am j"><div class="fd am j acr"><div class="acs e ej"><div><div class="e"><div tabindex="-1" class="cq"><a href="https://medium.com/data-science-collective" rel="noopener follow"><div class="dm"><img alt="Data Science Collective" class="ei ip e acu act" src="https://miro.medium.com/v2/resize:fill:40:40/1*0nV0Q-FBHj94Kggq00pG2Q.jpeg" width="20" height="20" /></div></a></div></div></div></div><div class="wq e ej dd"><p class="z b y u w">In</p></div><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai ls am j" href="https://medium.com/data-science-collective" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">Data Science Collective</p></a></div></div></div></div></div><div class="acv e ej dd"><p class="z b y u w">by</p></div><div class="fd am j acr"><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai ls am j" href="https://medium.com/@UdaykiranEstari" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">Udaykiran Estari</p></a></div></div></div></div></div><div class="ny am j ej db dd"><p class="z b y u w">·</p><p class="z b y u w">Jul 3</p></div></div><div class="acw e acx acy acz ada adb it"><div class="adc add ade adf adg adh adi adj"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/data-science-collective/stop-wasting-llm-tokens-building-a-self-updating-codebase-knowledge-graph-with-okf-20284060c1b1" rel="noopener follow"><div title="Stop Wasting LLM Tokens: Building a Self-Updating Codebase Knowledge Graph with OKF"><h2 class="z kb qc qe adk adl adm adn ado adp fl qf qh adq adr ads adt adu adv fn fo fp adw adx ady adz aea aeb fq fr fs aec aed aee aef aeg aeh ft fu fv aei aej aek ael aem aen fw hv cv">Stop Wasting LLM Tokens: Building a Self-Updating Codebase Knowledge Graph with OKF</h2></div><div class="aeo e"><h3 class="z b fz u dv aep es hr aeq ht hv w">Google's new Open Knowledge Format (OKF) standardizes context. Here's how to build a pipeline to keep your codebase memory self-updating.</h3></div></a></div></div><div class="wh am kx ce"><div class="am j li"><div class="vi am"><button class="e cf af ac" aria-label="Member-only story"><div class=""><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></div></button></div><div class="ck cz da r s"><div class="dm aer xr am j"><div class="dr go aes am j li"><a class="dr hd aes am j li" tabindex="-1" href="https://medium.com/data-science-collective/stop-wasting-llm-tokens-building-a-self-updating-codebase-knowledge-graph-with-okf-20284060c1b1" rel="noopener follow"><div><div class="am"><div tabindex="-1" class="cq"><div class="am j db">A response icon7</div></div></div></div></a></div></div></div><div class="am j aeu aev"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fgenrec-towards-llm-native-recommendation-at-netflix-f20be6f643e3" rel="noopener follow"><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></a><div class="ck cz da r s"><div><div class="bt"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2F20284060c1b1&amp;operation=register&amp;redirect=https%3A%2F%2Fmedium.com%2F%40UdaykiranEstari%2Fstop-wasting-llm-tokens-building-a-self-updating-codebase-knowledge-graph-with-okf-20284060c1b1" rel="noopener follow"></a></div></div></div></div></div></div><div class="da r s"></div></div></div></div></div></div></div></div></article><article aria-label="MCP is Dead" class="al" data-testid="post-preview"><div class="al vi e"><div class="cs"><div class="al e"><div class="dm al abl abm abn abo abp abq abr abs abt abu abv abw abx"><div class="aby"><div aria-label="MCP is Dead"><div class="aca acb acc acd ek"><img alt="MCP is Dead" class="cs ace acf acg ach" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*Oj5PiyfEi8DadSC8Jy374w.png" /></div></div></div><div class="abz am ew gh"><div class="am gh us cs aci acj ack acl"><div class="acm acn aco acp acq fd am j"><div class="fd am j acr"><div class="acs e ej"><div><div class="e"><div tabindex="-1" class="cq"><a href="https://uxplanet.org" rel="noopener follow"><div class="dm"><img alt="UX Planet" class="ei ip e acu act" src="https://miro.medium.com/v2/resize:fill:40:40/1*A0FnBy5FBoVQC02SZXLXPg.png" width="20" height="20" /></div></a></div></div></div></div><div class="wq e ej dd"><p class="z b y u w">In</p></div><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai ls am j" href="https://uxplanet.org" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">UX Planet</p></a></div></div></div></div></div><div class="acv e ej dd"><p class="z b y u w">by</p></div><div class="fd am j acr"><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai ls am j" href="https://medium.com/@101" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">Nick Babich</p><div class="agi agj e"><div class="am agk"><div class="am"></div></div></div></a></div></div></div></div></div><div class="ny am j ej db dd"><p class="z b y u w">·</p><p class="z b y u w">Apr 6</p></div></div><div class="acw e acx acy acz ada adb it"><div class="adc add ade adf adg adh adi adj"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://uxplanet.org/mcp-is-dead-cf16b667ba6d" rel="noopener follow"><div title=""><h2 class="z kb qc qe adk adl adm adn ado adp fl qf qh adq adr ads adt adu adv fn fo fp adw adx ady adz aea aeb fq fr fs aec aed aee aef aeg aeh ft fu fv aei aej aek ael aem aen fw hv cv">MCP is Dead</h2></div><div class="aeo e"><h3 class="z b fz u dv aep es hr aeq ht hv w">Why you should avoid using MCP in Claude Code and what to use instead</h3></div></a></div></div><div class="wh am kx ce"><div class="am j li"><div class="vi am"><button class="e cf af ac" aria-label="Member-only story"><div class=""><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></div></button></div><div class="ck cz da r s"><div class="dm aer xr am j"><div class="dr go aes am j li"><a class="dr hd aes am j li" tabindex="-1" href="https://uxplanet.org/mcp-is-dead-cf16b667ba6d" rel="noopener follow"><div><div class="am"><div tabindex="-1" class="cq"><div class="am j db">A response icon289</div></div></div></div></a></div></div></div><div class="am j aeu aev"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fgenrec-towards-llm-native-recommendation-at-netflix-f20be6f643e3" rel="noopener follow"><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></a><div class="ck cz da r s"><div><div class="bt"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2Fcf16b667ba6d&amp;operation=register&amp;redirect=https%3A%2F%2Fuxplanet.org%2Fmcp-is-dead-cf16b667ba6d" rel="noopener follow"></a></div></div></div></div></div></div></div></div></div></div></div></div></div></article>]]></description>
      <link>https://netflixtechblog.com/genrec-towards-llm-native-recommendation-at-netflix-f20be6f643e3</link>
      <guid>https://netflixtechblog.com/genrec-towards-llm-native-recommendation-at-netflix-f20be6f643e3</guid>
      <pubDate>Thu, 30 Jul 2026 22:10:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[In-House LLM Serving at Netflix]]></title>
      <description><![CDATA[<article><div class="e"><div class="e"><section><div><div class="ik jd je jf jg"><div class="am ew"><div class="fd cs iq ir is it"><div><div></div><p id="5e4f" class="pw-post-body-paragraph ot ou jj ov b ow ox oy oz pa pb pc pd fo pe pf pg fr ph pi pj fu pk pl pm pn ik cv"><em class="po">By AI Platform’s Model Runtime team and Inference team</em></p><h2 id="e501" class="pp pq jj z pr ps pt pu fl pv pw px fn py pz qa qb qc qd qe qf qg qh qi qj qk cv">Introduction</h2><p id="5e53" class="pw-post-body-paragraph ot ou jj ov b ow ql oy oz pa qm pc pd fo qn pf pg fr qo pi pj fu qp pl pm pn ik cv">Most organizations consume LLMs through hosted APIs. Netflix went further — we run the full stack ourselves, from model deployment through inference, inside our existing production environment rather than a separate ML silo. Some of those decisions weren’t obvious, and a few revealed their trade-offs only under production load.</p><p id="f1cd" class="pw-post-body-paragraph ot ou jj ov b ow ox oy oz pa pb pc pd fo pe pf pg fr ph pi pj fu pk pl pm pn ik cv">This post focuses on the choices where alternatives were seriously considered: engine selection, model packaging, API surface design, deployment strategy, and output constraints enforcement. The goal is to share not just what was built, but why — and what production revealed that the design phase didn’t anticipate.</p><h2 id="b9f1" class="pp pq jj z pr ps pt pu fl pv pw px fn py pz qa qb qc qd qe qf qg qh qi qj qk cv">Architecture Overview</h2><p id="286b" class="pw-post-body-paragraph ot ou jj ov b ow ql oy oz pa qm pc pd fo qn pf pg fr qo pi pj fu qp pl pm pn ik cv">Member-scale ML at Netflix is fronted by a unified JVM-based serving system that handles the end-to-end flow for downstream consumers: routing and A/B test logic, candidate generation, feature fetching, inference, post-processing, and logging at each stage. Both real-time and cached batch paths are supported. Figure 1 shows the two ways callers reach inference today: the gRPC path through this serving system and a direct HTTP path used by newer LLM-driven applications.</p><p id="9db1" class="pw-post-body-paragraph ot ou jj ov b ow ox oy oz pa pb pc pd fo pe pf pg fr ph pi pj fu pk pl pm pn ik cv">Where inference runs depends on the model. Small CPU models run in-process, avoiding remote-call overhead. Larger models need GPUs — the serving system handles pre- and post-processing locally but delegates inference to a remote service, <strong class="ov jk">Model Scoring Service (MSS)</strong>. MSS is the shared inference backend, supporting XGBoost, TensorFlow, PyTorch, and LLMs behind a unified interface, with NVIDIA Triton Inference Server underneath managing model loading, batching, and GPU scheduling.</p><p id="9176" class="pw-post-body-paragraph ot ou jj ov b ow ox oy oz pa pb pc pd fo pe pf pg fr ph pi pj fu pk pl pm pn ik cv">On top of Triton sits a Java control plane that handles deployment, versioning, health checking, autoscaling, and multi-region rollout. Model authors package their artifacts and configure the deployment; the control plane provisions GPU instances, configures Triton, and orchestrates zero-downtime upgrades.</p><figure class="qt qu qv qw qx qy qq qr paragraph-image"><div role="button" tabindex="0" class="qz ra dm rb cs rc">Press enter or click to view image in full size<div class="qq qr qs"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*GKGOrp0xddZwHomMhiSeCA.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*GKGOrp0xddZwHomMhiSeCA.png" /><img alt="" class="cs er rd re" width="700" height="392" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="rf bv rg qq qr rh ri z b cr u w">Figure 1. Serving Architecture Overview</figcaption></figure><h2 id="dd34" class="pp pq jj z pr ps pt pu fl pv pw px fn py pz qa qb qc qd qe qf qg qh qi qj qk cv">Design Decisions and Implementation</h2><p id="f947" class="pw-post-body-paragraph ot ou jj ov b ow ql oy oz pa qm pc pd fo qn pf pg fr qo pi pj fu qp pl pm pn ik cv">Four decisions shape this platform — engine, packaging, API surface, and rollout — presented in dependency order, since each one constrains the next.</p><h3 id="398b" class="rj pq jj z pr fk rk ao fl fm rl aq fn fo rm fp fq fr rn fs ft fu ro fv fw rp cv">vLLM as the Paved-Path Engine</h3><p id="954f" class="pw-post-body-paragraph ot ou jj ov b ow ql oy oz pa qm pc pd fo qn pf pg fr qo pi pj fu qp pl pm pn ik cv">The platform was originally built on TensorRT-LLM, a performant inference engine at the time and already integrated with Triton — the compute backend in use within MSS.</p><p id="0e48" class="pw-post-body-paragraph ot ou jj ov b ow ox oy oz pa pb pc pd fo pe pf pg fr ph pi pj fu pk pl pm pn ik cv">By summer 2025, two things had shifted: open-source engines had largely closed the performance gap with specialized stacks, and our workload mix had broadened to include embedding generation, prefill-only inference for ranking and retrieval, autoregressive decoding, and custom models with non-trivial per-step constraint logic. We re-benchmarked against this mix and <strong class="ov jk">selected vLLM as our paved-path engine</strong> on operational fit:</p><ul class=""><li id="a052" class="ot ou jj ov b ow ox oy oz pa pb pc pd fo pe pf pg fr ph pi pj fu pk pl pm pn rq rr rs cv"><strong class="ov jk">Loads custom model architectures without a multi-step compilation pipeline</strong> — faster iteration on non-standard models.</li><li id="01a5" class="ot ou jj ov b ow rt oy oz pa ru pc pd fo rv pf pg fr rw pi pj fu rx pl pm pn rq rr rs cv"><strong class="ov jk">Extensibility hooks for custom decoding logic</strong> — necessary for the constrained-decoding work described later.</li><li id="2385" class="ot ou jj ov b ow rt oy oz pa ru pc pd fo rv pf pg fr rw pi pj fu rx pl pm pn rq rr rs cv"><strong class="ov jk">Debuggability</strong> — easier to inspect failures and intermediate state than with a compiled engine in earlier TensorRT-LLM.</li><li id="c2b9" class="ot ou jj ov b ow rt oy oz pa ru pc pd fo rv pf pg fr rw pi pj fu rx pl pm pn rq rr rs cv"><strong class="ov jk">Familiarity</strong> — many ML practitioners were already using vLLM in research, which cut the research-to-production handoff cost.</li></ul><h3 id="6419" class="rj pq jj z pr fk rk ao fl fm rl aq fn fo rm fp fq fr rn fs ft fu ro fv fw rp cv">Integrating vLLM into Triton</h3><p id="b5bc" class="pw-post-body-paragraph ot ou jj ov b ow ql oy oz pa qm pc pd fo qn pf pg fr qo pi pj fu qp pl pm pn ik cv">With vLLM picked, the next decision was how to package models for it. Triton supports two ways, and the choice has significant implications for maintainability — specifically, how tightly model artifacts are coupled to frontend upgrades.</p><ul class=""><li id="499d" class="ot ou jj ov b ow ox oy oz pa pb pc pd fo pe pf pg fr ph pi pj fu pk pl pm pn rq rr rs cv"><strong class="ov jk">Python backend.</strong> The author defines explicit input/output tensor specs at packaging time. These specs are frozen in the artifact and must match what the third-party vendor’s frontend’s request builder expects, so every frontend upgrade that touches I/O specs requires a coordinated change to packaging code; otherwise, requests fail at runtime.</li><li id="e464" class="ot ou jj ov b ow rt oy oz pa ru pc pd fo rv pf pg fr rw pi pj fu rx pl pm pn rq rr rs cv"><strong class="ov jk">vLLM backend.</strong> The artifact is just a JSON config pointing to the model weights and tokenizer. Triton’s vLLM backend reads this config and generates I/O tensor specs dynamically at deployment time — the author never defines them. Models and frontend evolve independently.</li></ul><p id="bd3c" class="pw-post-body-paragraph ot ou jj ov b ow ox oy oz pa pb pc pd fo pe pf pg fr ph pi pj fu pk pl pm pn ik cv">The vLLM backend is the architecturally correct default. Two things bit us in production:</p><ul class=""><li id="264e" class="ot ou jj ov b ow ox oy oz pa pb pc pd fo pe pf pg fr ph pi pj fu pk pl pm pn rq rr rs cv"><strong class="ov jk">Triton/vLLM version mismatch.</strong> Triton’s vLLM backend is compiled against a specific vLLM API surface. When the two drift — for example, Triton 25.09 importing vllm.engine.metrics, a module removed in vLLM 0.11.2 — the backend fails to load entirely. The platform has to pin compatible versions when baking the service image, and prevent model authors from overriding the vLLM version at packaging time.</li><li id="4740" class="ot ou jj ov b ow rt oy oz pa ru pc pd fo rv pf pg fr rw pi pj fu rx pl pm pn rq rr rs cv"><strong class="ov jk">Custom model logic.</strong> The vLLM backend expects a standard HuggingFace-compatible model and handles the full inference lifecycle. Models needing custom preprocessing, postprocessing, or non-standard execution — ensemble pipelines, custom tokenization — must use the Python backend, which gives full control over execute(). This escape hatch will likely remain necessary for a subset of models.</li></ul><h3 id="595f" class="rj pq jj z pr fk rk ao fl fm rl aq fn fo rm fp fq fr rn fs ft fu ro fv fw rp cv">Ecosystem-Compatible HTTP Frontend</h3><p id="f244" class="pw-post-body-paragraph ot ou jj ov b ow ql oy oz pa qm pc pd fo qn pf pg fr qo pi pj fu qp pl pm pn ik cv">With engine and packaging settled, the next question is how callers reach the system. A key design goal of our system was that LLM models should <strong class="ov jk">NOT</strong> be special snowflakes. Every model — XGBoost ensemble or large-scale LLMs — is scored via the same gRPC call, so we reuse the same client libraries, health checking, and deployment pipelines. Given that the OpenAI-compatible API interface has become the de facto interface for the LLM ecosystem — inference engines, orchestration frameworks, evaluation tools, and client libraries all speak it — so we <strong class="ov jk">expose the OpenAI-compatible API as an additional frontend alongside gRPC</strong>.</p><p id="c236" class="pw-post-body-paragraph ot ou jj ov b ow ox oy oz pa pb pc pd fo pe pf pg fr ph pi pj fu pk pl pm pn ik cv">The payoff shows up in the experimentation-to-production path: graduating from a hosted model to a fine-tuned self-hosted one — for quality, latency, cost, or data privacy — is nearly seamless. Same API, minimal code changes.</p><div class="ry et rz sa am j gh"><div class="sb sc sd se sf e bv"><h2 class="sg b sh si sj sk sl">Get Netflix Technology Blog’s stories in your inbox</h2></div><div class="sm e bv"><p class="z b cr u w">Join Medium for free to get updates from this writer.</p></div><div class="sq sr ss st su sv am"><div class="sw am gh"><div class="am sw br bq iw sx sy ei go sz ta tb tc td te tf"><input class="tg ca by cb th ti dx na cw hp cs" placeholder="Enter your email" type="text" value="" /></div></div><div class="ck cz tj ej"><div class="sn so sp"><button class="z b cr u tk tl tm tn to tp tq bk bl tr ts tt bp bq br bs bt bu bv">Subscribe</button></div></div><div class="cs da r s"><button class="z b cr u tk tl tm tn to tp tq bk bl tr ts tt bp cs bq br bs bt bu bv">Subscribe</button></div></div><div class="ij e"><label class="am j"><input class="tu tv tw dr hd tx ty tz ua ub uc ud" type="checkbox" checked="checked" /></label></div><div class="e"><p class="z b cr u cv">Remember me for faster sign in</p></div></div><div class="ul e"></div><p id="630d" class="pw-post-body-paragraph ot ou jj ov b ow oy oz pa pc pd fo pf pg fr pi pj fu pl pm sa pn ik cv">Behind the API, the implementation reuses <a class="bx um" href="https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/client_guide/openai_readme.html" rel="noopener ugc nofollow" target="_blank">NVIDIA’s Triton OpenAI-compatible frontend</a>. It starts an embedded Triton server, wraps it in a TritonLLMEngine that converts request schemas into Triton inference requests, and serves responses through FastAPI. KServe HTTP/gRPC frontends are enabled alongside, so the same Triton instance remains accessible to the Java control plane over gRPC. Adopting Triton’s frontend directly exposed one gap: response_format — accepted by the schema — was silently dropped before reaching vLLM, so that a caller requesting JSON output proceeded without guided decoding constraints and could receive malformed JSON with no error surfaced by the platform. We git-subtreed and patched the frontend to translate response_format into vLLM’s guided decoding parameters at request time.</p><h3 id="243a" class="rj pq jj z pr fk rk ao fl fm rl aq fn fo rm fp fq fr rn fs ft fu ro fv fw rp cv">Deployment Strategies</h3><p id="74cd" class="pw-post-body-paragraph ot ou jj ov b ow ql oy oz pa qm pc pd fo qn pf pg fr qo pi pj fu qp pl pm pn ik cv">With API surface and engine in place, the question that remains is how new versions roll out without dropping requests. GPU deployments take longer to bring up than CPU services, and the I/O schema may change between model versions — adding a coordination problem on top. The platform offers two strategies:</p><ul class=""><li id="f8c5" class="ot ou jj ov b ow ox oy oz pa pb pc pd fo pe pf pg fr ph pi pj fu pk pl pm pn rq rr rs cv"><strong class="ov jk">Red-Black</strong> deploys a new version alongside the current one. Once the new instance passes health checks, traffic shifts in phases — the new version scales up while the old scales down at the same rate. If any step fails, the system triggers an atomic rollback. Red-Black is the right choice when the model interface is stable. Production revealed a <strong class="ov jk">coordination gap</strong> when a new version requires an I/O schema change (e.g., new tensor dimensions): the upstream consumer can’t update its config until the new model is fully live, so it inevitably sends “old” requests to a “new” deployment during the migration window, and those fail.</li><li id="4508" class="ot ou jj ov b ow rt oy oz pa ru pc pd fo rv pf pg fr rw pi pj fu rx pl pm pn rq rr rs cv"><strong class="ov jk">Versioned</strong> solves that gap by maintaining an independent deployment for every (modelId, modelVersion) pair. Multiple versions serve simultaneously, decoupling model deployment from consumer updates: the consumer waits for the new version to be fully ready before switching its config, while the old version keeps serving legacy traffic. The platform cleans up older deployments after inactivity but always preserves the latest. The trade-off is a <strong class="ov jk">temporary increase in GPU cost</strong> during the transition overlap.</li></ul><p id="30c7" class="pw-post-body-paragraph ot ou jj ov b ow ox oy oz pa pb pc pd fo pe pf pg fr ph pi pj fu pk pl pm pn ik cv"><strong class="ov jk">We recommend embedding variable configurations</strong> (e.g., tensor shapes) <strong class="ov jk">directly into the inference model</strong> to make it version-agnostic, so it can use the cheaper <strong class="ov jk">Red-Black</strong> path. <strong class="ov jk">Versioned</strong> is reserved for the rare cases where a breaking interface change is unavoidable.</p><h2 id="1966" class="pp pq jj z pr ps pt pu fl pv pw px fn py pz qa qb qc qd qe qf qg qh qi qj qk cv">Operational Notes</h2><p id="0bd2" class="pw-post-body-paragraph ot ou jj ov b ow ql oy oz pa qm pc pd fo qn pf pg fr qo pi pj fu qp pl pm pn ik cv">Beyond those four decisions, two operational details are worth flagging — both hit production gaps the design phase didn’t anticipate.</p><h3 id="6bbd" class="rj pq jj z pr fk rk ao fl fm rl aq fn fo rm fp fq fr rn fs ft fu ro fv fw rp cv">Boot sequence</h3><p id="0ad1" class="pw-post-body-paragraph ot ou jj ov b ow ql oy oz pa qm pc pd fo qn pf pg fr qo pi pj fu qp pl pm pn ik cv">Bringing a vLLM-on-Triton instance up involves several coordinated steps before the gRPC port opens. Two are non-routine.</p><ul class=""><li id="8935" class="ot ou jj ov b ow ox oy oz pa pb pc pd fo pe pf pg fr ph pi pj fu pk pl pm pn rq rr rs cv"><strong class="ov jk">Model caching.</strong> Downloading large LLMs directly from S3 or Hugging Face at startup is slow enough to inflate cold-start latency past what schedulers tolerate. We materialize models on Amazon FSx at the time of model announcement, so warm starts hit a high-performance file system instead of object storage.</li><li id="073e" class="ot ou jj ov b ow rt oy oz pa ru pc pd fo rv pf pg fr rw pi pj fu rx pl pm pn rq rr rs cv"><strong class="ov jk">Embedded vs standalone Triton.</strong> When consumers need the OpenAI-compatible API, Triton runs as an embedded server inside the OpenAI-compatible frontend process; otherwise, it runs standalone. This is configured per-deployment at packaging time.</li></ul><p id="baf4" class="pw-post-body-paragraph ot ou jj ov b ow ox oy oz pa pb pc pd fo pe pf pg fr ph pi pj fu pk pl pm pn ik cv">The rest of the boot sequence is mechanical: extracting the model package, installing custom vLLM plugins via Python entry_points, cleaning the Prometheus multiprocess directory, and gating the gRPC port until the engine is ready.</p><h3 id="8033" class="rj pq jj z pr fk rk ao fl fm rl aq fn fo rm fp fq fr rn fs ft fu ro fv fw rp cv">Unified metrics endpoint</h3><p id="1e0d" class="pw-post-body-paragraph ot ou jj ov b ow ql oy oz pa qm pc pd fo qn pf pg fr qo pi pj fu qp pl pm pn ik cv">The Prometheus cleanup above hints at a wider observability gap. vLLM writes metrics to PROMETHEUS_MULTIPROC_DIR as .db files; Triton reports server-level metrics through its own Prometheus endpoint. Neither is aware of the other, and Triton’s built-in bridge surfaces only 9 of 40+ vLLM metrics — missing critical ones like token throughput, KV cache utilization, and prefix cache hit rates.</p><p id="2801" class="pw-post-body-paragraph ot ou jj ov b ow ox oy oz pa pb pc pd fo pe pf pg fr ph pi pj fu pk pl pm pn ik cv">We added a lightweight HTTP proxy that merges both into a single /metrics endpoint: it fetches Triton metrics via HTTP, reads vLLM metrics from disk using Prometheus’s MultiProcessCollector, and returns the combined output. Existing dashboards and alerts work without modification.</p><h2 id="e3a9" class="pp pq jj z pr ps pt pu fl pv pw px fn py pz qa qb qc qd qe qf qg qh qi qj qk cv">Deep-Dive: Constrained Decoding at Scale</h2><p id="0f1e" class="pw-post-body-paragraph ot ou jj ov b ow ql oy oz pa qm pc pd fo qn pf pg fr qo pi pj fu qp pl pm pn ik cv">Some Netflix production workloads rely heavily on fine-grained control over token generation. Rather than applying business logic after inference — paying for invalid generations, then retrying or repairing — we push constraints inside the decode loop, so the model generates outputs that are compliant by construction. We implement this via vLLM’s custom logits processor interface, modeling each constraint as a state machine that evolves with the generated token history and emits token-eligibility masks at each step. Each request gets its own configured processor, since different requests apply different rules.</p><p id="a09a" class="pw-post-body-paragraph ot ou jj ov b ow ox oy oz pa pb pc pd fo pe pf pg fr ph pi pj fu pk pl pm pn ik cv">Getting this to scale ran across two engine versions: we initially deployed on vLLM V0 (V1 had feature gaps), then migrated to V1 in Q4 2025 once it matured. The two subsections that follow are the before-and-after.</p><h3 id="87db" class="rj pq jj z pr fk rk ao fl fm rl aq fn fo rm fp fq fr rn fs ft fu ro fv fw rp cv">Why the first implementation didn’t scale</h3><p id="c810" class="pw-post-body-paragraph ot ou jj ov b ow ql oy oz pa qm pc pd fo qn pf pg fr qo pi pj fu qp pl pm pn ik cv">Our initial pure-Python implementation worked functionally but hit a scaling bottleneck. In vLLM V0, custom logits processors run per-request: the GPU produces logits for the whole batch, the CPU copies them across and waits for the transfer, and then constraint logic runs sequentially for each request — sequentially because the GIL prevents Python from parallelizing the per-request work. CPU time in logit processing therefore grows linearly with batch size, hitting tail latencies. End-to-end latency becomes CPU-bound even though the model’s forward pass is batched efficiently on GPU. It’s a bottleneck invisible in single-request benchmarks that only surfaces under realistic concurrency. Figure 2 makes the serial pattern visible.</p><figure class="qt qu qv qw qx qy qq qr paragraph-image"><div role="button" tabindex="0" class="qz ra dm rb cs rc">Press enter or click to view image in full size<div class="qq qr un"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*KuTzU3O2pm6A4HrUog1c5Q.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*KuTzU3O2pm6A4HrUog1c5Q.png" /><img alt="" class="cs er rd re" width="700" height="194" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="rf bv rg qq qr rh ri z b cr u w">Figure 2: Logits processor serial execution on CPU with vLLM V0</figcaption></figure><h3 id="efd6" class="rj pq jj z pr fk rk ao fl fm rl aq fn fo rm fp fq fr rn fs ft fu ro fv fw rp cv">vLLM V1 enabled a batch-level design</h3><p id="ede5" class="pw-post-body-paragraph ot ou jj ov b ow ql oy oz pa qm pc pd fo qn pf pg fr qo pi pj fu qp pl pm pn ik cv">The structural fix arrived in vLLM V1, which moved logits processing to batch level. We rewrote our custom processor to operate on batch-level data structures, computing masks across many requests together, and reimplemented the hot path in C++ with multi-threading to step around the GIL. The V1 API requires explicit tracking of batch membership changes via update_state(batch_update) — more complex than V0’s per-request interface, but necessary to maintain correct state in a dynamically evolving batch. Figure 3 shows logits processing time staying flat as batch size grows.</p><figure class="qt qu qv qw qx qy qq qr paragraph-image"><div role="button" tabindex="0" class="qz ra dm rb cs rc">Press enter or click to view image in full size<div class="qq qr uo"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*xqPCzoPySS31VGUnJC0Igg.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*xqPCzoPySS31VGUnJC0Igg.png" /><img alt="" class="cs er rd re" width="700" height="259" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="rf bv rg qq qr rh ri z b cr u w">Figure 3: Batched logits processor execution on CPU with vLLM V1</figcaption></figure><h3 id="ad57" class="rj pq jj z pr fk rk ao fl fm rl aq fn fo rm fp fq fr rn fs ft fu ro fv fw rp cv">Operational hardening</h3><p id="f498" class="pw-post-body-paragraph ot ou jj ov b ow ql oy oz pa qm pc pd fo qn pf pg fr qo pi pj fu qp pl pm pn ik cv">Now, performance was no longer the bottleneck. But stateful constraint logic in the decode loop introduced two issues the design phase didn’t anticipate:</p><ul class=""><li id="ad3f" class="ot ou jj ov b ow ox oy oz pa pb pc pd fo pe pf pg fr ph pi pj fu pk pl pm pn rq rr rs cv"><strong class="ov jk">Partial prefills.</strong> V1 performs chunked prefilling, so a request can be prefilled over multiple engine steps. BatchUpdate lacks the granularity to tell whether a request was fully or only partially prefilled, so we added internal tracking.</li><li id="0024" class="ot ou jj ov b ow rt oy oz pa ru pc pd fo rv pf pg fr rw pi pj fu rx pl pm pn rq rr rs cv"><strong class="ov jk">Preemption.</strong> Under memory pressure, vLLM may evict a partially completed request’s KV cache and reschedule it later with a different prompt and output token list. This breaks the state machine’s assumption that the output token list grows monotonically. We detect when the token history shrinks between decode steps, reset the state machine, and reinitialize from the new prompt.</li></ul><h2 id="9e40" class="pp pq jj z pr ps pt pu fl pv pw px fn py pz qa qb qc qd qe qf qg qh qi qj qk cv">Wrap up</h2><p id="41e1" class="pw-post-body-paragraph ot ou jj ov b ow ql oy oz pa qm pc pd fo qn pf pg fr qo pi pj fu qp pl pm pn ik cv">We set out to build an LLM serving platform for broad production ML requirements — low latency, deep customization, and integration with existing infrastructure. The result is a system on vLLM and Triton, unified behind a consistent API, designed to give ML practitioners a fast path from experimentation to production.</p><p id="1b64" class="pw-post-body-paragraph ot ou jj ov b ow ox oy oz pa pb pc pd fo pe pf pg fr ph pi pj fu pk pl pm pn ik cv">The lessons were often in the details — version pinning, silent API gaps, packaging trade-offs — but addressing them has made the platform meaningfully more robust and the developer experience smoother. Next investments reflect where we expect friction:</p><ul class=""><li id="cb82" class="ot ou jj ov b ow ox oy oz pa pb pc pd fo pe pf pg fr ph pi pj fu pk pl pm pn rq rr rs cv"><strong class="ov jk">System prompt compression</strong> to reduce prompt length without sacrificing quality.</li><li id="a29f" class="ot ou jj ov b ow rt oy oz pa ru pc pd fo rv pf pg fr rw pi pj fu rx pl pm pn rq rr rs cv"><strong class="ov jk">Asynchronous scheduling</strong> of vLLM V1.</li><li id="6cc8" class="ot ou jj ov b ow rt oy oz pa ru pc pd fo rv pf pg fr rw pi pj fu rx pl pm pn rq rr rs cv"><strong class="ov jk">Vectorized logits processors</strong> that run as fused GPU kernels instead of CPU code.</li><li id="a2c4" class="ot ou jj ov b ow rt oy oz pa ru pc pd fo rv pf pg fr rw pi pj fu rx pl pm pn rq rr rs cv"><strong class="ov jk">Lower-precision model variants</strong> to decrease memory footprint and increase throughput.</li></ul><p id="b0fe" class="pw-post-body-paragraph ot ou jj ov b ow ox oy oz pa pb pc pd fo pe pf pg fr ph pi pj fu pk pl pm pn ik cv">We’ll continue working closely with the open-source community as this space evolves.</p><h2 id="3120" class="pp pq jj z pr ps pt pu fl pv pw px fn py pz qa qb qc qd qe qf qg qh qi qj qk cv">Contributions</h2><p id="27bd" class="pw-post-body-paragraph ot ou jj ov b ow ql oy oz pa qm pc pd fo qn pf pg fr qo pi pj fu qp pl pm pn ik cv">This system is the result of close collaboration and contributions from many teams within the AI Platform org at Netflix. In particular, Liping Peng designed and developed the model packaging workflow and drove the integration of Triton and vLLM with MSS to enable a unified pathway for serving LLMs. Hakan Baba, Nicolas Hortiguera, and ZQ Zhang led GPU capacity planning, system performance tuning, application integration and observability, as well as A/B test readiness and operational excellence efforts for all production models. Santino Ramos enabled vLLM for production models and optimized constrained decoding performance. Binh Tang developed the initial version of custom model serving and benchmarked different LLM serving frameworks. Lanxi Huang and Daneo Zhang built the serving development tools to enable user self-service. Lingyi Liu drove the overall system architecture and core technical decisions. Abhishek Agrawal and Shaojing Li provide management leadership to ensure alignment, prioritization and execution.</p><h2 id="6f6a" class="pp pq jj z pr ps pt pu fl pv pw px fn py pz qa qb qc qd qe qf qg qh qi qj qk cv">Acknowledgements</h2><p id="3897" class="pw-post-body-paragraph ot ou jj ov b ow ql oy oz pa qm pc pd fo qn pf pg fr qo pi pj fu qp pl pm pn ik cv">This work heavily leverages open-source ML libraries, such as Triton, vLLM and PyTorch, etc. We’re especially grateful to the teams and contributors from the community. We also thank our partner teams in Netflix AI for Member Systems for their close collaborations and innovation on the modeling side.</p></div></div></div></div></div></section></div></div></article><article aria-label="GenPage: Towards End-to-End Generative Homepage Construction at Netflix" class="al" data-testid="post-preview"><div class="al uy e"><div class="cs"><div class="al e"><div class="dm al abb abc abd abe abf abg abh abi abj abk abl abm abn"><div class="abo"><div aria-label="GenPage: Towards End-to-End Generative Homepage Construction at Netflix"><div class="abq abr abs abt ek"><img alt="GenPage: Towards End-to-End Generative Homepage Construction at Netflix" class="cs abu abv abw abx" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*NurrizMgC7_QbsGuuW42Dg.gif" /></div></div></div><div class="abp am ew gh"><div class="am gh ug cs aby abz aca acb"><div class="acc acd ace acf acg fd am j"><div class="fd am j ach"><div class="aci e ej"><div><div class="e"><div tabindex="-1" class="cq"><a href="https://netflixtechblog.com" rel="noopener follow"><div class="dm"><img alt="Netflix TechBlog" class="ei if e ack acj" src="https://miro.medium.com/v2/resize:fill:40:40/1*ty4NvNrGg4ReETxqU2N3Og.png" width="20" height="20" /></div></a></div></div></div></div><div class="wg e ej dd"><p class="z b y u w">In</p></div><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai lf am j" href="https://netflixtechblog.com" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">Netflix TechBlog</p></a></div></div></div></div></div><div class="acl e ej dd"><p class="z b y u w">by</p></div><div class="fd am j ach"><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai lf am j" href="https://netflixtechblog.medium.com" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">Netflix Technology Blog</p></a></div></div></div></div></div><div class="nn am j ej db dd"><p class="z b y u w">·</p><p class="z b y u w">Jun 29</p></div></div><div class="acm e acn aco acp acq acr ik"><div class="acs act acu acv acw acx acy acz"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" rel="noopener follow" href="https://netflixtechblog.com/genpage-towards-end-to-end-generative-homepage-construction-at-netflix-77146fba8a08" data-discover="true"><div title="GenPage: Towards End-to-End Generative Homepage Construction at Netflix"><h2 class="z jk ps pu ada adb adc add ade adf fl pv px adg adh adi adj adk adl fn fo fp adm adn ado adp adq adr fq fr fs ads adt adu adv adw adx ft fu fv ady adz aea aeb aec aed fw hv cv">GenPage: Towards End-to-End Generative Homepage Construction at Netflix</h2></div><div class="aee e"><h3 class="z b fz u dv aef es hr aeg ht hv w">Authors: Lequn Wang, Jiangwei Pan, and Linas Baltrunas</h3></div></a></div></div><div class="vx am kk ce"><div class="am j kv"><div class="ck cz da r s"><div class="dm aeh xh am j"><div class="dr go aei am j kv"><a class="dr hd aei am j kv" tabindex="-1" rel="noopener follow" href="https://netflixtechblog.com/genpage-towards-end-to-end-generative-homepage-construction-at-netflix-77146fba8a08" data-discover="true"><div><div class="am"><div tabindex="-1" class="cq"><div class="am j db">A response icon11</div></div></div></div></a></div></div></div><div class="am j aek ael"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fin-house-llm-serving-at-netflix-a5a8e799ea2c" rel="noopener follow"><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></a><div class="ck cz da r s"><div><div class="bt"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2F77146fba8a08&amp;operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fgenpage-towards-end-to-end-generative-homepage-construction-at-netflix-77146fba8a08" rel="noopener follow"></a></div></div></div></div></div></div><div class="da r s"></div></div></div></div></div></div></div></div></article><article aria-label="How Netflix Simplified Batch Compute with Kueue" class="al" data-testid="post-preview"><div class="al uy e"><div class="cs"><div class="al e"><div class="dm al abb abc abd abe abf abg abh abi abj abk abl abm abn"><div class="abo"><div aria-label="How Netflix Simplified Batch Compute with Kueue"><div class="abq abr abs abt ek"><img alt="How Netflix Simplified Batch Compute with Kueue" class="cs abu abv aeq abx" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*GL_sNsqautjh4lnu5XmRww.png" /></div></div></div><div class="abp am ew gh"><div class="am gh ug cs aby abz aca acb"><div class="acc acd ace acf acg fd am j"><div class="fd am j ach"><div class="aci e ej"><div><div class="e"><div tabindex="-1" class="cq"><a href="https://netflixtechblog.com" rel="noopener follow"><div class="dm"><img alt="Netflix TechBlog" class="ei if e ack acj" src="https://miro.medium.com/v2/resize:fill:40:40/1*ty4NvNrGg4ReETxqU2N3Og.png" width="20" height="20" /></div></a></div></div></div></div><div class="wg e ej dd"><p class="z b y u w">In</p></div><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai lf am j" href="https://netflixtechblog.com" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">Netflix TechBlog</p></a></div></div></div></div></div><div class="acl e ej dd"><p class="z b y u w">by</p></div><div class="fd am j ach"><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai lf am j" href="https://netflixtechblog.medium.com" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">Netflix Technology Blog</p></a></div></div></div></div></div><div class="nn am j ej db dd"><p class="z b y u w">·</p><p class="z b y u w">Jun 22</p></div></div><div class="acm e acn aco acp acq acr ik"><div class="acs act acu acv acw acx acy acz"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" rel="noopener follow" href="https://netflixtechblog.com/how-netflix-simplified-batch-compute-with-kueue-87860682629c" data-discover="true"><div title=""><h2 class="z jk ps pu ada adb adc add ade adf fl pv px adg adh adi adj adk adl fn fo fp adm adn ado adp adq adr fq fr fs ads adt adu adv adw adx ft fu fv ady adz aea aeb aec aed fw hv cv">How Netflix Simplified Batch Compute with Kueue</h2></div><div class="aee e"><h3 class="z b fz u dv aef es hr aeg ht hv w">By Alvin Bao, Alex Petrov, Jennifer Lai, Aidan Sherr, and Samartha Chandrashekar</h3></div></a></div></div><div class="vx am kk ce"><div class="am j kv"><div class="ck cz da r s"><div class="dm aeh xh am j"><div class="dr go aei am j kv"></div></div></div><div class="am j aek ael"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fin-house-llm-serving-at-netflix-a5a8e799ea2c" rel="noopener follow"><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></a><div class="ck cz da r s"><div><div class="bt"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2F87860682629c&amp;operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fhow-netflix-simplified-batch-compute-with-kueue-87860682629c" rel="noopener follow"></a></div></div></div></div></div></div><div class="da r s"></div></div></div></div></div></div></div></div></article><article aria-label="Building Service Topology at Scale: Architecture, Challenges, and Lessons Learned" class="al" data-testid="post-preview"><div class="al uy e"><div class="cs"><div class="al e"><div class="dm al abb abc abd abe abf abg abh abi abj abk abl abm abn"><div class="abo"><div aria-label="Diagram showing two separate ingestion pipelines — a flow log pipeline and an IPC pipeline, each fed by data enrichment — writing to their own graph store, with a shared API serving UI and backend clients"><div class="abq abr abs abt ek"><img alt="Diagram showing two separate ingestion pipelines — a flow log pipeline and an IPC pipeline, each fed by data enrichment — writing to their own graph store, with a shared API serving UI and backend clients" class="cs abu abv abw abx" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*5_09VEnyPtFR0tbO1zzzMQ.png" /></div></div></div><div class="abp am ew gh"><div class="am gh ug cs aby abz aca acb"><div class="acc acd ace acf acg fd am j"><div class="fd am j ach"><div class="aci e ej"><div><div class="e"><div tabindex="-1" class="cq"><a href="https://netflixtechblog.com" rel="noopener follow"><div class="dm"><img alt="Netflix TechBlog" class="ei if e ack acj" src="https://miro.medium.com/v2/resize:fill:40:40/1*ty4NvNrGg4ReETxqU2N3Og.png" width="20" height="20" /></div></a></div></div></div></div><div class="wg e ej dd"><p class="z b y u w">In</p></div><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai lf am j" href="https://netflixtechblog.com" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">Netflix TechBlog</p></a></div></div></div></div></div><div class="acl e ej dd"><p class="z b y u w">by</p></div><div class="fd am j ach"><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai lf am j" href="https://netflixtechblog.medium.com" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">Netflix Technology Blog</p></a></div></div></div></div></div><div class="nn am j ej db dd"><p class="z b y u w">·</p><p class="z b y u w">3d ago</p></div></div><div class="acm e acn aco acp acq acr ik"><div class="acs act acu acv acw acx acy acz"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" rel="noopener follow" href="https://netflixtechblog.com/building-service-topology-at-scale-architecture-challenges-and-lessons-learned-f4b792f3f0d8" data-discover="true"><div title="Building Service Topology at Scale: Architecture, Challenges, and Lessons Learned"><h2 class="z jk ps pu ada adb adc add ade adf fl pv px adg adh adi adj adk adl fn fo fp adm adn ado adp adq adr fq fr fs ads adt adu adv adw adx ft fu fv ady adz aea aeb aec aed fw hv cv">Building Service Topology at Scale: Architecture, Challenges, and Lessons Learned</h2></div><div class="aee e"><h3 class="z b fz u dv aef es hr aeg ht hv w">By Parth Jain, Rakesh Sukumar, Yingwu Zhao, Renzo Sanchez-Silva &amp; Nathan Fisher. A deep dive into the engineering challenges of building a…</h3></div></a></div></div><div class="vx am kk ce"><div class="am j kv"><div class="ck cz da r s"><div class="dm aeh xh am j"><div class="dr go aei am j kv"><a class="dr hd aei am j kv" tabindex="-1" rel="noopener follow" href="https://netflixtechblog.com/building-service-topology-at-scale-architecture-challenges-and-lessons-learned-f4b792f3f0d8" data-discover="true"><div><div class="am"><div tabindex="-1" class="cq"><div class="am j db">A response icon5</div></div></div></div></a></div></div></div><div class="am j aek ael"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fin-house-llm-serving-at-netflix-a5a8e799ea2c" rel="noopener follow"><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></a><div class="ck cz da r s"><div><div class="bt"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2Ff4b792f3f0d8&amp;operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fbuilding-service-topology-at-scale-architecture-challenges-and-lessons-learned-f4b792f3f0d8" rel="noopener follow"></a></div></div></div></div></div></div><div class="da r s"></div></div></div></div></div></div></div></div></article><article aria-label="VMAF v1: Good Is Not Good Enough" class="al" data-testid="post-preview"><div class="al uy e"><div class="cs"><div class="al e"><div class="dm al abb abc abd abe abf abg abh abi abj abk abl abm abn"><div class="abo"><div aria-label="VMAF v1: Good Is Not Good Enough"><div class="abq abr abs abt ek"><img alt="VMAF v1: Good Is Not Good Enough" class="cs abu abv abw abx" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*4Ymg7tqfTpT3SUSQX92smw.png" /></div></div></div><div class="abp am ew gh"><div class="am gh ug cs aby abz aca acb"><div class="acc acd ace acf acg fd am j"><div class="fd am j ach"><div class="aci e ej"><div><div class="e"><div tabindex="-1" class="cq"><a href="https://netflixtechblog.com" rel="noopener follow"><div class="dm"><img alt="Netflix TechBlog" class="ei if e ack acj" src="https://miro.medium.com/v2/resize:fill:40:40/1*ty4NvNrGg4ReETxqU2N3Og.png" width="20" height="20" /></div></a></div></div></div></div><div class="wg e ej dd"><p class="z b y u w">In</p></div><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai lf am j" href="https://netflixtechblog.com" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">Netflix TechBlog</p></a></div></div></div></div></div><div class="acl e ej dd"><p class="z b y u w">by</p></div><div class="fd am j ach"><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai lf am j" href="https://netflixtechblog.medium.com" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">Netflix Technology Blog</p></a></div></div></div></div></div><div class="nn am j ej db dd"><p class="z b y u w">·</p><p class="z b y u w">Jun 19</p></div></div><div class="acm e acn aco acp acq acr ik"><div class="acs act acu acv acw acx acy acz"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" rel="noopener follow" href="https://netflixtechblog.com/vmaf-v1-good-is-not-good-enough-60d7e4244ea8" data-discover="true"><div title=""><h2 class="z jk ps pu ada adb adc add ade adf fl pv px adg adh adi adj adk adl fn fo fp adm adn ado adp adq adr fq fr fs ads adt adu adv adw adx ft fu fv ady adz aea aeb aec aed fw hv cv">VMAF v1: Good Is Not Good Enough</h2></div><div class="aee e"><h3 class="z b fz u dv aef es hr aeg ht hv w">By Christos G. Bampis, Zhi Li, Kyle Swanson, Nil Fons Miret and Pavan Madhusudanarao</h3></div></a></div></div><div class="vx am kk ce"><div class="am j kv"><div class="ck cz da r s"><div class="dm aeh xh am j"><div class="dr go aei am j kv"></div></div></div><div class="am j aek ael"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fin-house-llm-serving-at-netflix-a5a8e799ea2c" rel="noopener follow"><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></a><div class="ck cz da r s"><div><div class="bt"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2F60d7e4244ea8&amp;operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fvmaf-v1-good-is-not-good-enough-60d7e4244ea8" rel="noopener follow"></a></div></div></div></div></div></div></div></div></div></div></div></div></div></article><article aria-label="MCP is Dead" class="al" data-testid="post-preview"><div class="al uy e"><div class="cs"><div class="al e"><div class="dm al abb abc abd abe abf abg abh abi abj abk abl abm abn"><div class="abo"><div aria-label="MCP is Dead"><div class="abq abr abs abt ek"><img alt="MCP is Dead" class="cs abu abv abw abx" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*Oj5PiyfEi8DadSC8Jy374w.png" /></div></div></div><div class="abp am ew gh"><div class="am gh ug cs aby abz aca acb"><div class="acc acd ace acf acg fd am j"><div class="fd am j ach"><div class="aci e ej"><div><div class="e"><div tabindex="-1" class="cq"><a href="https://uxplanet.org" rel="noopener follow"><div class="dm"><img alt="UX Planet" class="ei if e ack acj" src="https://miro.medium.com/v2/resize:fill:40:40/1*A0FnBy5FBoVQC02SZXLXPg.png" width="20" height="20" /></div></a></div></div></div></div><div class="wg e ej dd"><p class="z b y u w">In</p></div><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai lf am j" href="https://uxplanet.org" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">UX Planet</p></a></div></div></div></div></div><div class="acl e ej dd"><p class="z b y u w">by</p></div><div class="fd am j ach"><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai lf am j" href="https://medium.com/@101" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">Nick Babich</p><div class="afx afy e"><div class="am afz"><div class="am"></div></div></div></a></div></div></div></div></div><div class="nn am j ej db dd"><p class="z b y u w">·</p><p class="z b y u w">Apr 6</p></div></div><div class="acm e acn aco acp acq acr ik"><div class="acs act acu acv acw acx acy acz"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://uxplanet.org/mcp-is-dead-cf16b667ba6d" rel="noopener follow"><div title=""><h2 class="z jk ps pu ada adb adc add ade adf fl pv px adg adh adi adj adk adl fn fo fp adm adn ado adp adq adr fq fr fs ads adt adu adv adw adx ft fu fv ady adz aea aeb aec aed fw hv cv">MCP is Dead</h2></div><div class="aee e"><h3 class="z b fz u dv aef es hr aeg ht hv w">Why you should avoid using MCP in Claude Code and what to use instead</h3></div></a></div></div><div class="vx am kk ce"><div class="am j kv"><div class="uy am"><button class="e cf af ac" aria-label="Member-only story"><div class=""><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></div></button></div><div class="ck cz da r s"><div class="dm aeh xh am j"><div class="dr go aei am j kv"><a class="dr hd aei am j kv" tabindex="-1" href="https://uxplanet.org/mcp-is-dead-cf16b667ba6d" rel="noopener follow"><div><div class="am"><div tabindex="-1" class="cq"><div class="am j db">A response icon275</div></div></div></div></a></div></div></div><div class="am j aek ael"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fin-house-llm-serving-at-netflix-a5a8e799ea2c" rel="noopener follow"><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></a><div class="ck cz da r s"><div><div class="bt"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2Fcf16b667ba6d&amp;operation=register&amp;redirect=https%3A%2F%2Fuxplanet.org%2Fmcp-is-dead-cf16b667ba6d" rel="noopener follow"></a></div></div></div></div></div></div><div class="da r s"></div></div></div></div></div></div></div></div></article><article aria-label="A Step-by-Step Guide for Developing Your Personal Agentic System." class="al" data-testid="post-preview"><div class="al uy e"><div class="cs"><div class="al e"><div class="dm al abb abc abd abe abf abg abh abi abj abk abl abm abn"><div class="abo"><div aria-label="A Step-by-Step Guide for Developing Your Personal Agentic System."><div class="abq abr abs abt ek"><img alt="A Step-by-Step Guide for Developing Your Personal Agentic System." class="cs abu abv abw abx" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*po63mpOfUFmrajkBRmDkGw.png" /></div></div></div><div class="abp am ew gh"><div class="am gh ug cs aby abz aca acb"><div class="acc acd ace acf acg fd am j"><div class="fd am j ach"><div class="aci e ej"><div><div class="e"><div tabindex="-1" class="cq"><a href="https://medium.com/data-science-collective" rel="noopener follow"><div class="dm"><img alt="Data Science Collective" class="ei if e ack acj" src="https://miro.medium.com/v2/resize:fill:40:40/1*0nV0Q-FBHj94Kggq00pG2Q.jpeg" width="20" height="20" /></div></a></div></div></div></div><div class="wg e ej dd"><p class="z b y u w">In</p></div><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai lf am j" href="https://medium.com/data-science-collective" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">Data Science Collective</p></a></div></div></div></div></div><div class="acl e ej dd"><p class="z b y u w">by</p></div><div class="fd am j ach"><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai lf am j" href="https://erdogant.medium.com" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">Erdogan T</p></a></div></div></div></div></div><div class="nn am j ej db dd"><p class="z b y u w">·</p><p class="z b y u w">Jun 12</p></div></div><div class="acm e acn aco acp acq acr ik"><div class="acs act acu acv acw acx acy acz"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/data-science-collective/a-step-by-step-guide-for-developing-your-personal-agentic-system-24c6cd6fa849" rel="noopener follow"><div title=""><h2 class="z jk ps pu ada adb adc add ade adf fl pv px adg adh adi adj adk adl fn fo fp adm adn ado adp adq adr fq fr fs ads adt adu adv adw adx ft fu fv ady adz aea aeb aec aed fw hv cv">A Step-by-Step Guide for Developing Your Personal Agentic System.</h2></div><div class="aee e"><h3 class="z b fz u dv aef es hr aeg ht hv w">A complete guide to learn how to set up and create your own agentic LLM system with local databases and specialized for your task.</h3></div></a></div></div><div class="vx am kk ce"><div class="am j kv"><div class="uy am"><button class="e cf af ac" aria-label="Member-only story"><div class=""><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></div></button></div><div class="ck cz da r s"><div class="dm aeh xh am j"><div class="dr go aei am j kv"><a class="dr hd aei am j kv" tabindex="-1" href="https://medium.com/data-science-collective/a-step-by-step-guide-for-developing-your-personal-agentic-system-24c6cd6fa849" rel="noopener follow"><div><div class="am"><div tabindex="-1" class="cq"><div class="am j db">A response icon28</div></div></div></div></a></div></div></div><div class="am j aek ael"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fin-house-llm-serving-at-netflix-a5a8e799ea2c" rel="noopener follow"><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></a><div class="ck cz da r s"><div><div class="bt"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2F24c6cd6fa849&amp;operation=register&amp;redirect=https%3A%2F%2Fmedium.com%2Fdata-science-collective%2Fa-step-by-step-guide-for-developing-your-personal-agentic-system-24c6cd6fa849" rel="noopener follow"></a></div></div></div></div></div></div></div></div></div></div></div></div></div></article><article aria-label="How to Create Loops with Claude Code: A Practical Guide to Agentic Automation" class="al" data-testid="post-preview"><div class="al uy e"><div class="cs"><div class="al e"><div class="dm al abb abc abd abe abf abg abh abi abj abk abl abm abn"><div class="abo"><div aria-label="How to Create Loops with Claude Code: A Practical Guide to Agentic Automation"><div class="abq abr abs abt ek"><img alt="How to Create Loops with Claude Code: A Practical Guide to Agentic Automation" class="cs abu abv abw abx" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*WbMmltJlf6xsaKxctAiUpg.png" /></div></div></div><div class="abp am ew gh"><div class="am gh ug cs aby abz aca acb"><div class="acc acd ace acf acg fd am j"><div class="fd am j ach"><div class="aci e ej"><div><div class="e"><div tabindex="-1" class="cq"><a href="https://levelup.gitconnected.com" rel="noopener follow"><div class="dm"><img alt="Level Up Coding" class="ei if e ack acj" src="https://miro.medium.com/v2/resize:fill:40:40/1*5D9oYBd58pyjMkV_5-zXXQ.jpeg" width="20" height="20" /></div></a></div></div></div></div><div class="wg e ej dd"><p class="z b y u w">In</p></div><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai lf am j" href="https://levelup.gitconnected.com" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">Level Up Coding</p></a></div></div></div></div></div><div class="acl e ej dd"><p class="z b y u w">by</p></div><div class="fd am j ach"><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai lf am j" href="https://yousefhosni.medium.com" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">Youssef Hosni</p></a></div></div></div></div></div><div class="nn am j ej db dd"><p class="z b y u w">·</p><p class="z b y u w">Jun 24</p></div></div><div class="acm e acn aco acp acq acr ik"><div class="acs act acu acv acw acx acy acz"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://levelup.gitconnected.com/how-to-create-loops-with-claude-code-a-practical-guide-to-agentic-automation-6f422390a143" rel="noopener follow"><div title="How to Create Loops with Claude Code: A Practical Guide to Agentic Automation"><h2 class="z jk ps pu ada adb adc add ade adf fl pv px adg adh adi adj adk adl fn fo fp adm adn ado adp adq adr fq fr fs ads adt adu adv adw adx ft fu fv ady adz aea aeb aec aed fw hv cv">How to Create Loops with Claude Code: A Practical Guide to Agentic Automation</h2></div><div class="aee e"><h3 class="z b fz u dv aef es hr aeg ht hv w">Claude Loop Engineering: How to Design Agents That Iterate, Verify, and Remember</h3></div></a></div></div><div class="vx am kk ce"><div class="am j kv"><div class="uy am"><button class="e cf af ac" aria-label="Member-only story"><div class=""><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></div></button></div><div class="ck cz da r s"><div class="dm aeh xh am j"><div class="dr go aei am j kv"><a class="dr hd aei am j kv" tabindex="-1" href="https://levelup.gitconnected.com/how-to-create-loops-with-claude-code-a-practical-guide-to-agentic-automation-6f422390a143" rel="noopener follow"><div><div class="am"><div tabindex="-1" class="cq"><div class="am j db">A response icon13</div></div></div></div></a></div></div></div><div class="am j aek ael"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fin-house-llm-serving-at-netflix-a5a8e799ea2c" rel="noopener follow"><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></a><div class="ck cz da r s"><div><div class="bt"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2F6f422390a143&amp;operation=register&amp;redirect=https%3A%2F%2Flevelup.gitconnected.com%2Fhow-to-create-loops-with-claude-code-a-practical-guide-to-agentic-automation-6f422390a143" rel="noopener follow"></a></div></div></div></div></div></div><div class="da r s"></div></div></div></div></div></div></div></div></article><article aria-label="You only have weeks left to vibe code" class="al" data-testid="post-preview"><div class="al uy e"><div class="cs"><div class="al e"><div class="dm al abb abc abd abe abf abg abh abi abj abk abl abm abn"><div class="abo"><div aria-label="Vibe coding is really over this time, claude terminal window telling you to stop"><div class="abq abr abs abt ek"><img alt="Vibe coding is really over this time, claude terminal window telling you to stop" class="cs abu abv abw abx" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*zKR3rUwFqBImPUfH9OdFAw.png" /></div></div></div><div class="abp am ew gh"><div class="am gh ug cs aby abz aca acb"><div class="acc acd ace acf acg fd am j"><div class="fd am j ach"><div class="aci e ej"><div><div class="e"><div tabindex="-1" class="cq"><a tabindex="-1" href="https://michalmalewicz.medium.com" rel="noopener follow"><div class="e dm"><img alt="Michal Malewicz" class="e bs ee acj ack ei" src="https://miro.medium.com/v2/resize:fill:40:40/1*149zXrb2FXvS_mctL4NKSg.png" width="20" height="20" /></div></a></div></div></div></div><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai lf am j" href="https://michalmalewicz.medium.com" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">Michal Malewicz</p><div class="afx afy e"><div class="am afz"><div class="am"></div></div></div></a></div></div></div></div></div><div class="nn am j ej db dd"><p class="z b y u w">·</p><p class="z b y u w">Jun 29</p></div></div><div class="acm e acn aco acp acq acr ik"><div class="acs act acu acv acw acx acy acz"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://michalmalewicz.medium.com/you-only-have-weeks-left-to-vibe-code-09a89c3d3f9b" rel="noopener follow"><div title=""><h2 class="z jk ps pu ada adb adc add ade adf fl pv px adg adh adi adj adk adl fn fo fp adm adn ado adp adq adr fq fr fs ads adt adu adv adw adx ft fu fv ady adz aea aeb aec aed fw hv cv">You only have weeks left to vibe code</h2></div><div class="aee e"><h3 class="z b fz u dv aef es hr aeg ht hv w">Then it’s over. You better hurry up!</h3></div></a></div></div><div class="vx am kk ce"><div class="am j kv"><div class="uy am"><button class="e cf af ac" aria-label="Member-only story"><div class=""><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></div></button></div><div class="ck cz da r s"><div class="dm aeh xh am j"><div class="dr go aei am j kv"><a class="dr hd aei am j kv" tabindex="-1" href="https://michalmalewicz.medium.com/you-only-have-weeks-left-to-vibe-code-09a89c3d3f9b" rel="noopener follow"><div><div class="am"><div tabindex="-1" class="cq"><div class="am j db">A response icon205</div></div></div></div></a></div></div></div><div class="am j aek ael"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fin-house-llm-serving-at-netflix-a5a8e799ea2c" rel="noopener follow"><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></a><div class="ck cz da r s"><div><div class="bt"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2F09a89c3d3f9b&amp;operation=register&amp;redirect=https%3A%2F%2Fmichalmalewicz.medium.com%2Fyou-only-have-weeks-left-to-vibe-code-09a89c3d3f9b" rel="noopener follow"></a></div></div></div></div></div></div><div class="da r s"></div></div></div></div></div></div></div></div></article><article aria-label="Building tools to manage cross-cutting concerns in Kubernetes" class="al" data-testid="post-preview"><div class="al uy e"><div class="cs"><div class="al e"><div class="dm al abb abc abd abe abf abg abh abi abj abk abl abm abn"><div class="abo"><div aria-label="Building tools to manage cross-cutting concerns in Kubernetes"><div class="abq abr abs abt ek"><img alt="Building tools to manage cross-cutting concerns in Kubernetes" class="cs abu abv abw abx" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*W5wsjHgClkt-m3b5jIAU0Q.png" /></div></div></div><div class="abp am ew gh"><div class="am gh ug cs aby abz aca acb"><div class="acc acd ace acf acg fd am j"><div class="fd am j ach"><div class="aci e ej"><div><div class="e"><div tabindex="-1" class="cq"><a href="https://itnext.io" rel="noopener follow"><div class="dm"><img alt="ITNEXT" class="ei if e ack acj" src="https://miro.medium.com/v2/resize:fill:40:40/1*yAqDFIFA5F_NXalOJKz4TA.png" width="20" height="20" /></div></a></div></div></div></div><div class="wg e ej dd"><p class="z b y u w">In</p></div><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai lf am j" href="https://itnext.io" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">ITNEXT</p></a></div></div></div></div></div><div class="acl e ej dd"><p class="z b y u w">by</p></div><div class="fd am j ach"><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai lf am j" href="https://medium.com/@bgrant0607" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">Brian Grant</p></a></div></div></div></div></div><div class="nn am j ej db dd"><p class="z b y u w">·</p><p class="z b y u w">4d ago</p></div></div><div class="acm e acn aco acp acq acr ik"><div class="acs act acu acv acw acx acy acz"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://itnext.io/building-tools-to-manage-cross-cutting-concerns-in-kubernetes-by-using-confighub-6bd64afcdd69" rel="noopener follow"><div title=""><h2 class="z jk ps pu ada adb adc add ade adf fl pv px adg adh adi adj adk adl fn fo fp adm adn ado adp adq adr fq fr fs ads adt adu adv adw adx ft fu fv ady adz aea aeb aec aed fw hv cv">Building tools to manage cross-cutting concerns in Kubernetes</h2></div><div class="aee e"><h3 class="z b fz u dv aef es hr aeg ht hv w">We can build composable, interoperable tools to automate cross-cutting configuration changes on top of configuration as data.</h3></div></a></div></div><div class="vx am kk ce"><div class="am j kv"><div class="ck cz da r s"><div class="dm aeh xh am j"><div class="dr go aei am j kv"></div></div></div><div class="am j aek ael"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fin-house-llm-serving-at-netflix-a5a8e799ea2c" rel="noopener follow"><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></a><div class="ck cz da r s"><div><div class="bt"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2F6bd64afcdd69&amp;operation=register&amp;redirect=https%3A%2F%2Fitnext.io%2Fbuilding-tools-to-manage-cross-cutting-concerns-in-kubernetes-by-using-confighub-6bd64afcdd69" rel="noopener follow"></a></div></div></div></div></div></div><div class="da r s"></div></div></div></div></div></div></div></div></article><article aria-label="The Day a Google L7 Engineer Tore My System Design to Shreds" class="al" data-testid="post-preview"><div class="al uy e"><div class="cs"><div class="al e"><div class="dm al abb abc abd abe abf abg abh abi abj abk abl abm abn"><div class="abo"><div aria-label="The Day a Google L7 Engineer Tore My System Design to Shreds"><div class="abq abr abs abt ek"><img alt="The Day a Google L7 Engineer Tore My System Design to Shreds" class="cs abu abv abw abx" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*NVtpplR8n3E_HuwZG88Lmw.png" /></div></div></div><div class="abp am ew gh"><div class="am gh ug cs aby abz aca acb"><div class="acc acd ace acf acg fd am j"><div class="fd am j ach"><div class="aci e ej"><div><div class="e"><div tabindex="-1" class="cq"><a tabindex="-1" href="https://cloudwithazeem.medium.com" rel="noopener follow"><div class="e dm"><img alt="Cloud With Azeem" class="e bs ee acj ack ei" src="https://miro.medium.com/v2/resize:fill:40:40/1*oJWwUx75Cf5oGoEfAefJpw.png" width="20" height="20" /></div></a></div></div></div></div><div class="fd e"><div><div class="e"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai lf am j" href="https://cloudwithazeem.medium.com" rel="noopener follow"><p class="z b y u dv hq es hr hs ht hu hv cv">Cloud With Azeem</p></a></div></div></div></div></div><div class="nn am j ej db dd"><p class="z b y u w">·</p><p class="z b y u w">Feb 16</p></div></div><div class="acm e acn aco acp acq acr ik"><div class="acs act acu acv acw acx acy acz"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://cloudwithazeem.medium.com/google-l7-system-design-interview-lessons-0b3834fded07" rel="noopener follow"><div title=""><h2 class="z jk ps pu ada adb adc add ade adf fl pv px adg adh adi adj adk adl fn fo fp adm adn ado adp adq adr fq fr fs ads adt adu adv adw adx ft fu fv ady adz aea aeb aec aed fw hv cv">The Day a Google L7 Engineer Tore My System Design to Shreds</h2></div><div class="aee e"><h3 class="z b fz u dv aef es hr aeg ht hv w">If you are not a medium member, read here for free</h3></div></a></div></div><div class="vx am kk ce"><div class="am j kv"><div class="uy am"><button class="e cf af ac" aria-label="Member-only story"><div class=""><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></div></button></div><div class="ck cz da r s"><div class="dm aeh xh am j"><div class="dr go aei am j kv"><a class="dr hd aei am j kv" tabindex="-1" href="https://cloudwithazeem.medium.com/google-l7-system-design-interview-lessons-0b3834fded07" rel="noopener follow"><div><div class="am"><div tabindex="-1" class="cq"><div class="am j db">A response icon49</div></div></div></div></a></div></div></div><div class="am j aek ael"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fin-house-llm-serving-at-netflix-a5a8e799ea2c" rel="noopener follow"><div><div class="bt"><div tabindex="-1" class="cq"></div></div></div></a><div class="ck cz da r s"><div><div class="bt"><div tabindex="-1" class="cq"><a class="bx x by bz ca ab cb ac ae af ag ah ai aj ak" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2F0b3834fded07&amp;operation=register&amp;redirect=https%3A%2F%2Fcloudwithazeem.medium.com%2Fgoogle-l7-system-design-interview-lessons-0b3834fded07" rel="noopener follow"></a></div></div></div></div></div></div></div></div></div></div></div></div></div></article>]]></description>
      <link>https://netflixtechblog.com/in-house-llm-serving-at-netflix-a5a8e799ea2c</link>
      <guid>https://netflixtechblog.com/in-house-llm-serving-at-netflix-a5a8e799ea2c</guid>
      <pubDate>Fri, 17 Jul 2026 23:32:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Building Service Topology at Scale: Architecture, Challenges, and Lessons Learned]]></title>
      <description><![CDATA[<article><div class="e"><div class="e"><section><div><div><div class="an ex"><div class="fe ct ir is it iu"><div class="je e"><div class="jf an dp"><div><div class="bu"><div tabindex="-1" class="cr"><div class="jh k ji iw hc dp ag"><p class="ab b cs v x">Featured</p></div></div></div></div></div></div></div></div></div><div class="il jj jk jl jm"><div class="an ex"><div class="fe ct ir is it iu"><div><div></div><p id="0aa9" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><em class="po">By </em><a class="by pp" href="https://www.linkedin.com/in/parth-jain-8a09abb6/" rel="noopener ugc nofollow" target="_blank"><em class="po">Parth Jain</em></a>, <a class="by pp" href="https://www.linkedin.com/in/raskuma/" rel="noopener ugc nofollow" target="_blank"><em class="po">Rakesh Sukumar</em></a><em class="po">, </em><a class="by pp" href="https://www.linkedin.com/in/yingwu-zhao-62037418/" rel="noopener ugc nofollow" target="_blank"><em class="po">Yingwu Zhao</em></a><em class="po">, </em><a class="by pp" href="https://www.linkedin.com/in/renzosanchezsilva/" rel="noopener ugc nofollow" target="_blank"><em class="po">Renzo Sanchez-Silva</em></a><em class="po"> &amp; </em><a class="by pp" href="https://www.linkedin.com/in/nathfisher/" rel="noopener ugc nofollow" target="_blank"><em class="po">Nathan Fisher</em></a><em class="po"><br />A deep dive into the engineering challenges of building a real-time service dependency map at Netflix scale — from streaming architectures and distributed aggregation pipelines to time-travel queries and the methodology that made it work.</em></p><h2 id="4bb4" class="pq pr jp ab ps pt pu pv fm pw px py fo pz qa qb qc qd qe qf qg qh qi qj qk ql cw">Introduction</h2><p id="7446" class="pw-post-body-paragraph ot ou jp ov b ow qm oy oz pa qn pc pd fp qo pf pg fs qp pi pj fv qq pl pm pn il cw">In our <a class="by pp" rel="noopener ugc nofollow" href="https://netflixtechblog.com/from-silos-to-service-topology-why-netflix-built-a-real-time-service-map-0165ba13a7bc" data-discover="true">first post</a>, we introduced the problem: engineers at Netflix needed a unified, real-time view of service dependencies to troubleshoot faster, understand blast radius, and navigate our distributed architecture. We described our multi-source approach — combining eBPF network flows, IPC metrics, and distributed tracing into physically separate graph layers that can be queried independently or merged into a comprehensive view.</p><p id="b44c" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">That post explained <em class="po">what</em> we built and <em class="po">why</em>. This post is about <em class="po">how</em> — the engineering reality of building this system at Netflix scale.</p><p id="9499" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">Here’s the truth: the first version worked perfectly… in our local environment. Production was a different story. Kafka consumers fell behind. Instances ran out of memory. Some nodes received 100x the traffic of others. Garbage collection pauses consumed more CPU than actual business logic.</p><p id="038f" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">What you’ll learn in this post isn’t a success story — it’s a learning journey. We’ll walk through the architecture decisions that enabled scale, the production challenges that tested those decisions, the optimization methodology that guided us through, and the lessons that apply to any distributed system. Along the way, we’ll share the innovations that made it possible to process millions of flow records per second, reconstruct topology at any point in time, and provide sub-second query responses — all while maintaining near real-time freshness.</p><h2 id="da69" class="pq pr jp ab ps pt pu pv fm pw px py fo pz qa qb qc qd qe qf qg qh qi qj qk ql cw">Architecture Deep-Dive: Building for Streaming and Scale</h2><h3 id="ada1" class="qr pr jp ab ps fl qs ap fm fn qt ar fo fp qu fq fr fs qv ft fu fv qw fw fx qx cw">Streaming-First: Why Real-Time Matters</h3><p id="5efd" class="pw-post-body-paragraph ot ou jp ov b ow qm oy oz pa qn pc pd fp qo pf pg fs qp pi pj fv qq pl pm pn il cw">Traditional service topology systems use batch processing — aggregating data hourly or daily, then storing complete snapshots. This approach works at a modest scale but has a fundamental problem: by the time you see the data, it’s already old. During a production incident at 3am, an hour-old dependency map is archaeology, not observability.</p><p id="f5da" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">Our key architectural decision was to build streaming-first. Instead of batch jobs that process historical data, we continuously ingest flow records from multi-region Kafka streams and IPC metrics as Server-Sent Events, process them through reactive pipelines with backpressure handling, and provide near real-time topology updates — typically within tens of minutes, compared to the hours-old or day-old data that batch processing approaches provide.</p><p id="7e9e" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">This wasn’t just about freshness — it was essential for our use cases. Live events can’t wait for the next hourly batch. Incident response needs current data. Change validation requires seeing immediate impact. The architecture had to support continuous updates while handling massive scale without falling behind.</p><p id="c25c" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">How Backpressure Enables Real-Time Processing<br /></strong>The streaming approach created new challenges, but also required solving a fundamental problem: how do you process millions of flow records per second in real-time without losing data when downstream systems slow down?</p><p id="3071" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">Traditional approaches fall short at our scale:</p><ul class=""><li id="d587" class="ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn qy qz ra cw"><strong class="ov jq">Unbounded queues</strong>: Simple but dangerous. Keep buffering until you run out of memory, then the instance crashes.</li><li id="de56" class="ot ou jp ov b ow rb oy oz pa rc pc pd fp rd pf pg fs re pi pj fv rf pl pm pn qy qz ra cw"><strong class="ov jq">Drop-based flow control</strong>: Discard data when buffers fill. Fast, but now your topology is incomplete — you’ve lost connection information.</li><li id="cfc2" class="ot ou jp ov b ow rb oy oz pa rc pc pd fp rd pf pg fs re pi pj fv rf pl pm pn qy qz ra cw"><strong class="ov jq">Batch processing</strong>: Process everything, but hours later. By then, the incident is over (or worse, still happening with stale data).</li></ul><p id="9abe" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">We needed something different: the ability to slow down gracefully under load without losing data. This is where reactive streams with backpressure became essential.</p><p id="45c7" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">Here’s how it works: when Stage 3 can’t write to the graph database fast enough, it signals Stage 2 to slow down. Stage 2 signals Stage 1. Stage 1 signals the Kafka consumer to pause. The data waits in Kafka until downstream capacity returns.</p><figure class="rj rk rl rm rn ro rg rh paragraph-image"><div role="button" tabindex="0" class="rp rq dn rr ct rs">Press enter or click to view image in full size<div class="rg rh ri"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*dgCgcPQv-EvBNb_AXqtnhA.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*dgCgcPQv-EvBNb_AXqtnhA.png" /><img alt="Diagram showing backpressure propagating backward through a pipeline — from Stage 3 to Stage 2 to Stage 1 to the message stream — each stage signaling the previous one to slow dow" class="ct es rt c" width="700" height="60" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="ru bw rv rg rh rw rx ab b cs v x">When a downstream stage can’t keep up, it signals upstream to slow down — backpressure flows in the opposite direction of the data</figcaption></figure><p id="837f" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">Backpressure propagates naturally through the entire system. When any stage becomes overwhelmed — from traffic spikes, GC pauses, or external slowdowns — the pipeline automatically slows to a sustainable rate. No data is lost in most cases, no instances crash, the system degrades gracefully.</p><p id="a2e3" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">This is what enables “real-time” at our scale. During normal operation, we process with minimal latency. During load spikes or temporary slowdowns, we slow down rather than fall over. The data still gets processed — just a few seconds or minutes later instead of immediately. For topology updates, this trade-off is acceptable: slightly delayed real-time updates are vastly better than hour-old batch data or incomplete topology from dropped records.</p><p id="7124" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">The cost of this approach is complexity. Reactive streams are harder to reason about compared to traditional synchronous blocking models (we’ll discuss this more in the challenges section). But at Netflix scale, backpressure isn’t optional — it’s the mechanism that keeps the system running reliably under production load.</p><h3 id="99ba" class="qr pr jp ab ps fl qs ap fm fn qt ar fo fp qu fq fr fs qv ft fu fv qw fw fx qx cw">Multi-Layer Architecture: Physical Separation for Independent Optimization</h3><p id="2f03" class="pw-post-body-paragraph ot ou jp ov b ow qm oy oz pa qn pc pd fp qo pf pg fs qp pi pj fv qq pl pm pn il cw">As we covered in our <a class="by pp" rel="noopener ugc nofollow" href="https://netflixtechblog.com/from-silos-to-service-topology-why-netflix-built-a-real-time-service-map-0165ba13a7bc" data-discover="true">first post</a>, our multi-source approach uses three physically separate topology layers with different storage optimized for each:</p><ul class=""><li id="e3ce" class="ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn qy qz ra cw"><strong class="ov jq">Network Layer</strong>: eBPF flow logs in graph database partition — comprehensive coverage but lacks application context</li><li id="d1e9" class="ot ou jp ov b ow rb oy oz pa rc pc pd fp rd pf pg fs re pi pj fv rf pl pm pn qy qz ra cw"><strong class="ov jq">IPC Layer</strong>: Application metrics in a different graph database isolated from the one for Network Layer — rich endpoint details but only instrumented services</li><li id="c923" class="ot ou jp ov b ow rb oy oz pa rc pc pd fp rd pf pg fs re pi pj fv rf pl pm pn qy qz ra cw"><strong class="ov jq">Tracing Layer</strong>: Distributed traces in columnar storage (Parquet) — actual request paths but sampled.(<em class="po">We cover the tracing layer and its integration in our next post</em><strong class="ov jq"><em class="po">)</em></strong>.</li></ul><figure class="rj rk rl rm rn ro rg rh paragraph-image"><div role="button" tabindex="0" class="rp rq dn rr ct rs">Press enter or click to view image in full size<div class="rg rh ry"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*5_09VEnyPtFR0tbO1zzzMQ.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*5_09VEnyPtFR0tbO1zzzMQ.png" /><img alt="Diagram showing two separate ingestion pipelines — a flow log pipeline and an IPC pipeline, each fed by data enrichment — writing to their own graph store, with a shared API serving UI and backend clients" class="ct es rt c" width="700" height="263" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="ru bw rv rg rh rw rx ab b cs v x">Flow logs and IPC metrics travel through two independently-optimized pipelines into separate graph stores, unified behind a single API</figcaption></figure><p id="4afa" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">Physical storage isolation enables independent optimization — each layer has different throughput, query patterns, and evolution timelines. At query time, we execute parallel queries across relevant storage systems and merge results, providing unified views with sub-second latency while maintaining flexibility to evolve each layer independently.</p><h3 id="32db" class="qr pr jp ab ps fl qs ap fm fn qt ar fo fp qu fq fr fs qv ft fu fv qw fw fx qx cw">The Three-Stage Distributed Aggregation Pipeline</h3><p id="df25" class="pw-post-body-paragraph ot ou jp ov b ow qm oy oz pa qn pc pd fp qo pf pg fs qp pi pj fv qq pl pm pn il cw">The heart of the network layer ingestion is a three-stage distributed pipeline. This architecture solves a fundamental challenge with network flow logs: <strong class="ov jq">they only show individual network hops, not the true application-level connections we need to build a useful topology</strong>.</p><p id="afd0" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">The Core Problem: Network Intermediaries</strong></p><p id="5ca6" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">In cloud environments, traffic between applications rarely flows directly — it traverses intermediate network components like load balancers, NAT gateways, API gateways, and proxies. Network flow logs show individual hops: App A → Load Balancer and Load Balancer → App B appear as separate flows. But what engineers need is the logical dependency: App A → App B. Without resolving these intermediaries, our topology would be cluttered with infrastructure components rather than showing the service-to-service relationships that matter for troubleshooting.</p><p id="47ff" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">The three-stage pipeline solves this:</p><figure class="rj rk rl rm rn ro rg rh paragraph-image"><div role="button" tabindex="0" class="rp rq dn rr ct rs">Press enter or click to view image in full size<div class="rg rh rz"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*LWO54sgNSsNjNda_vG3OhQ.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*LWO54sgNSsNjNda_vG3OhQ.png" /><img alt="Diagram of the flow log pipeline showing a message stream flowing through Stage 1, Stage 2, and Stage 3 via SSE, with data enrichment feeding into Stage 3 before writing to the network graph store" class="ct es rt c" width="700" height="129" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="ru bw rv rg rh rw rx ab b cs v x">The flow log pipeline in detail — three stages connected by SSE, with enrichment applied just before the final graph write</figcaption></figure><p id="8cd5" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Stage 1: Initial Aggregation (FlowLog Ingestion Service)</strong></p><pre class="rj rk rl rm rn sa sb sc ix sd co cw">Multi-Region Kafka (4 regions)<br />  → Filter invalid flow logs<br />  → 5-minute time-window batching<br />  → Create initial aggregators per window<br />  → Distribute via consistent hashing<br />  → Stream to Stage 2 via SSE</pre><p id="a3b6" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">Stage 1 consumes flow logs from multi-region Kafka, filters invalid records, batches them into 5-minute time windows, and creates initial aggregator objects. At this stage, we’re still working with raw network hops — identifying which flows involve intermediaries but not yet resolving them. Aggregators stream to Stage 2 for resolution.</p><p id="d6d0" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Stage 2: Network Intermediary Resolution Layer (Intermediate GraphEntity Ingestion Service)</strong></p><pre class="rj rk rl rm rn sa sb sc ix sd co cw">Stage 1 Aggregators (via SSE streams)<br />  → Group flows by intermediary (load balancer, NAT gateway, proxy, etc.)<br />  → Identify pairs: (Source → Intermediary) + (Intermediary → Destination)<br />  → Resolve to direct edges: Source → Destination<br />  → Track which intermediaries were traversed<br />  → Aggregate metrics across both hops<br />  → Re-distribute via consistent hashing<br />  → Stream to Stage 3 via SSE</pre><p id="a77d" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">This is the key step<strong class="ov jq">.</strong> Stage 2 performs graph resolution:</p><ol class=""><li id="82f1" class="ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn sj qz ra cw"><strong class="ov jq">Collect flows by intermediary</strong>: Group aggregators where an intermediary is either source or destination — creating maps of flows going TO intermediaries (Source → Intermediary) and FROM intermediaries (Intermediary → Destination)</li><li id="1a28" class="ot ou jp ov b ow rb oy oz pa rc pc pd fp rd pf pg fs re pi pj fv rf pl pm pn sj qz ra cw"><strong class="ov jq">Resolve direct edges</strong>: For each intermediary, join its incoming and outgoing flows to create direct application edges (App A → App B), combining metrics from both hops</li><li id="e283" class="ot ou jp ov b ow rb oy oz pa rc pc pd fp rd pf pg fs re pi pj fv rf pl pm pn sj qz ra cw"><strong class="ov jq">Result</strong>: Clean application-level topology showing App A → App B instead of App A → Load Balancer → App B</li></ol><p id="bf74" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">This resolution happens at aggregation time, not query time, with resolved edges flowing to Stage 3.</p><p id="70e3" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Why can’t we do this in a single stage?</strong> The fundamental issue is <strong class="ov jq">data locality</strong>. To resolve App A → Load Balancer → App B into App A → App B, we need both flows on the same instance to perform the join. But in Stage 1, flows are scattered across instances based on Kafka’s partitioning. Stage 2’s critical function is to redistribute aggregators by intermediary identifier — all flows involving “Load Balancer X” route to the same instance for resolution. This is the classic map-reduce pattern: Stage 1 maps, Stage 2 shuffles and reduces by intermediary, Stage 3 performs final aggregation.</p><figure class="rj rk rl rm rn ro rg rh paragraph-image"><div role="button" tabindex="0" class="rp rq dn rr ct rs">Press enter or click to view image in full size<div class="rg rh sk"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*thv_UxYLRCf_IbLJPe_3Sw.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*thv_UxYLRCf_IbLJPe_3Sw.png" /><img alt="Three-panel diagram showing how flow records for services A, B, C, D and load balancers LB1 and LB2 are scattered across instances in Stage 1, reshuffled and resolved into direct edges in Stage 2, and combined and persisted to the graph store in Stage 3." class="ct es rt c" width="700" height="79" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="ru bw rv rg rh rw rx ab b cs v x">A concrete example of why a single stage isn’t enough — Stage 1 scatters flows by partition, Stage 2 reshuffles by intermediary to resolve direct edges, and Stage 3 persists the final result.</figcaption></figure><p id="4e43" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Stage 3: Final Aggregation and Enrichment (GraphEntity Ingestion Service)</strong></p><pre class="rj rk rl rm rn sa sb sc ix sd co cw">Stage 2 Aggregators (via SSE streams)Flow<br />  → Final aggregation across time windows<br />  → Enrich with external data (query key-value stores)<br />  → Convert to graph entities<br />  → Persist to graph database (throttled writes)</pre><p id="867a" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">Stage 3 performs final aggregation of resolved edges, enriches graph nodes with external data sources (application health, ownership, metadata), converts aggregators to concrete graph entities (nodes and edges with all properties populated), and persists them to the distributed graph database with controlled throttling to respect storage system limits.</p><p id="6ef2" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Why Three Stages, Not Two?</strong></p><p id="3055" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">We initially used two stages: aggregate in Stage 1, resolve and persist in Stage 2. This worked in testing but failed at production scale — Stage 2 became overwhelmed by data concentration.</p><p id="b26b" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">The problem: intermediary resolution requires collecting ALL flows involving an intermediary on the same instance.As a result, the instances handling flow logs for popular applications and their intermediaries became ‘hot nodes’ due to significant data concentrationCompounding this, data enrichment (querying external stores for health and metadata) meant the busiest instances were also doing the most I/O.</p><p id="ee70" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">The solution: split responsibilities into three stages. Stage 2 focuses purely on resolution and redistributes. Stage 3 handles enrichment and persistence. This graduated redistribution — distribute, resolve, distribute again, persist — spreads load across multiple instances and isolates compute-heavy resolution from I/O-heavy enrichment. Even when intermediaries see 100x typical traffic, no single instance becomes a bottleneck.</p><p id="9bd7" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Why Server-Sent Events Instead of gRPC or Message Queues?</strong></p><p id="ffba" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">We initially used gRPC but it became a performance bottleneck — serialization overhead, connection pool management, and memory pressure for streaming responses consumed more CPU than business logic. Message queues added infrastructure complexity without benefit for our use case.</p><p id="bc4b" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">SSE proved ideal: lightweight HTTP-based protocol with minimal serialization, natural backpressure integration with reactive streams, and simpler connection model. The lesson: industry best practices like “use gRPC for service communication” don’t apply universally. For streaming large volumes of pre-aggregated data, lighter-weight alternatives may be more appropriate. Measure, don’t assume.</p><p id="1859" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Why IPC Doesn’t Need Three Stages</strong></p><figure class="rj rk rl rm rn ro rg rh paragraph-image"><div role="button" tabindex="0" class="rp rq dn rr ct rs">Press enter or click to view image in full size<div class="rg rh sl"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*vHVMSq80gccWg78809kkDg.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*vHVMSq80gccWg78809kkDg.png" /><img alt="Diagram of the IPC pipeline showing an IPC metrics stream flowing via SSE into a single aggregation stage, with data enrichment feeding into that stage, before writing to the IPC graph store." class="ct es rt c" width="700" height="175" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="ru bw rv rg rh rw rx ab b cs v x">The IPC pipeline mirrors the same pattern as the flow log pipeline, but needs only a single stage.</figcaption></figure><p id="f4a4" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">The IPC layer uses single-stage aggregation because: (1) IPC metrics are already at application level — no intermediaries to resolve, and (2) data is partitioned correctly from the start — each node receives all IPC metrics for its assigned applications via consistent hashing, eliminating the need for redistribution. This highlights a key principle: <strong class="ov jq">data partitioning strategy determines processing architecture</strong>. When data arrives with the right partitioning, you can aggregate directly; when it doesn’t (like network flows requiring intermediary resolution), you need shuffle/redistribution stages.</p><h3 id="76f9" class="qr pr jp ab ps fl qs ap fm fn qt ar fo fp qu fq fr fs qv ft fu fv qw fw fx qx cw">Dynamic Load Distribution: How Hashing Works with Auto-Scaling</h3><p id="8912" class="pw-post-body-paragraph ot ou jp ov b ow qm oy oz pa qn pc pd fp qo pf pg fs qp pi pj fv qq pl pm pn il cw">How do we decide which instance receives which aggregator when our Auto Scaling Groups dynamically add or remove instances? Traditional approaches assume static clusters — requiring explicit rebalancing, coordination services, or manual data movement when cluster size changes.</p><p id="9980" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Our Approach: Dynamic Consistent Hashing</strong></p><p id="6e7c" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">We use consistent hashing with dynamic instance discovery from our service registry. Each instance queries the registry to get the current list of healthy ASG instances, maintains them in sorted order (ensuring all instances have the same view), and uses this list for the hash function <code class="ej sm sn so sb b">findOwnerInstance(aggregator.primaryKey)</code>. When ASG scales up or down, the hash function naturally redistributes aggregators based on the updated instance list — no explicit coordination needed.</p><p id="038e" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">The key insight: leverage existing infrastructure. Our service registry already tracks ASG membership for health checking. Using it as our source of truth gives us dynamic cluster membership for free. Consistent hashing provides stable partitioning (most aggregators stay on the same instance during membership changes), while the sorted list ensures consistency.</p><p id="5007" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">The Result</strong></p><p id="032c" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">Load follows infrastructure automatically. During traffic spikes or live events, new instances immediately receive their share. During deployments, aggregators seamlessly shift to healthy instances. This pattern proved crucial for production stability — no manual intervention, no coordination protocol, just automatic rebalancing.</p><h2 id="18e2" class="pq pr jp ab ps pt pu pv fm pw px py fo pz qa qb qc qd qe qf qg qh qi qj qk ql cw">The V1 Journey: Major Challenges at Production Scale</h2><p id="53b4" class="pw-post-body-paragraph ot ou jp ov b ow qm oy oz pa qn pc pd fp qo pf pg fs qp pi pj fv qq pl pm pn il cw">Getting the initial version (V1) to production taught us that scale changes everything. What works in development breaks in production. Every assumption gets tested. And fixing one bottleneck reveals the next.</p><h3 id="cfd2" class="qr pr jp ab ps fl qs ap fm fn qt ar fo fp qu fq fr fs qv ft fu fv qw fw fx qx cw">Challenge 1: Kafka Consumer Lag</h3><p id="551b" class="pw-post-body-paragraph ot ou jp ov b ow qm oy oz pa qn pc pd fp qo pf pg fs qp pi pj fv qq pl pm pn il cw"><strong class="ov jq">The Problem</strong>: Our multi-region Kafka consumers started falling behind — consumer lag grew from seconds to minutes, then hours. Flow logs were arriving faster than we could process them. If this continued, we’d never catch up, and our “real-time” topology would become increasingly stale.</p><p id="7415" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Investigation</strong>: We instrumented Kafka consumer metrics heavily. Key findings:</p><ul class=""><li id="42ec" class="ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn qy qz ra cw">Kafka had fewer partitions than optimal for our consumer group size</li><li id="573d" class="ot ou jp ov b ow rb oy oz pa rc pc pd fp rd pf pg fs re pi pj fv rf pl pm pn qy qz ra cw">Each fetch operation retrieved relatively few records</li><li id="1580" class="ot ou jp ov b ow rb oy oz pa rc pc pd fp rd pf pg fs re pi pj fv rf pl pm pn qy qz ra cw">Network socket buffers weren’t right-sized for our throughput</li><li id="96e6" class="ot ou jp ov b ow rb oy oz pa rc pc pd fp rd pf pg fs re pi pj fv rf pl pm pn qy qz ra cw">Cross-region read latency added overhead</li></ul><p id="4576" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Solutions Applied</strong>:</p><ol class=""><li id="07f2" class="ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn sj qz ra cw"><strong class="ov jq">Increased Kafka partitions</strong>: More partitions enabled more parallel consumers in our consumer group, distributing load across more instances.</li><li id="4544" class="ot ou jp ov b ow rb oy oz pa rc pc pd fp rd pf pg fs re pi pj fv rf pl pm pn sj qz ra cw"><strong class="ov jq">Tuned fetch parameters</strong>: Increased records per fetch operation, reducing the number of network round-trips. This trades off per-message latency (we fetch larger batches) for throughput (more records processed per second).</li><li id="0ebd" class="ot ou jp ov b ow rb oy oz pa rc pc pd fp rd pf pg fs re pi pj fv rf pl pm pn sj qz ra cw"><strong class="ov jq">Increased socket receive buffer size</strong>: Ensured network buffers never limited fetch operations. At our scale, default buffer sizes were too small.</li></ol><p id="5b4e" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Results</strong>: Throughput improved significantly, and lag reduced to acceptable levels — typically under a minute even during peak traffic.</p><p id="db35" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Lesson</strong>: At scale, you can’t optimize in isolation. Fixing Kafka lag revealed the next bottleneck: our instances themselves couldn’t keep up with the higher ingest rate. The pipeline moved faster, which exposed downstream capacity problems.</p><h3 id="5ed1" class="qr pr jp ab ps fl qs ap fm fn qt ar fo fp qu fq fr fs qv ft fu fv qw fw fx qx cw">Challenge 2: Hot Nodes and Data Amplification</h3><p id="8f7b" class="pw-post-body-paragraph ot ou jp ov b ow qm oy oz pa qn pc pd fp qo pf pg fs qp pi pj fv qq pl pm pn il cw"><strong class="ov jq">The Problem</strong>: This was the most severe production issue we faced. Some instances in our Auto Scaling Group were receiving 100x more traffic than others. Memory usage spiked. Garbage collection pauses became frequent and long. More CPU time was spent in GC than in business logic. Eventually, hot instances would go DOWN, triggering cascading failures as their load redistributed to other instances.</p><p id="2cda" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Root Cause Investigation</strong>:<br />Flow logs for popular services dominate traffic volume. A service like our authentication layer or recommendation API is called by hundreds of other services, generating orders of magnitude more flow records than typical services.</p><p id="939f" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">Our initial architecture used consistent hashing to determine which instance owned aggregation for each destination service. All flow logs for a given destination are routed to the same instance — the “owner” for that destination. This design seemed reasonable: group related data for efficient aggregation.</p><p id="d597" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">But popular destinations created hot nodes. One instance might own authentication services, another might own a rarely-used backend service. The load distribution was wildly uneven — some instances handled 100x the flow records of others.</p><p id="bd39" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">Worse, data amplification occurred during redistribution. Consider a service called by 100 upstream services across 10 instances. All 10 instances receive flow logs for that destination (because they all have local clients calling it). When they route aggregators to the owner instance, that instance receives 10 separate aggregators it must merge. The data volume multiplied during shuffling.</p><figure class="rj rk rl rm rn ro rg rh paragraph-image"><div class="rg rh sp"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*MmWlcA1C2o-z_Zui1wCarA.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*MmWlcA1C2o-z_Zui1wCarA.png" /><img alt="Diagram showing many instances each sending aggregators for the same destination into a single owner instance, illustrating how data volume multiplies at the point of convergence" class="ct es rt c" width="594" height="590" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div><figcaption class="ru bw rv rg rh rw rx ab b cs v x">When many instances route data for the same key to one owner, the volume multiplies right where it lands — the root cause of hot nodes.</figcaption></figure><p id="24be" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">We profiled extensively using async-profiler and heap dump analysis. The results were clear: hot instances spent most of their CPU on garbage collection, trying to manage the rapid allocation and deallocation of aggregator objects as flow logs poured in faster than they could be processed. Memory pressure led to GC thrashing, which consumed CPU, which slowed processing, which increased memory pressure — a vicious cycle.</p><p id="707a" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Solution: The Three-Stage Pipeline’s Dual Benefits<br /></strong>The three-stage pipeline we described earlier — designed primarily for proxy resolution — turned out to be exactly what we needed to solve the hot nodes problem as well. Here’s why:</p><p id="d78c" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Stage 1</strong> performs initial aggregation locally before any distribution. Instead of sending every flow log to a remote instance immediately. Each instance performs online aggregation of raw flow logs into time-windowed aggregators (over 5-minute periods) directly in memory; this allows the raw flow to be discarded and garbage collected quickly,, significantly reducing memory pressure, and ensures only the aggregation results are transferred across the network to downstream stages.</p><div class="sq eu sr ss an k gi"><div class="st su sv sw sx e bw"><h2 class="sy b sz ta tb tc td">Get Netflix Technology Blog’s stories in your inbox</h2></div><div class="te e bw"><p class="ab b cs v x">Join Medium for free to get updates from this writer.</p></div><div class="ti tj tk tl tm tn an"><div class="to an gi"><div class="an to bs br ix tp tq ej gp tr ts tt tu tv tw tx"><input class="ty cb bz cc tz ua dy nc cx hq ct" placeholder="Enter your email" type="text" value="" /></div></div><div class="cl da ub ek"><div class="tf tg th"><button class="ab b cs v uc ud ue uf ug uh ui bl bm uj uk ul bq br bs bt bu bv bw">Subscribe</button></div></div><div class="ct db s t"><button class="ab b cs v uc ud ue uf ug uh ui bl bm uj uk ul bq ct br bs bt bu bv bw">Subscribe</button></div></div><div class="ik e"><label class="an k"><input class="um un uo ds he up uq ur us ut uu uv" type="checkbox" checked="checked" /></label></div><div class="e"><p class="ab b cs v cw">Remember me for faster sign in</p></div></div><div class="vd e"></div><p id="e7cc" class="pw-post-body-paragraph ot ou jp ov b ow oy oz pa pc pd fp pf pg fs pi pj fv pl pm ss pn il cw"><strong class="ov jq">Stage 2</strong> focuses on proxy resolution but also provides intermediate redistribution. Aggregators from Stage 1 distribute via consistent hashing to Stage 2 instances. Now we’re moving compressed aggregators, not individual flow logs. After resolution, Stage 2 redistributes resolved edges again to Stage 3, providing a second hashing operation that further spreads load.</p><p id="291d" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Stage 3</strong> receives resolved aggregators that have been compressed twice and distributed twice. Even for extremely popular services, load has been spread across enough distribution points that no single instance becomes overwhelmed.</p><p id="1fa5" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">The key insight: architectural decisions driven by one requirement (proxy resolution) often solve other problems (load distribution) as beneficial side effects. The three-stage pipeline with graduated redistribution achieves both goals — it resolves proxies to show clean application-level topology AND prevents hot nodes by spreading load across multiple distribution points.</p><p id="221a" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Switching from gRPC to SSE<br /></strong>As described earlier, this challenge also revealed that gRPC wasn’t the right protocol for inter-stage communication at our scale. We replaced gRPC with Server-Sent Events, dramatically reducing resource consumption on both sender and receiver sides.</p><p id="18e0" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Results</strong>:</p><ul class=""><li id="0818" class="ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn qy qz ra cw">CPU usage became evenly distributed across instances — no more hot nodes with 10x the load of others</li><li id="4165" class="ot ou jp ov b ow rb oy oz pa rc pc pd fp rd pf pg fs re pi pj fv rf pl pm pn qy qz ra cw">Network bandwidth usage dropped significantly due to better aggregation and lighter-weight protocol</li><li id="abc0" class="ot ou jp ov b ow rb oy oz pa rc pc pd fp rd pf pg fs re pi pj fv rf pl pm pn qy qz ra cw">Memory pressure decreased as we reduced the object allocation rate</li><li id="aacc" class="ot ou jp ov b ow rb oy oz pa rc pc pd fp rd pf pg fs re pi pj fv rf pl pm pn qy qz ra cw">The system scaled gracefully with Auto Scaling Group changes</li></ul><p id="11f6" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Lesson</strong>: Technology choices must match your specific use case. gRPC is excellent for request-response RPC patterns. For streaming large volumes of aggregated data in a pipeline, lighter-weight alternatives can be more appropriate. Let measurements guide the decision, not industry hype or existing team expertise.</p><h3 id="409b" class="qr pr jp ab ps fl qs ap fm fn qt ar fo fp qu fq fr fs qv ft fu fv qw fw fx qx cw">Challenge 3: Memory and Garbage Collection</h3><p id="75bc" class="pw-post-body-paragraph ot ou jp ov b ow qm oy oz pa qn pc pd fp qo pf pg fs qp pi pj fv qq pl pm pn il cw"><strong class="ov jq">The Problem</strong>: Even after fixing hot nodes, we still saw high heap usage, frequent garbage collection pauses, and instances occasionally going DOWN. GC logs showed pauses consuming significant CPU time — in some cases, more than our business logic.</p><p id="3f68" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Root Cause</strong>: Multiple factors contributed: objects accumulating in heap while waiting for 5-minute aggregation windows to complete, unnecessary conversions between different object types as data flowed through stages, and immutability overhead — following Scala best practices, we used immutable data structures for aggregators, but every update created new objects, overwhelming the garbage collector at millions of records per second.</p><p id="520d" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Investigation</strong>: Heap dumps and GC logs revealed flow log objects retained beyond their useful lifetime, unnecessary intermediate conversion objects, and constant creation/disposal of immutable aggregator versions. Minor GCs occurred every few seconds, major GCs took hundreds of milliseconds — the JVM spent more time on garbage collection than business logic.</p><p id="5080" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Solutions Applied</strong>:</p><ol class=""><li id="6242" class="ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn sj qz ra cw"><strong class="ov jq">Faster processing</strong>: Process flow logs immediately, aggregate quickly, release references. Optimized Pekko stream stages to minimize object lifetime.</li><li id="6b70" class="ot ou jp ov b ow rb oy oz pa rc pc pd fp rd pf pg fs re pi pj fv rf pl pm pn sj qz ra cw"><strong class="ov jq">Eliminate unnecessary conversions</strong>: Route aggregators directly between stages instead of converting to intermediate types.</li><li id="93bc" class="ot ou jp ov b ow rb oy oz pa rc pc pd fp rd pf pg fs re pi pj fv rf pl pm pn sj qz ra cw"><strong class="ov jq">Mutable structures on hotpath</strong>: This was controversial — Scala best practices emphasize immutability. But at our scale, immutability created too many objects. We pragmatically chose mutable aggregators on the hotpath (immutability elsewhere), prioritizing performance over convention. Switching to mutable aggregators reduced heap allocation by over 50% and cut GC pause time significantly, though it required more careful code review.</li><li id="d3e3" class="ot ou jp ov b ow rb oy oz pa rc pc pd fp rd pf pg fs re pi pj fv rf pl pm pn sj qz ra cw"><strong class="ov jq">Tuned time windows</strong>: Balanced data freshness against memory pressure.</li></ol><p id="ebab" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Results</strong>:</p><ul class=""><li id="f6ba" class="ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn qy qz ra cw">Heap usage decreased substantially</li><li id="9116" class="ot ou jp ov b ow rb oy oz pa rc pc pd fp rd pf pg fs re pi pj fv rf pl pm pn qy qz ra cw">GC pauses reduced to acceptable levels (tens of milliseconds instead of hundreds)</li><li id="e502" class="ot ou jp ov b ow rb oy oz pa rc pc pd fp rd pf pg fs re pi pj fv rf pl pm pn qy qz ra cw">CPU freed up for business logic instead of garbage collection</li><li id="db52" class="ot ou jp ov b ow rb oy oz pa rc pc pd fp rd pf pg fs re pi pj fv rf pl pm pn qy qz ra cw">Instance stability improved — no more instances going DOWN due to memory issues</li></ul><p id="f842" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Lesson</strong>: “Best practices” are starting points, not absolute rules. At unique scale, you may need to diverge from conventions. But do it deliberately, with measurement justifying the decision, and with awareness of the trade-offs. Don’t abandon immutability everywhere — just where performance data proves it’s necessary.</p><h2 id="1b07" class="pq pr jp ab ps pt pu pv fm pw px py fo pz qa qb qc qd qe qf qg qh qi qj qk ql cw">Challenge 4: Reactive Streams Complexity</h2><p id="44b5" class="pw-post-body-paragraph ot ou jp ov b ow qm oy oz pa qn pc pd fp qo pf pg fs qp pi pj fv qq pl pm pn il cw"><strong class="ov jq">The Problem</strong>: Our Pekko Streams pipelines would stall unexpectedly. Backpressure propagation didn’t work as expected. We struggled to debug why certain streams would stop processing without obvious errors. The reactive programming mental model — with its emphasis on async boundaries, backpressure, and demand-driven processing — proved harder to master than anticipated.</p><p id="aaac" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">What We Learned</strong>:<br />Reactive streams with backpressure are powerful tools for building systems that handle load spikes gracefully. When downstream consumers slow down (due to temporary load, GC pauses, or external system slowdowns), backpressure allows upstream producers to slow down rather than overflow buffers or drop data.</p><p id="8b9c" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">But this power comes with complexity:</p><ul class=""><li id="7190" class="ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn qy qz ra cw"><strong class="ov jq">Non-intuitive behavior</strong>: Traditional imperative code flows top-to-bottom. Reactive streams are demand-driven — downstream consumers pull from upstream producers. This inversion of control isn’t intuitive.</li><li id="0e4d" class="ot ou jp ov b ow rb oy oz pa rc pc pd fp rd pf pg fs re pi pj fv rf pl pm pn qy qz ra cw"><strong class="ov jq">Async boundaries</strong>: The .async operator in Pekko Streams creates a boundary where processing moves to a different thread. This can improve parallelism but also introduces complexity around buffer sizing, demand signaling, and error propagation. We initially misunderstood when to use .async and ended up with over-parallelized streams that created more overhead than benefit.</li><li id="a0bf" class="ot ou jp ov b ow rb oy oz pa rc pc pd fp rd pf pg fs re pi pj fv rf pl pm pn qy qz ra cw"><strong class="ov jq">Debugging difficulty</strong>: When a stream stalls, there’s no stack trace pointing to the problem. You must understand the internal mechanics — demand signals, buffer states, materializer state — to diagnose issues.</li></ul><p id="b49a" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Our Approach</strong>:</p><ol class=""><li id="8213" class="ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn sj qz ra cw"><strong class="ov jq">Deep learning investment</strong>: We invested significant time in understanding reactive streams concepts deeply. Reading documentation, experimenting with small examples, and building team expertise.</li><li id="acf5" class="ot ou jp ov b ow rb oy oz pa rc pc pd fp rd pf pg fs re pi pj fv rf pl pm pn sj qz ra cw"><strong class="ov jq">Simplified patterns</strong>: Where possible, we simplified our stream graphs. Complex branching and merging patterns are powerful but hard to debug. We preferred linear flows with clear stage boundaries.</li><li id="1e83" class="ot ou jp ov b ow rb oy oz pa rc pc pd fp rd pf pg fs re pi pj fv rf pl pm pn sj qz ra cw"><strong class="ov jq">Better monitoring</strong>: We added metrics at stream boundaries — tracking buffer sizes, element throughput, backpressure events. Visibility into stream internals helped diagnose issues.</li><li id="be8c" class="ot ou jp ov b ow rb oy oz pa rc pc pd fp rd pf pg fs re pi pj fv rf pl pm pn sj qz ra cw"><strong class="ov jq">Team education</strong>: We documented our learnings, shared patterns that worked, and built institutional knowledge about reactive streams.</li></ol><p id="2996" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Lesson</strong>: Powerful abstractions require investment. Don’t assume you understand a framework without validation. Build your mental model deliberately, test it with experiments, and be humble about your understanding. Reactive streams are worth mastering for systems that need to handle load gracefully, but expect a learning curve.</p><h2 id="468d" class="pq pr jp ab ps pt pu pv fm pw px py fo pz qa qb qc qd qe qf qg qh qi qj qk ql cw">V2 Evolution: Continuous Refinement</h2><p id="3056" class="pw-post-body-paragraph ot ou jp ov b ow qm oy oz pa qn pc pd fp qo pf pg fs qp pi pj fv qq pl pm pn il cw">V1 got us to production. The major architectural challenges — Kafka lag, hot nodes, memory pressure — were solved. But production at full scale revealed new optimization opportunities. V2 represents the continuous refinement that turns a working system into a production-ready system.</p><h3 id="3f16" class="qr pr jp ab ps fl qs ap fm fn qt ar fo fp qu fq fr fs qv ft fu fv qw fw fx qx cw">Challenge 5: Persistent Heap Pressure</h3><p id="6f7a" class="pw-post-body-paragraph ot ou jp ov b ow qm oy oz pa qn pc pd fp qo pf pg fs qp pi pj fv qq pl pm pn il cw"><strong class="ov jq">The Problem</strong>: Despite V1 optimizations, we still observed higher-than-desired heap usage. GC metrics improved but weren’t optimal. Memory profiling showed room for improvement.</p><p id="238c" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Root Cause</strong>: Deeper analysis revealed we were still doing unnecessary object conversions between stages. We’d convert aggregators to full graph entities (with all properties populated) before routing to the next stage, even though the next stage just needed the compressed aggregator state.</p><p id="a4fe" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Solution</strong>: Architectural change to route aggregators directly through all stages, only converting to final graph entities at Stage 3 immediately before persistence. This eliminated two intermediate conversion steps and the associated object allocation.</p><p id="eaa6" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Result</strong>: Heap usage dropped further, GC pauses became even less frequent, and memory headroom improved.</p><h3 id="751e" class="qr pr jp ab ps fl qs ap fm fn qt ar fo fp qu fq fr fs qv ft fu fv qw fw fx qx cw">Challenge 6: Serialization Complexity</h3><p id="9a22" class="pw-post-body-paragraph ot ou jp ov b ow qm oy oz pa qn pc pd fp qo pf pg fs qp pi pj fv qq pl pm pn il cw"><strong class="ov jq">The Problem</strong>: Custom serialization logic for SSE messages caused occasional erratic errors that were hard to reproduce and debug. Different parts of the codebase used inconsistent serialization approaches.</p><p id="8e3c" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Solution</strong>: Standardized on JSON encoding throughout the pipeline. While slightly less efficient than binary serialization, JSON’s human readability made debugging far easier, and the overhead was negligible compared to other operations. Consistency eliminated an entire class of bugs.</p><p id="69bc" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Result</strong>: Serialization-related errors disappeared. Debugging became easier because we could read SSE message contents directly.</p><h3 id="2c4d" class="qr pr jp ab ps fl qs ap fm fn qt ar fo fp qu fq fr fs qv ft fu fv qw fw fx qx cw">Challenge 7: Stream Processing Inefficiencies</h3><p id="dee5" class="pw-post-body-paragraph ot ou jp ov b ow qm oy oz pa qn pc pd fp qo pf pg fs qp pi pj fv qq pl pm pn il cw"><strong class="ov jq">The Problem</strong>: Even after understanding reactive streams better, our Pekko configurations weren’t optimal. We had over-parallelized some stages and under-parallelized others. The .async boundaries weren’t placed optimally.</p><p id="83a8" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Solution</strong>: Through continued profiling and experimentation, we tuned parallelism parameters, adjusted buffer sizes, and refined async boundary placement. We added monitoring at stream boundaries to identify bottlenecks.</p><p id="39ed" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Result</strong>: Throughput improvements and more consistent processing latency.</p><h3 id="6404" class="qr pr jp ab ps fl qs ap fm fn qt ar fo fp qu fq fr fs qv ft fu fv qw fw fx qx cw">Challenge 8: Uneven Graph Database Throughput</h3><p id="d2a1" class="pw-post-body-paragraph ot ou jp ov b ow qm oy oz pa qn pc pd fp qo pf pg fs qp pi pj fv qq pl pm pn il cw"><strong class="ov jq">The Problem</strong>: Write distribution to our graph database wasn’t even. Some partitions received heavy write traffic while others sat idle. This caused throttling to kick in unevenly and limited overall write throughput.</p><p id="1730" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Solution</strong>: Implemented batching of aggregators before writing to the graph database and improved distribution logic across partitions. Rather than writing each aggregator immediately, we batch them and write multiple entities in coordinated operations.</p><p id="443d" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Result</strong>: More consistent write throughput and better utilization of database capacity.</p><h3 id="f41c" class="qr pr jp ab ps fl qs ap fm fn qt ar fo fp qu fq fr fs qv ft fu fv qw fw fx qx cw">Challenge 9: Data Enrichment at Aggregation Time</h3><p id="9b68" class="pw-post-body-paragraph ot ou jp ov b ow qm oy oz pa qn pc pd fp qo pf pg fs qp pi pj fv qq pl pm pn il cw">Beyond the core topology graph, we needed to enrich nodes with additional context. At Stage 3, before persisting graph entities, we integrate enrichment data from external sources — application health status, ownership information, and other metadata. Performing this enrichment at aggregation time rather than at query time avoids the performance overhead of post-query joins and ensures every topology node has full context when queried.</p><h3 id="d053" class="qr pr jp ab ps fl qs ap fm fn qt ar fo fp qu fq fr fs qv ft fu fv qw fw fx qx cw">Pattern Recognition</h3><p id="2bcf" class="pw-post-body-paragraph ot ou jp ov b ow qm oy oz pa qn pc pd fp qo pf pg fs qp pi pj fv qq pl pm pn il cw">Each V2 challenge followed the same pattern: production revealed an assumption, profiling identified the root cause, targeted fixes improved specific metrics. Measure, hypothesize, validate, iterate. This is how you build at scale — not by getting everything right upfront, but by continuous learning and improvement.</p><h2 id="d6d8" class="pq pr jp ab ps pt pu pv fm pw px py fo pz qa qb qc qd qe qf qg qh qi qj qk ql cw">Time Travel: Continuous Topology Reconstruction</h2><p id="9737" class="pw-post-body-paragraph ot ou jp ov b ow qm oy oz pa qn pc pd fp qo pf pg fs qp pi pj fv qq pl pm pn il cw">One of the most powerful capabilities we built enables querying historical topology: “What did the call graph look like when this incident happened?” This time-travel feature required solving an interesting architectural challenge — how to efficiently store and reconstruct topology across time.</p><h3 id="0ccd" class="qr pr jp ab ps fl qs ap fm fn qt ar fo fp qu fq fr fs qv ft fu fv qw fw fx qx cw">The Problem</h3><p id="b300" class="pw-post-body-paragraph ot ou jp ov b ow qm oy oz pa qn pc pd fp qo pf pg fs qp pi pj fv qq pl pm pn il cw">Engineers need to answer temporal questions: What did the topology look like during an incident? How have dependencies evolved? Traditional approaches — full snapshots or event sourcing — either have exponential storage costs or require slow log replay.</p><h3 id="c103" class="qr pr jp ab ps fl qs ap fm fn qt ar fo fp qu fq fr fs qv ft fu fv qw fw fx qx cw">Our Approach: Time-Windowed Aggregators with Mutation Tracking</h3><p id="d686" class="pw-post-body-paragraph ot ou jp ov b ow qm oy oz pa qn pc pd fp qo pf pg fs qp pi pj fv qq pl pm pn il cw">We combine two mechanisms:</p><p id="44d4" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">1. Time-Windowed Aggregator Snapshots</strong>: Every aggregator stores startTs and endTs timestamps for its 5-minute window. These immutable aggregators persist in the graph database keyed by (entity_id, timestamp), providing checkpoint states every 5 minutes.</p><p id="1b9b" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">2. Property-Level Mutation Tracking</strong>: The graph database maintains mutation history at the property level — storing only changed properties with timestamps. This is much more efficient than full entity copies and provides sub-window precision beyond the 5-minute aggregation boundaries.</p><p id="3d5b" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">3. Query-Time Reconstruction</strong>: When querying historical topology, we query the mutation history API for the time range, retrieve all mutations, and reconstruct topology state by applying mutations in order.</p><p id="d3dc" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">This approach provides efficient storage (compressed aggregator states + sparse property mutations), fast retrieval (indexed mutation history, no log replay), and flexible analysis (arbitrary time ranges without pre-computing all possibilities).</p><p id="2f55" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><strong class="ov jq">Query-Time Re-Aggregation</strong>: We can further aggregate historical data at query time using the same aggregator classes from ingestion. This enables arbitrary groupby dimensions (availability tier, business domain, deployment cluster) that weren’t pre-computed, allowing exploratory analysis without exploding storage costs.</p><h2 id="c9fe" class="pq pr jp ab ps pt pu pv fm pw px py fo pz qa qb qc qd qe qf qg qh qi qj qk ql cw">Lessons for Distributed Systems</h2><p id="3b2f" class="pw-post-body-paragraph ot ou jp ov b ow qm oy oz pa qn pc pd fp qo pf pg fs qp pi pj fv qq pl pm pn il cw">While these challenges were specific to service topology, the lessons apply broadly to distributed systems at scale.</p><h3 id="2260" class="qr pr jp ab ps fl qs ap fm fn qt ar fo fp qu fq fr fs qv ft fu fv qw fw fx qx cw">Scale Changes Everything</h3><p id="5f44" class="pw-post-body-paragraph ot ou jp ov b ow qm oy oz pa qn pc pd fp qo pf pg fs qp pi pj fv qq pl pm pn il cw">What works at 100 requests per second fails at 100,000 requests per second. The change isn’t linear — it’s qualitative. Approaches that are fine at modest scale hit fundamental walls at extreme scale.</p><p id="c925" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">Examples from our journey: immutable data structures create GC pressure at millions of allocations per second; single-stage aggregation fails catastrophically with power-law traffic distribution; standard gRPC becomes heavyweight for streaming aggregation at volume.</p><p id="2ee9" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">The lesson: be willing to break conventional wisdom when scale justifies it. But do it based on measurement, not speculation.</p><h3 id="4b3a" class="qr pr jp ab ps fl qs ap fm fn qt ar fo fp qu fq fr fs qv ft fu fv qw fw fx qx cw">Optimize One Bottleneck at a Time</h3><p id="b20c" class="pw-post-body-paragraph ot ou jp ov b ow qm oy oz pa qn pc pd fp qo pf pg fs qp pi pj fv qq pl pm pn il cw">Distributed systems have cascading bottlenecks. Fix Kafka lag, and you discover hot node issues. Fix hot nodes, and you discover GC problems. Fix GC, and you discover serialization inefficiencies.</p><p id="03b9" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">This isn’t failure — it’s the nature of complex systems. Each optimization raises throughput, which stresses the next weakest point. The approach: prioritize based on impact, fix the current bottleneck thoroughly with measurement confirming resolution, then move to the next one. Optimization at scale is continuous, not one-time.</p><h3 id="055b" class="qr pr jp ab ps fl qs ap fm fn qt ar fo fp qu fq fr fs qv ft fu fv qw fw fx qx cw">Distribution Is Key to Scale</h3><p id="ad28" class="pw-post-body-paragraph ot ou jp ov b ow qm oy oz pa qn pc pd fp qo pf pg fs qp pi pj fv qq pl pm pn il cw">Single aggregation points are inevitable bottlenecks. Consistent hashing distributes load but doesn’t prevent concentration when data itself is unevenly distributed (power-law distributions like ours).</p><p id="35a6" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">Our three-stage pipeline with graduated redistribution solved this. Load spreads across multiple distribution points at each stage. Even with highly skewed data, no single instance becomes overwhelmed. The general principle: use multi-stage processing with redistribution at each stage when dealing with skewed data at scale.</p><h2 id="a8d3" class="pq pr jp ab ps pt pu pv fm pw px py fo pz qa qb qc qd qe qf qg qh qi qj qk ql cw">Current State and Impact</h2><p id="2dab" class="pw-post-body-paragraph ot ou jp ov b ow qm oy oz pa qn pc pd fp qo pf pg fs qp pi pj fv qq pl pm pn il cw">Service Topology operates in production today, processing flow logs, ipc metrics and traces from multiple regions and serving queries with sub-second latency. Teams across Netflix use it daily for incident investigation, blast radius analysis, dependency understanding, and production change management. The system has become essential infrastructure for maintaining reliability at scale.</p><h2 id="d7d2" class="pq pr jp ab ps pt pu pv fm pw px py fo pz qa qb qc qd qe qf qg qh qi qj qk ql cw">Conclusion</h2><p id="d5df" class="pw-post-body-paragraph ot ou jp ov b ow qm oy oz pa qn pc pd fp qo pf pg fs qp pi pj fv qq pl pm pn il cw">Service Topology at Netflix represents a journey through building distributed systems at scale. We started with engineers struggling to understand dependencies across scattered tools. We built a multi-layer architecture using streaming aggregation, network intermediary resolution, and time-travel capabilities. And we learned that optimization at scale is continuous — measure, iterate, validate, repeat.</p><p id="ebf1" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">The challenges we faced — Kafka lag, hot nodes, memory pressure — required breaking conventional wisdom when data justified it. Each fix revealed the next bottleneck. But that iterative process, guided by constant measurement, is what makes systems work at extreme scale.</p><p id="4908" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw">In our next post, we’ll explore the tracing layer integration, unified querying across heterogeneous storage, and how all three layers combine to provide comprehensive topology visibility.</p><h2 id="3caa" class="pq pr jp ab ps pt pu pv fm pw px py fo pz qa qb qc qd qe qf qg qh qi qj qk ql cw">Acknowledgements</h2><p id="1aa4" class="pw-post-body-paragraph ot ou jp ov b ow qm oy oz pa qn pc pd fp qo pf pg fs qp pi pj fv qq pl pm pn il cw"><em class="po">Service Topology was built by </em><a class="by pp" href="https://www.linkedin.com/in/parth-jain-8a09abb6/" rel="noopener ugc nofollow" target="_blank"><em class="po">Parth Jain</em></a><em class="po">, </em><a class="by pp" href="https://www.linkedin.com/in/raskuma/" rel="noopener ugc nofollow" target="_blank"><em class="po">Rakesh Sukumar</em></a><em class="po">, </em><a class="by pp" href="https://www.linkedin.com/in/yingwu-zhao-62037418/" rel="noopener ugc nofollow" target="_blank"><em class="po">Yingwu Zhao</em></a><em class="po">, </em><a class="by pp" href="https://www.linkedin.com/in/renzosanchezsilva/" rel="noopener ugc nofollow" target="_blank"><em class="po">Renzo Sanchez-Silva</em></a><em class="po">, and </em><a class="by pp" href="https://www.linkedin.com/in/nathfisher/" rel="noopener ugc nofollow" target="_blank"><em class="po">Nathan Fisher</em></a><em class="po">.</em></p><p id="3f4f" class="pw-post-body-paragraph ot ou jp ov b ow ox oy oz pa pb pc pd fp pe pf pg fs ph pi pj fv pk pl pm pn il cw"><em class="po">Special thanks to the many engineers across Netflix who made this possible — the Observability team who built the broader system, the graph database platform team who provided the storage foundation, and the Platform Modernization Engineering, and Live teams who provided invaluable feedback and use cases throughout development.</em></p></div></div></div></div></div></section></div></div></article><article aria-label="GenPage: Towards End-to-End Generative Homepage Construction at Netflix" class="am" data-testid="post-preview"><div class="am vn e"><div class="ct"><div class="am e"><div class="dn am zk zl zm zn zo zp zq zr zs zt zu zv zw"><div class="zx"><div aria-label="GenPage: Towards End-to-End Generative Homepage Construction at Netflix"><div class="zz aba abb abc el"><img alt="GenPage: Towards End-to-End Generative Homepage Construction at Netflix" class="ct abd abe abf abg" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*NurrizMgC7_QbsGuuW42Dg.gif" /></div></div></div><div class="zy an ex gi"><div class="an gi uy ct abh abi abj abk"><div class="abl abm abn abo abp fe an k"><div class="fe an k abq"><div class="abr e ek"><div><div class="e"><div tabindex="-1" class="cr"><a href="https://netflixtechblog.com" rel="noopener follow"><div class="dn"><img alt="Netflix TechBlog" class="ej ig e abt abs" src="https://miro.medium.com/v2/resize:fill:40:40/1*ty4NvNrGg4ReETxqU2N3Og.png" width="20" height="20" /></div></a></div></div></div></div><div class="wv e ek de"><p class="ab b z v x">In</p></div><div class="fe e"><div><div class="e"><div tabindex="-1" class="cr"><a class="by y bz ca cb ac cc ae af ag ah ai aj lh an k" href="https://netflixtechblog.com" rel="noopener follow"><p class="ab b z v dw hr et hs ht hu hv hw cw">Netflix TechBlog</p></a></div></div></div></div></div><div class="abu e ek de"><p class="ab b z v x">by</p></div><div class="fe an k abq"><div class="fe e"><div><div class="e"><div tabindex="-1" class="cr"><a class="by y bz ca cb ac cc ae af ag ah ai aj lh an k" href="https://netflixtechblog.medium.com" rel="noopener follow"><p class="ab b z v dw hr et hs ht hu hv hw cw">Netflix Technology Blog</p></a></div></div></div></div></div><div class="nu an k ek dc de"><p class="ab b z v x">·</p><p class="ab b z v x">Jun 29</p></div></div><div class="abv e abw abx aby abz aca il"><div class="acb acc acd ace acf acg ach aci"><a class="by y bz ca cb ac cc ae af ag ah ai aj ak al" rel="noopener follow" href="https://netflixtechblog.com/genpage-towards-end-to-end-generative-homepage-construction-at-netflix-77146fba8a08" data-discover="true"><div title="GenPage: Towards End-to-End Generative Homepage Construction at Netflix"><h2 class="ab jq pt pv acj ack acl acm acn aco fm pw py acp acq acr acs act acu fo fp fq acv acw acx acy acz ada fr fs ft adb adc add ade adf adg fu fv fw adh adi adj adk adl adm fx hw cw">GenPage: Towards End-to-End Generative Homepage Construction at Netflix</h2></div><div class="adn e"><h3 class="ab b ga v dw ado et hs adp hu hw x">Authors: Lequn Wang, Jiangwei Pan, and Linas Baltrunas</h3></div></a></div></div><div class="wm an km cf"><div class="an k kx"><div class="cl da db s t"><div class="dn adq adr an k"><div class="ds gp ads an k kx"><a class="ds he ads an k kx" tabindex="-1" rel="noopener follow" href="https://netflixtechblog.com/genpage-towards-end-to-end-generative-homepage-construction-at-netflix-77146fba8a08" data-discover="true"><div><div class="an"><div tabindex="-1" class="cr"><div class="an k dc">A response icon10</div></div></div></div></a></div></div></div><div class="an k adu adv"><a class="by y bz ca cb ac cc ae af ag ah ai aj ak al" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fbuilding-service-topology-at-scale-architecture-challenges-and-lessons-learned-f4b792f3f0d8" rel="noopener follow"><div><div class="bu"><div tabindex="-1" class="cr"></div></div></div></a><div class="cl da db s t"><div><div class="bu"><div tabindex="-1" class="cr"><a class="by y bz ca cb ac cc ae af ag ah ai aj ak al" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2F77146fba8a08&amp;operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fgenpage-towards-end-to-end-generative-homepage-construction-at-netflix-77146fba8a08" rel="noopener follow"></a></div></div></div></div></div></div><div class="db s t"></div></div></div></div></div></div></div></div></article><article aria-label="How Netflix Simplified Batch Compute with Kueue" class="am" data-testid="post-preview"><div class="am vn e"><div class="ct"><div class="am e"><div class="dn am zk zl zm zn zo zp zq zr zs zt zu zv zw"><div class="zx"><div aria-label="How Netflix Simplified Batch Compute with Kueue"><div class="zz aba abb abc el"><img alt="How Netflix Simplified Batch Compute with Kueue" class="ct abd abe aea abg" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*GL_sNsqautjh4lnu5XmRww.png" /></div></div></div><div class="zy an ex gi"><div class="an gi uy ct abh abi abj abk"><div class="abl abm abn abo abp fe an k"><div class="fe an k abq"><div class="abr e ek"><div><div class="e"><div tabindex="-1" class="cr"><a href="https://netflixtechblog.com" rel="noopener follow"><div class="dn"><img alt="Netflix TechBlog" class="ej ig e abt abs" src="https://miro.medium.com/v2/resize:fill:40:40/1*ty4NvNrGg4ReETxqU2N3Og.png" width="20" height="20" /></div></a></div></div></div></div><div class="wv e ek de"><p class="ab b z v x">In</p></div><div class="fe e"><div><div class="e"><div tabindex="-1" class="cr"><a class="by y bz ca cb ac cc ae af ag ah ai aj lh an k" href="https://netflixtechblog.com" rel="noopener follow"><p class="ab b z v dw hr et hs ht hu hv hw cw">Netflix TechBlog</p></a></div></div></div></div></div><div class="abu e ek de"><p class="ab b z v x">by</p></div><div class="fe an k abq"><div class="fe e"><div><div class="e"><div tabindex="-1" class="cr"><a class="by y bz ca cb ac cc ae af ag ah ai aj lh an k" href="https://netflixtechblog.medium.com" rel="noopener follow"><p class="ab b z v dw hr et hs ht hu hv hw cw">Netflix Technology Blog</p></a></div></div></div></div></div><div class="nu an k ek dc de"><p class="ab b z v x">·</p><p class="ab b z v x">Jun 22</p></div></div><div class="abv e abw abx aby abz aca il"><div class="acb acc acd ace acf acg ach aci"><a class="by y bz ca cb ac cc ae af ag ah ai aj ak al" rel="noopener follow" href="https://netflixtechblog.com/how-netflix-simplified-batch-compute-with-kueue-87860682629c" data-discover="true"><div title=""><h2 class="ab jq pt pv acj ack acl acm acn aco fm pw py acp acq acr acs act acu fo fp fq acv acw acx acy acz ada fr fs ft adb adc add ade adf adg fu fv fw adh adi adj adk adl adm fx hw cw">How Netflix Simplified Batch Compute with Kueue</h2></div><div class="adn e"><h3 class="ab b ga v dw ado et hs adp hu hw x">By Alvin Bao, Alex Petrov, Jennifer Lai, Aidan Sherr, and Samartha Chandrashekar</h3></div></a></div></div><div class="wm an km cf"><div class="an k kx"><div class="cl da db s t"><div class="dn adq adr an k"><div class="ds gp ads an k kx"></div></div></div><div class="an k adu adv"><a class="by y bz ca cb ac cc ae af ag ah ai aj ak al" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fbuilding-service-topology-at-scale-architecture-challenges-and-lessons-learned-f4b792f3f0d8" rel="noopener follow"><div><div class="bu"><div tabindex="-1" class="cr"></div></div></div></a><div class="cl da db s t"><div><div class="bu"><div tabindex="-1" class="cr"><a class="by y bz ca cb ac cc ae af ag ah ai aj ak al" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2F87860682629c&amp;operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fhow-netflix-simplified-batch-compute-with-kueue-87860682629c" rel="noopener follow"></a></div></div></div></div></div></div><div class="db s t"></div></div></div></div></div></div></div></div></article><article aria-label="A Human-Augmenting Agentic Workflow for Causal Inference" class="am" data-testid="post-preview"><div class="am vn e"><div class="ct"><div class="am e"><div class="dn am zk zl zm zn zo zp zq zr zs zt zu zv zw"><div class="zx"><div aria-label="A Human-Augmenting Agentic Workflow for Causal Inference"><div class="zz aba abb abc el"><img alt="A Human-Augmenting Agentic Workflow for Causal Inference" class="ct abd abe abf abg" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*7midjxNL3H3BlZvQy4seSA.png" /></div></div></div><div class="zy an ex gi"><div class="an gi uy ct abh abi abj abk"><div class="abl abm abn abo abp fe an k"><div class="fe an k abq"><div class="abr e ek"><div><div class="e"><div tabindex="-1" class="cr"><a href="https://netflixtechblog.com" rel="noopener follow"><div class="dn"><img alt="Netflix TechBlog" class="ej ig e abt abs" src="https://miro.medium.com/v2/resize:fill:40:40/1*ty4NvNrGg4ReETxqU2N3Og.png" width="20" height="20" /></div></a></div></div></div></div><div class="wv e ek de"><p class="ab b z v x">In</p></div><div class="fe e"><div><div class="e"><div tabindex="-1" class="cr"><a class="by y bz ca cb ac cc ae af ag ah ai aj lh an k" href="https://netflixtechblog.com" rel="noopener follow"><p class="ab b z v dw hr et hs ht hu hv hw cw">Netflix TechBlog</p></a></div></div></div></div></div><div class="abu e ek de"><p class="ab b z v x">by</p></div><div class="fe an k abq"><div class="fe e"><div><div class="e"><div tabindex="-1" class="cr"><a class="by y bz ca cb ac cc ae af ag ah ai aj lh an k" href="https://netflixtechblog.medium.com" rel="noopener follow"><p class="ab b z v dw hr et hs ht hu hv hw cw">Netflix Technology Blog</p></a></div></div></div></div></div><div class="nu an k ek dc de"><p class="ab b z v x">·</p><p class="ab b z v x">Jun 8</p></div></div><div class="abv e abw abx aby abz aca il"><div class="acb acc acd ace acf acg ach aci"><a class="by y bz ca cb ac cc ae af ag ah ai aj ak al" rel="noopener follow" href="https://netflixtechblog.com/a-human-augmenting-agentic-workflow-for-causal-inference-4623f0a9c5af" data-discover="true"><div title=""><h2 class="ab jq pt pv acj ack acl acm acn aco fm pw py acp acq acr acs act acu fo fp fq acv acw acx acy acz ada fr fs ft adb adc add ade adf adg fu fv fw adh adi adj adk adl adm fx hw cw">A Human-Augmenting Agentic Workflow for Causal Inference</h2></div><div class="adn e"><h3 class="ab b ga v dw ado et hs adp hu hw x">By Winston Chou, Adrien Alexandre, Lars Olds, Yi Zhang, Garrett Hagemann, and Nathan Kallus</h3></div></a></div></div><div class="wm an km cf"><div class="an k kx"><div class="cl da db s t"><div class="dn adq adr an k"><div class="ds gp ads an k kx"><a class="ds he ads an k kx" tabindex="-1" rel="noopener follow" href="https://netflixtechblog.com/a-human-augmenting-agentic-workflow-for-causal-inference-4623f0a9c5af" data-discover="true"><div><div class="an"><div tabindex="-1" class="cr"><div class="an k dc">A response icon9</div></div></div></div></a></div></div></div><div class="an k adu adv"><a class="by y bz ca cb ac cc ae af ag ah ai aj ak al" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fbuilding-service-topology-at-scale-architecture-challenges-and-lessons-learned-f4b792f3f0d8" rel="noopener follow"><div><div class="bu"><div tabindex="-1" class="cr"></div></div></div></a><div class="cl da db s t"><div><div class="bu"><div tabindex="-1" class="cr"><a class="by y bz ca cb ac cc ae af ag ah ai aj ak al" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2F4623f0a9c5af&amp;operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fa-human-augmenting-agentic-workflow-for-causal-inference-4623f0a9c5af" rel="noopener follow"></a></div></div></div></div></div></div><div class="db s t"></div></div></div></div></div></div></div></div></article><article aria-label="VMAF v1: Good Is Not Good Enough" class="am" data-testid="post-preview"><div class="am vn e"><div class="ct"><div class="am e"><div class="dn am zk zl zm zn zo zp zq zr zs zt zu zv zw"><div class="zx"><div aria-label="VMAF v1: Good Is Not Good Enough"><div class="zz aba abb abc el"><img alt="VMAF v1: Good Is Not Good Enough" class="ct abd abe abf abg" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*4Ymg7tqfTpT3SUSQX92smw.png" /></div></div></div><div class="zy an ex gi"><div class="an gi uy ct abh abi abj abk"><div class="abl abm abn abo abp fe an k"><div class="fe an k abq"><div class="abr e ek"><div><div class="e"><div tabindex="-1" class="cr"><a href="https://netflixtechblog.com" rel="noopener follow"><div class="dn"><img alt="Netflix TechBlog" class="ej ig e abt abs" src="https://miro.medium.com/v2/resize:fill:40:40/1*ty4NvNrGg4ReETxqU2N3Og.png" width="20" height="20" /></div></a></div></div></div></div><div class="wv e ek de"><p class="ab b z v x">In</p></div><div class="fe e"><div><div class="e"><div tabindex="-1" class="cr"><a class="by y bz ca cb ac cc ae af ag ah ai aj lh an k" href="https://netflixtechblog.com" rel="noopener follow"><p class="ab b z v dw hr et hs ht hu hv hw cw">Netflix TechBlog</p></a></div></div></div></div></div><div class="abu e ek de"><p class="ab b z v x">by</p></div><div class="fe an k abq"><div class="fe e"><div><div class="e"><div tabindex="-1" class="cr"><a class="by y bz ca cb ac cc ae af ag ah ai aj lh an k" href="https://netflixtechblog.medium.com" rel="noopener follow"><p class="ab b z v dw hr et hs ht hu hv hw cw">Netflix Technology Blog</p></a></div></div></div></div></div><div class="nu an k ek dc de"><p class="ab b z v x">·</p><p class="ab b z v x">Jun 19</p></div></div><div class="abv e abw abx aby abz aca il"><div class="acb acc acd ace acf acg ach aci"><a class="by y bz ca cb ac cc ae af ag ah ai aj ak al" rel="noopener follow" href="https://netflixtechblog.com/vmaf-v1-good-is-not-good-enough-60d7e4244ea8" data-discover="true"><div title=""><h2 class="ab jq pt pv acj ack acl acm acn aco fm pw py acp acq acr acs act acu fo fp fq acv acw acx acy acz ada fr fs ft adb adc add ade adf adg fu fv fw adh adi adj adk adl adm fx hw cw">VMAF v1: Good Is Not Good Enough</h2></div><div class="adn e"><h3 class="ab b ga v dw ado et hs adp hu hw x">By Christos G. Bampis, Zhi Li, Kyle Swanson, Nil Fons Miret and Pavan Madhusudanarao</h3></div></a></div></div><div class="wm an km cf"><div class="an k kx"><div class="cl da db s t"><div class="dn adq adr an k"><div class="ds gp ads an k kx"></div></div></div><div class="an k adu adv"><a class="by y bz ca cb ac cc ae af ag ah ai aj ak al" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fbuilding-service-topology-at-scale-architecture-challenges-and-lessons-learned-f4b792f3f0d8" rel="noopener follow"><div><div class="bu"><div tabindex="-1" class="cr"></div></div></div></a><div class="cl da db s t"><div><div class="bu"><div tabindex="-1" class="cr"><a class="by y bz ca cb ac cc ae af ag ah ai aj ak al" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2F60d7e4244ea8&amp;operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fvmaf-v1-good-is-not-good-enough-60d7e4244ea8" rel="noopener follow"></a></div></div></div></div></div></div></div></div></div></div></div></div></div></article><article aria-label="Should You Still Learn to Code in 2026?" class="am" data-testid="post-preview"><div class="am vn e"><div class="ct"><div class="am e"><div class="dn am zk zl zm zn zo zp zq zr zs zt zu zv zw"><div class="zx"><div aria-label="Should You Still Learn to Code in 2026?"><div class="zz aba abb abc el"><img alt="Should You Still Learn to Code in 2026?" class="ct abd abe abf abg" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*gasaRNc3_X8K96F58zKyvw.jpeg" /></div></div></div><div class="zy an ex gi"><div class="an gi uy ct abh abi abj abk"><div class="abl abm abn abo abp fe an k"><div class="fe an k abq"><div class="abr e ek"><div><div class="e"><div tabindex="-1" class="cr"><a href="https://medium.com/data-science-collective" rel="noopener follow"><div class="dn"><img alt="Data Science Collective" class="ej ig e abt abs" src="https://miro.medium.com/v2/resize:fill:40:40/1*0nV0Q-FBHj94Kggq00pG2Q.jpeg" width="20" height="20" /></div></a></div></div></div></div><div class="wv e ek de"><p class="ab b z v x">In</p></div><div class="fe e"><div><div class="e"><div tabindex="-1" class="cr"><a class="by y bz ca cb ac cc ae af ag ah ai aj lh an k" href="https://medium.com/data-science-collective" rel="noopener follow"><p class="ab b z v dw hr et hs ht hu hv hw cw">Data Science Collective</p></a></div></div></div></div></div><div class="abu e ek de"><p class="ab b z v x">by</p></div><div class="fe an k abq"><div class="fe e"><div><div class="e"><div tabindex="-1" class="cr"><a class="by y bz ca cb ac cc ae af ag ah ai aj lh an k" href="https://medium.com/@gratitudedriven" rel="noopener follow"><p class="ab b z v dw hr et hs ht hu hv hw cw">Marina Wyss</p></a></div></div></div></div></div><div class="nu an k ek dc de"><p class="ab b z v x">·</p><p class="ab b z v x">Feb 23</p></div></div><div class="abv e abw abx aby abz aca il"><div class="acb acc acd ace acf acg ach aci"><a class="by y bz ca cb ac cc ae af ag ah ai aj ak al" href="https://medium.com/data-science-collective/should-you-still-learn-to-code-in-2026-034685e17707" rel="noopener follow"><div title=""><h2 class="ab jq pt pv acj ack acl acm acn aco fm pw py acp acq acr acs act acu fo fp fq acv acw acx acy acz ada fr fs ft adb adc add ade adf adg fu fv fw adh adi adj adk adl adm fx hw cw">Should You Still Learn to Code in 2026?</h2></div><div class="adn e"><h3 class="ab b ga v dw ado et hs adp hu hw x">The answer isn’t as obvious as I used to believe.</h3></div></a></div></div><div class="wm an km cf"><div class="an k kx"><div class="vn an"><button class="e cg ag ae" aria-label="Member-only story"><div class=""><div><div class="bu"><div tabindex="-1" class="cr"></div></div></div></div></button></div><div class="cl da db s t"><div class="dn adq adr an k"><div class="ds gp ads an k kx"><a class="ds he ads an k kx" tabindex="-1" href="https://medium.com/data-science-collective/should-you-still-learn-to-code-in-2026-034685e17707" rel="noopener follow"><div><div class="an"><div tabindex="-1" class="cr"><div class="an k dc">A response icon283</div></div></div></div></a></div></div></div><div class="an k adu adv"><a class="by y bz ca cb ac cc ae af ag ah ai aj ak al" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fbuilding-service-topology-at-scale-architecture-challenges-and-lessons-learned-f4b792f3f0d8" rel="noopener follow"><div><div class="bu"><div tabindex="-1" class="cr"></div></div></div></a><div class="cl da db s t"><div><div class="bu"><div tabindex="-1" class="cr"><a class="by y bz ca cb ac cc ae af ag ah ai aj ak al" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2F034685e17707&amp;operation=register&amp;redirect=https%3A%2F%2Fmedium.com%2Fdata-science-collective%2Fshould-you-still-learn-to-code-in-2026-034685e17707" rel="noopener follow"></a></div></div></div></div></div></div><div class="db s t"></div></div></div></div></div></div></div></div></article><article aria-label="Stop Memorizing Design Patterns: Use This Decision Tree Instead" class="am" data-testid="post-preview"><div class="am vn e"><div class="ct"><div class="am e"><div class="dn am zk zl zm zn zo zp zq zr zs zt zu zv zw"><div class="zx"><div aria-label="Stop Memorizing Design Patterns: Use This Decision Tree Instead"><div class="zz aba abb abc el"><img alt="Stop Memorizing Design Patterns: Use This Decision Tree Instead" class="ct abd abe abf abg" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*xfboC-sVIT2hzWkgQZT_7w.png" /></div></div></div><div class="zy an ex gi"><div class="an gi uy ct abh abi abj abk"><div class="abl abm abn abo abp fe an k"><div class="fe an k abq"><div class="abr e ek"><div><div class="e"><div tabindex="-1" class="cr"><a href="https://medium.com/womenintechnology" rel="noopener follow"><div class="dn"><img alt="Women in Technology" class="ej ig e abt abs" src="https://miro.medium.com/v2/resize:fill:40:40/1*kd0DvPkLdn59Emtg_rnsqg.png" width="20" height="20" /></div></a></div></div></div></div><div class="wv e ek de"><p class="ab b z v x">In</p></div><div class="fe e"><div><div class="e"><div tabindex="-1" class="cr"><a class="by y bz ca cb ac cc ae af ag ah ai aj lh an k" href="https://medium.com/womenintechnology" rel="noopener follow"><p class="ab b z v dw hr et hs ht hu hv hw cw">Women in Technology</p></a></div></div></div></div></div><div class="abu e ek de"><p class="ab b z v x">by</p></div><div class="fe an k abq"><div class="fe e"><div><div class="e"><div tabindex="-1" class="cr"><a class="by y bz ca cb ac cc ae af ag ah ai aj lh an k" href="https://medium.com/@akovtun" rel="noopener follow"><p class="ab b z v dw hr et hs ht hu hv hw cw">Alina Kovtun✨</p></a></div></div></div></div></div><div class="nu an k ek dc de"><p class="ab b z v x">·</p><p class="ab b z v x">Jan 29</p></div></div><div class="abv e abw abx aby abz aca il"><div class="acb acc acd ace acf acg ach aci"><a class="by y bz ca cb ac cc ae af ag ah ai aj ak al" href="https://medium.com/womenintechnology/stop-memorizing-design-patterns-use-this-decision-tree-instead-e84f22fca9fa" rel="noopener follow"><div title=""><h2 class="ab jq pt pv acj ack acl acm acn aco fm pw py acp acq acr acs act acu fo fp fq acv acw acx acy acz ada fr fs ft adb adc add ade adf adg fu fv fw adh adi adj adk adl adm fx hw cw">Stop Memorizing Design Patterns: Use This Decision Tree Instead</h2></div><div class="adn e"><h3 class="ab b ga v dw ado et hs adp hu hw x">Choose design patterns based on pain points: apply the right pattern with minimal over-engineering in any OO language.</h3></div></a></div></div><div class="wm an km cf"><div class="an k kx"><div class="vn an"><button class="e cg ag ae" aria-label="Member-only story"><div class=""><div><div class="bu"><div tabindex="-1" class="cr"></div></div></div></div></button></div><div class="cl da db s t"><div class="dn adq adr an k"><div class="ds gp ads an k kx"><a class="ds he ads an k kx" tabindex="-1" href="https://medium.com/womenintechnology/stop-memorizing-design-patterns-use-this-decision-tree-instead-e84f22fca9fa" rel="noopener follow"><div><div class="an"><div tabindex="-1" class="cr"><div class="an k dc">A response icon94</div></div></div></div></a></div></div></div><div class="an k adu adv"><a class="by y bz ca cb ac cc ae af ag ah ai aj ak al" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fbuilding-service-topology-at-scale-architecture-challenges-and-lessons-learned-f4b792f3f0d8" rel="noopener follow"><div><div class="bu"><div tabindex="-1" class="cr"></div></div></div></a><div class="cl da db s t"><div><div class="bu"><div tabindex="-1" class="cr"><a class="by y bz ca cb ac cc ae af ag ah ai aj ak al" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2Fe84f22fca9fa&amp;operation=register&amp;redirect=https%3A%2F%2Fmedium.com%2Fwomenintechnology%2Fstop-memorizing-design-patterns-use-this-decision-tree-instead-e84f22fca9fa" rel="noopener follow"></a></div></div></div></div></div></div></div></div></div></div></div></div></div></article><article aria-label="7 Coding Patterns I Stole From Senior Engineers" class="am" data-testid="post-preview"><div class="am vn e"><div class="ct"><div class="am e"><div class="dn am zk zl zm zn zo zp zq zr zs zt zu zv zw"><div class="zx"><div aria-label="7 Coding Patterns I Stole From Senior Engineers"><div class="zz aba abb abc el"><img alt="7 Coding Patterns I Stole From Senior Engineers" class="ct abd abe abf abg" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*ocCz_MiZvCdnqNNFBaA_gA.png" /></div></div></div><div class="zy an ex gi"><div class="an gi uy ct abh abi abj abk"><div class="abl abm abn abo abp fe an k"><div class="fe an k abq"><div class="abr e ek"><div><div class="e"><div tabindex="-1" class="cr"><a href="https://medium.com/skillstuff" rel="noopener follow"><div class="dn"><img alt="Skill Stuff" class="ej ig e abt abs" src="https://miro.medium.com/v2/resize:fill:40:40/1*T3BugcdkX-841qoxqSdQiQ.jpeg" width="20" height="20" /></div></a></div></div></div></div><div class="wv e ek de"><p class="ab b z v x">In</p></div><div class="fe e"><div><div class="e"><div tabindex="-1" class="cr"><a class="by y bz ca cb ac cc ae af ag ah ai aj lh an k" href="https://medium.com/skillstuff" rel="noopener follow"><p class="ab b z v dw hr et hs ht hu hv hw cw">Skill Stuff</p></a></div></div></div></div></div><div class="abu e ek de"><p class="ab b z v x">by</p></div><div class="fe an k abq"><div class="fe e"><div><div class="e"><div tabindex="-1" class="cr"><a class="by y bz ca cb ac cc ae af ag ah ai aj lh an k" href="https://codebyumar.medium.com" rel="noopener follow"><p class="ab b z v dw hr et hs ht hu hv hw cw">CodeByUmar</p></a></div></div></div></div></div><div class="nu an k ek dc de"><p class="ab b z v x">·</p><p class="ab b z v x">Jun 8</p></div></div><div class="abv e abw abx aby abz aca il"><div class="acb acc acd ace acf acg ach aci"><a class="by y bz ca cb ac cc ae af ag ah ai aj ak al" href="https://medium.com/skillstuff/7-coding-patterns-i-stole-from-senior-engineers-c95f757e52a6" rel="noopener follow"><div title=""><h2 class="ab jq pt pv acj ack acl acm acn aco fm pw py acp acq acr acs act acu fo fp fq acv acw acx acy acz ada fr fs ft adb adc add ade adf adg fu fv fw adh adi adj adk adl adm fx hw cw">7 Coding Patterns I Stole From Senior Engineers</h2></div><div class="adn e"><h3 class="ab b ga v dw ado et hs adp hu hw x">Most developers do not become better because they learn more syntax. They become better because they stop making code harder than it needs…</h3></div></a></div></div><div class="wm an km cf"><div class="an k kx"><div class="vn an"><button class="e cg ag ae" aria-label="Member-only story"><div class=""><div><div class="bu"><div tabindex="-1" class="cr"></div></div></div></div></button></div><div class="cl da db s t"><div class="dn adq adr an k"><div class="ds gp ads an k kx"><a class="ds he ads an k kx" tabindex="-1" href="https://medium.com/skillstuff/7-coding-patterns-i-stole-from-senior-engineers-c95f757e52a6" rel="noopener follow"><div><div class="an"><div tabindex="-1" class="cr"><div class="an k dc">A response icon60</div></div></div></div></a></div></div></div><div class="an k adu adv"><a class="by y bz ca cb ac cc ae af ag ah ai aj ak al" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fbuilding-service-topology-at-scale-architecture-challenges-and-lessons-learned-f4b792f3f0d8" rel="noopener follow"><div><div class="bu"><div tabindex="-1" class="cr"></div></div></div></a><div class="cl da db s t"><div><div class="bu"><div tabindex="-1" class="cr"><a class="by y bz ca cb ac cc ae af ag ah ai aj ak al" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2Fc95f757e52a6&amp;operation=register&amp;redirect=https%3A%2F%2Fmedium.com%2Fskillstuff%2F7-coding-patterns-i-stole-from-senior-engineers-c95f757e52a6" rel="noopener follow"></a></div></div></div></div></div></div><div class="db s t"></div></div></div></div></div></div></div></div></article><article aria-label="How to Create Loops with Claude Code: A Practical Guide to Agentic Automation" class="am" data-testid="post-preview"><div class="am vn e"><div class="ct"><div class="am e"><div class="dn am zk zl zm zn zo zp zq zr zs zt zu zv zw"><div class="zx"><div aria-label="How to Create Loops with Claude Code: A Practical Guide to Agentic Automation"><div class="zz aba abb abc el"><img alt="How to Create Loops with Claude Code: A Practical Guide to Agentic Automation" class="ct abd abe abf abg" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*WbMmltJlf6xsaKxctAiUpg.png" /></div></div></div><div class="zy an ex gi"><div class="an gi uy ct abh abi abj abk"><div class="abl abm abn abo abp fe an k"><div class="fe an k abq"><div class="abr e ek"><div><div class="e"><div tabindex="-1" class="cr"><a href="https://levelup.gitconnected.com" rel="noopener follow"><div class="dn"><img alt="Level Up Coding" class="ej ig e abt abs" src="https://miro.medium.com/v2/resize:fill:40:40/1*5D9oYBd58pyjMkV_5-zXXQ.jpeg" width="20" height="20" /></div></a></div></div></div></div><div class="wv e ek de"><p class="ab b z v x">In</p></div><div class="fe e"><div><div class="e"><div tabindex="-1" class="cr"><a class="by y bz ca cb ac cc ae af ag ah ai aj lh an k" href="https://levelup.gitconnected.com" rel="noopener follow"><p class="ab b z v dw hr et hs ht hu hv hw cw">Level Up Coding</p></a></div></div></div></div></div><div class="abu e ek de"><p class="ab b z v x">by</p></div><div class="fe an k abq"><div class="fe e"><div><div class="e"><div tabindex="-1" class="cr"><a class="by y bz ca cb ac cc ae af ag ah ai aj lh an k" href="https://yousefhosni.medium.com" rel="noopener follow"><p class="ab b z v dw hr et hs ht hu hv hw cw">Youssef Hosni</p></a></div></div></div></div></div><div class="nu an k ek dc de"><p class="ab b z v x">·</p><p class="ab b z v x">Jun 24</p></div></div><div class="abv e abw abx aby abz aca il"><div class="acb acc acd ace acf acg ach aci"><a class="by y bz ca cb ac cc ae af ag ah ai aj ak al" href="https://levelup.gitconnected.com/how-to-create-loops-with-claude-code-a-practical-guide-to-agentic-automation-6f422390a143" rel="noopener follow"><div title="How to Create Loops with Claude Code: A Practical Guide to Agentic Automation"><h2 class="ab jq pt pv acj ack acl acm acn aco fm pw py acp acq acr acs act acu fo fp fq acv acw acx acy acz ada fr fs ft adb adc add ade adf adg fu fv fw adh adi adj adk adl adm fx hw cw">How to Create Loops with Claude Code: A Practical Guide to Agentic Automation</h2></div><div class="adn e"><h3 class="ab b ga v dw ado et hs adp hu hw x">Claude Loop Engineering: How to Design Agents That Iterate, Verify, and Remember</h3></div></a></div></div><div class="wm an km cf"><div class="an k kx"><div class="vn an"><button class="e cg ag ae" aria-label="Member-only story"><div class=""><div><div class="bu"><div tabindex="-1" class="cr"></div></div></div></div></button></div><div class="cl da db s t"><div class="dn adq adr an k"><div class="ds gp ads an k kx"><a class="ds he ads an k kx" tabindex="-1" href="https://levelup.gitconnected.com/how-to-create-loops-with-claude-code-a-practical-guide-to-agentic-automation-6f422390a143" rel="noopener follow"><div><div class="an"><div tabindex="-1" class="cr"><div class="an k dc">A response icon11</div></div></div></div></a></div></div></div><div class="an k adu adv"><a class="by y bz ca cb ac cc ae af ag ah ai aj ak al" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fbuilding-service-topology-at-scale-architecture-challenges-and-lessons-learned-f4b792f3f0d8" rel="noopener follow"><div><div class="bu"><div tabindex="-1" class="cr"></div></div></div></a><div class="cl da db s t"><div><div class="bu"><div tabindex="-1" class="cr"><a class="by y bz ca cb ac cc ae af ag ah ai aj ak al" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2F6f422390a143&amp;operation=register&amp;redirect=https%3A%2F%2Flevelup.gitconnected.com%2Fhow-to-create-loops-with-claude-code-a-practical-guide-to-agentic-automation-6f422390a143" rel="noopener follow"></a></div></div></div></div></div></div><div class="db s t"></div></div></div></div></div></div></div></div></article><article aria-label="The Documentation is the Code" class="am" data-testid="post-preview"><div class="am vn e"><div class="ct"><div class="am e"><div class="dn am zk zl zm zn zo zp zq zr zs zt zu zv zw"><div class="zx"><div aria-label="The Documentation is the Code"><div class="zz aba abb abc el"><img alt="The Documentation is the Code" class="ct abd abe abf abg" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*mvwohyoavkIDBA0PlMoh4w.png" /></div></div></div><div class="zy an ex gi"><div class="an gi uy ct abh abi abj abk"><div class="abl abm abn abo abp fe an k"><div class="fe an k abq"><div class="abr e ek"><div><div class="e"><div tabindex="-1" class="cr"><a href="https://medium.com/intuitively-and-exhaustively-explained" rel="noopener follow"><div class="dn"><img alt="Intuitively and Exhaustively Explained" class="ej ig e abt abs" src="https://miro.medium.com/v2/resize:fill:40:40/1*QB9f-TX5vNkxURqqruU0jw.png" width="20" height="20" /></div></a></div></div></div></div><div class="wv e ek de"><p class="ab b z v x">In</p></div><div class="fe e"><div><div class="e"><div tabindex="-1" class="cr"><a class="by y bz ca cb ac cc ae af ag ah ai aj lh an k" href="https://medium.com/intuitively-and-exhaustively-explained" rel="noopener follow"><p class="ab b z v dw hr et hs ht hu hv hw cw">Intuitively and Exhaustively Explained</p></a></div></div></div></div></div><div class="abu e ek de"><p class="ab b z v x">by</p></div><div class="fe an k abq"><div class="fe e"><div><div class="e"><div tabindex="-1" class="cr"><a class="by y bz ca cb ac cc ae af ag ah ai aj lh an k" href="https://medium.com/@danielwarfield1" rel="noopener follow"><p class="ab b z v dw hr et hs ht hu hv hw cw">Daniel Warfield</p></a></div></div></div></div></div><div class="nu an k ek dc de"><p class="ab b z v x">·</p><p class="ab b z v x">Jun 29</p></div></div><div class="abv e abw abx aby abz aca il"><div class="acb acc acd ace acf acg ach aci"><a class="by y bz ca cb ac cc ae af ag ah ai aj ak al" href="https://medium.com/intuitively-and-exhaustively-explained/the-documentation-is-the-code-52b66da25d29" rel="noopener follow"><div title=""><h2 class="ab jq pt pv acj ack acl acm acn aco fm pw py acp acq acr acs act acu fo fp fq acv acw acx acy acz ada fr fs ft adb adc add ade adf adg fu fv fw adh adi adj adk adl adm fx hw cw">The Documentation is the Code</h2></div><div class="adn e"><h3 class="ab b ga v dw ado et hs adp hu hw x">A great inversion in software development</h3></div></a></div></div><div class="wm an km cf"><div class="an k kx"><div class="vn an"><button class="e cg ag ae" aria-label="Member-only story"><div class=""><div><div class="bu"><div tabindex="-1" class="cr"></div></div></div></div></button></div><div class="cl da db s t"><div class="dn adq adr an k"><div class="ds gp ads an k kx"><a class="ds he ads an k kx" tabindex="-1" href="https://medium.com/intuitively-and-exhaustively-explained/the-documentation-is-the-code-52b66da25d29" rel="noopener follow"><div><div class="an"><div tabindex="-1" class="cr"><div class="an k dc">A response icon13</div></div></div></div></a></div></div></div><div class="an k adu adv"><a class="by y bz ca cb ac cc ae af ag ah ai aj ak al" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fbuilding-service-topology-at-scale-architecture-challenges-and-lessons-learned-f4b792f3f0d8" rel="noopener follow"><div><div class="bu"><div tabindex="-1" class="cr"></div></div></div></a><div class="cl da db s t"><div><div class="bu"><div tabindex="-1" class="cr"><a class="by y bz ca cb ac cc ae af ag ah ai aj ak al" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2F52b66da25d29&amp;operation=register&amp;redirect=https%3A%2F%2Fmedium.com%2Fintuitively-and-exhaustively-explained%2Fthe-documentation-is-the-code-52b66da25d29" rel="noopener follow"></a></div></div></div></div></div></div><div class="db s t"></div></div></div></div></div></div></div></div></article><article aria-label="System Design Interview: How Would You Send 1 Million Notifications Without Overwhelming Your…" class="am" data-testid="post-preview"><div class="am vn e"><div class="ct"><div class="am e"><div class="dn am zk zl zm zn zo zp zq zr zs zt zu zv zw"><div class="zx"><div aria-label="System Design Interview: How Would You Send 1 Million Notifications Without Overwhelming Your…"><div class="zz aba abb abc el"><img alt="System Design Interview: How Would You Send 1 Million Notifications Without Overwhelming Your…" class="ct abd abe abf abg" src="https://miro.medium.com/v2/resize:fit:1358/format:webp/1*_a0MnVSLuc2Ivt2OEXytGw.png" /></div></div></div><div class="zy an ex gi"><div class="an gi uy ct abh abi abj abk"><div class="abl abm abn abo abp fe an k"><div class="fe an k abq"><div class="abr e ek"><div><div class="e"><div tabindex="-1" class="cr"><a tabindex="-1" href="https://codefarm0.medium.com" rel="noopener follow"><div class="e dn"><img alt="Arvind Kumar" class="e bt ef abs abt ej" src="https://miro.medium.com/v2/resize:fill:40:40/1*qLgT62h04Vn1WA1vdYL9lg.png" width="20" height="20" /></div></a></div></div></div></div><div class="fe e"><div><div class="e"><div tabindex="-1" class="cr"><a class="by y bz ca cb ac cc ae af ag ah ai aj lh an k" href="https://codefarm0.medium.com" rel="noopener follow"><p class="ab b z v dw hr et hs ht hu hv hw cw">Arvind Kumar</p></a></div></div></div></div></div><div class="nu an k ek dc de"><p class="ab b z v x">·</p><p class="ab b z v x">Jun 20</p></div></div><div class="abv e abw abx aby abz aca il"><div class="acb acc acd ace acf acg ach aci"><a class="by y bz ca cb ac cc ae af ag ah ai aj ak al" href="https://codefarm0.medium.com/system-design-interview-how-would-you-send-1-million-notifications-without-overwhelming-your-5a59b723127d" rel="noopener follow"><div title="System Design Interview: How Would You Send 1 Million Notifications Without Overwhelming Your…"><h2 class="ab jq pt pv acj ack acl acm acn aco fm pw py acp acq acr acs act acu fo fp fq acv acw acx acy acz ada fr fs ft adb adc add ade adf adg fu fv fw adh adi adj adk adl adm fx hw cw">System Design Interview: How Would You Send 1 Million Notifications Without Overwhelming Your…</h2></div><div class="adn e"><h3 class="ab b ga v dw ado et hs adp hu hw x">It’s Black Friday.</h3></div></a></div></div><div class="wm an km cf"><div class="an k kx"><div class="vn an"><button class="e cg ag ae" aria-label="Member-only story"><div class=""><div><div class="bu"><div tabindex="-1" class="cr"></div></div></div></div></button></div><div class="cl da db s t"><div class="dn adq adr an k"><div class="ds gp ads an k kx"><a class="ds he ads an k kx" tabindex="-1" href="https://codefarm0.medium.com/system-design-interview-how-would-you-send-1-million-notifications-without-overwhelming-your-5a59b723127d" rel="noopener follow"><div><div class="an"><div tabindex="-1" class="cr"><div class="an k dc">A response icon19</div></div></div></div></a></div></div></div><div class="an k adu adv"><a class="by y bz ca cb ac cc ae af ag ah ai aj ak al" href="https://medium.com/m/signin?operation=register&amp;redirect=https%3A%2F%2Fnetflixtechblog.com%2Fbuilding-service-topology-at-scale-architecture-challenges-and-lessons-learned-f4b792f3f0d8" rel="noopener follow"><div><div class="bu"><div tabindex="-1" class="cr"></div></div></div></a><div class="cl da db s t"><div><div class="bu"><div tabindex="-1" class="cr"><a class="by y bz ca cb ac cc ae af ag ah ai aj ak al" href="https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2F_%2Fbookmark%2Fp%2F5a59b723127d&amp;operation=register&amp;redirect=https%3A%2F%2Fcodefarm0.medium.com%2Fsystem-design-interview-how-would-you-send-1-million-notifications-without-overwhelming-your-5a59b723127d" rel="noopener follow"></a></div></div></div></div></div></div></div></div></div></div></div></div></div></article>]]></description>
      <link>https://netflixtechblog.com/building-service-topology-at-scale-architecture-challenges-and-lessons-learned-f4b792f3f0d8</link>
      <guid>https://netflixtechblog.com/building-service-topology-at-scale-architecture-challenges-and-lessons-learned-f4b792f3f0d8</guid>
      <pubDate>Tue, 14 Jul 2026 00:44:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[GenPage: Towards End-to-End Generative Homepage Construction at Netflix]]></title>
      <description><![CDATA[<p>Authors: <a href="https://www.linkedin.com/in/lequn-luke-wang-9226b2129/">Lequn Wang</a>, J<a href="https://www.linkedin.com/in/jiangwei-pan-66a62a13/">iangwei Pan</a>, and <a href="https://www.linkedin.com/in/linasbaltrunas/">Linas Baltrunas</a></p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*NurrizMgC7_QbsGuuW42Dg.gif"><figcaption><strong>Figure 1. </strong>Autoregressive homepage generation. GenPage builds a Netflix homepage one row or entity at a time, each one conditioned on what’s already on the page and the user’s context.</figcaption></figure><h3>Introduction</h3><p>The Netflix homepage is the first thing users see when they open the app and the primary way they discover content to enjoy. Almost every part of it is personalized, including which rows appear, which entities show up within those rows, and how everything is arranged on the page.</p><p>Constructing that homepage is a genuinely hard problem. It is not simply producing one ranked list. The homepage is a structured, two-dimensional layout, made up of recommendation rows and the entities within them. Here, an entity can be a movie, show, game, live event, or other recommendable item. Each choice can affect the value of the others. Traditionally, it is built through a complex, multi-stage pipeline, with separate components for candidate generation and ranking at both the row and entity levels.</p><p>We saw an opportunity to rethink this design. Large language models have shown that a single generative model can perform diverse tasks just by generating a response to a prompt. Inspired by this prompt-response paradigm, we trained a single generative model to build the homepage by directly answering one question:</p><blockquote>Given everything we know about this user and this request, what homepage should we generate to maximize user satisfaction?</blockquote><p>We call this approach GenPage. It treats the user history and request context as the prompt, and autoregressively generates the entire homepage as the response (Figure 1). Unlike most generative recommenders, such as<a href="https://arxiv.org/abs/2305.05065"> TIGER</a>,<a href="https://arxiv.org/abs/2402.17152"> HSTU</a>, and<a href="https://arxiv.org/abs/2506.13695"> OneRec</a>, which generate flat ranked lists, GenPage generates the rows, entities, and layout together.</p><p>This shift is motivated by several goals:</p><ul><li><strong>End-to-end modeling.</strong> A single transformer model that constructs the page from raw input signals can replace a complex multi-stage recommender stack. This reduces the number of ML models to maintain, avoids misaligned objectives across stages, and eliminates much of the traditional feature engineering.</li><li><strong>Whole-page optimization via reinforcement learning (RL).</strong> Autoregressive page generation makes it possible to optimize for page-level rewards with RL. This can capture interactions across rows and entities, such as diversity or the balance between rows with different <em>stopping power</em>. For example, a Continue Watching row near the top of the page may strongly satisfy a user’s immediate intent, but also reduce how much of the page they browse. Modeling these interactions at the page level lets us align the system more directly with user satisfaction than entity-level objectives alone.</li><li><strong>Better scaling behavior.</strong> A generative transformer model gives us a clearer path to improving quality through more data, compute, and model capacity, without repeatedly redesigning the system.</li><li><strong>Flexibility and extensibility.</strong> The prompt-response paradigm is flexible by design. By simplifying feature engineering and enabling whole-page optimization, GenPage makes it easier to support new product experiences, such as additional content types like live events, games, and podcasts; layouts beyond the current two-dimensional structure; personalized UI components; and per-entity artwork personalization, all with fewer architectural changes.</li></ul><p>Bringing GenPage into production at Netflix also required solving challenges specific to industry-scale recommender systems. Because the homepage is generated in real time, serving latency is a primary engineering constraint. We also need to handle entity cold start in a constantly evolving catalog, keep the model fresh as user interests and cultural trends shift, and enforce complex product and business rules on the generated output.</p><p>Despite these challenges, GenPage has already had substantial production impact. In an online A/B test against a mature, highly optimized multi-stage production recommender, GenPage delivered statistically significant gains on the core user engagement metric we use for launch decisions, while reducing end-to-end serving latency by 20%.</p><p>Offline, two findings stood out. First, enriching the prompt helped more than scaling model capacity in our current regime. Second, RL post-training increased homepage diversity even though diversity was not part of the objective.</p><p>We expect this approach to generalize to many personalization settings. In this post, we focus on Netflix homepage construction as a concrete case study, sharing our design, trade-offs, and lessons learned.</p><h3>Data</h3><p>Moving from a traditional recommender to a generative transformer requires us to rethink how the data is represented. Similar to how an LLM turns text into tokens, GenPage represents both the user context and the generated homepage as one sequence of discrete tokens (Figure 2). This sequence includes the full structured homepage layout, with multiple rows and the entities inside them, so the model can generate the page holistically rather than scoring each row or entity in isolation.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*RL-TMAFPo93aO3GocVBBvQ.png"><figcaption><strong>Figure 2. </strong>Tokenization of Netflix homepage construction data. The context tokens function as the prompt, drawing from diverse data sources including user history, profile attributes, and request context, with example tokens shown for each source. The page tokens represent the generated response, encoding the structured layout of rows and entities.</figcaption></figure><p>Each training example represents a homepage impression and consists of three components:</p><ul><li><strong>Context: </strong>user engagement history, profile attributes, and request context.</li><li><strong>Page:</strong> the recommended rows and entities shown on the homepage, in layout order.</li><li><strong>Feedback:</strong> user interactions with that page, such as play, thumbs-up, or abandonment for entities on the page.</li></ul><p>Only the context and page are tokenized as model inputs and outputs. Feedback is used to derive supervision signals via our internal reward system (see the Reward system section).</p><p>Instead of using an off-the-shelf text tokenizer, we build a domain-specific tokenizer for the homepage construction data. This is a proven approach in<a href="https://arxiv.org/abs/1808.09781"> recommender systems</a> and other specialized domains including<a href="https://arxiv.org/abs/2102.12092"> computer vision</a>,<a href="https://www.pnas.org/doi/10.1073/pnas.2016239118"> biology</a>, and<a href="https://arxiv.org/abs/1811.02633"> chemistry</a>, where the raw data is not naturally represented as text. Compared with generic text tokenization, this gives us two key advantages:</p><ul><li><strong>Computational efficiency.</strong> Custom tokenization significantly reduces sequence length, lowering inference cost and latency. For example, representing the event “User watched Orange Is the New Black for 50 minutes 30 days ago.” would require 16 tokens with the GPT-5 tokenizer, whereas our scheme compresses it to 4 tokens: [Entity_ID], [Action_Type], [Action_Time_Bucket], and [Action_Duration_Bucket].</li><li><strong>Product control.</strong> A direct mapping between tokens and product concepts, such as rows and entities, makes it easier to control what the model can generate. This is crucial for enforcing business rules on the final homepage.</li></ul><h4>Context tokens</h4><p>Context tokens encode user engagement history, user profile, and request context.</p><p>We represent user history as a sequence of user actions. For each action, we extract key metadata, including the action type, entity ID, timestamp, and duration. These actions include both explicit signals, such as play, add to My List, and thumbs-up, and implicit signals, such as trailer views or visits to a details page.</p><p>User profile tokens capture attributes such as language and profile type. Request context tokens encode signals like time of day, day of week, and device.</p><p>Some data sources are too long to include directly as raw token sequences. A user’s full impression history, for example, would be prohibitively expensive to represent in full. In these cases, we use a summarized version. This is a pragmatic trade-off: while GenPage aims to operate on raw inputs as much as possible, handcrafted summaries still introduce a form of prompt engineering into the pipeline. Learning to compress these long data sources end to end is an important direction for future work.</p><p>To help the model distinguish between data sources, we insert special tokens that mark the start of each segment. Continuous signals, such as timestamps and durations, are bucketized into discrete ranges to keep the vocabulary finite.</p><h4>Page tokens</h4><p>Each entity, such as a show, movie, or game, and each row, such as Korean TV Shows, is represented as a single token. The homepage is serialized in layout order: left to right, then top to bottom. We update the entity and row vocabulary daily to incorporate newly added entities and rows. Entities that are still out of vocabulary at serving time are handled through semantic embedding fusion and fallback tokens, both described later.</p><p>In principle, the same paradigm can extend to any output that can be expressed as a linear token sequence. This includes layouts beyond the current two-dimensional structure, such as one-dimensional feeds or mixed layouts, as well as personalized UI components and per-entity outputs such as personalized artwork. We leave these extensions to future work.</p><h4>Paginated recommendation</h4><p>To make recommendations responsive to in-session user preferences, the homepage is often generated incrementally, a few rows at a time. Before each pagination request, we append the page tokens from previously generated rows to the prompt, along with the user’s latest engagements on those rows from Netflix’s real-time event-logging infrastructure. This allows the model to generate the next set of recommendations using both the user’s long-term preferences and their most recent in-session behavior.</p><h3>Reward system</h3><p>To quantify the long-term value of a recommendation, we rely on an internal reward system described in<a href="https://dl.acm.org/doi/10.1145/3604915.3608873"> prior work</a>. The reward system is tuned through online A/B testing to align with long-term user satisfaction and serves as the primary supervision signal for both supervised and reinforcement learning.</p><p>The reward system processes user feedback and assigns a scalar reward for every impressed entity on the homepage. For instance, a TV show binge-watched in one night reflects stronger user satisfaction and receives a higher reward than a movie watched for only 10 minutes. An impressed entity that the user abandons receives a negative reward.</p><p>We define the page-level reward as the sum of rewards across all impressed entities on the homepage.</p><h3>Model architecture</h3><p>GenPage uses a standard decoder-only transformer architecture, the same general architecture behind many modern LLMs. This choice keeps the model simple and flexible, while also letting us benefit from the broad ecosystem of tooling around transformer training and serving.</p><p>One architectural detail is that we untie the input embedding and output projection weights. This is useful because pretraining and post-training place different demands on the logits. Next-token prediction pretraining optimizes a softmax over the vocabulary, while weighted binary classification (WBC) post-training optimizes per-token sigmoid scores, as described below. Untying the weights gives the model more flexibility to adapt to both objectives.</p><h3>Training recipe</h3><p>Our training pipeline mirrors the LLM recipe: we first teach the model the “language” of the Netflix homepage through pretraining, then align its outputs with user satisfaction through post-training. For post-training, we explore two alternative approaches: weighted binary classification (WBC) and reinforcement learning (RL).</p><p>WBC is simpler to optimize and aligns directly with the entity-level objectives of our production ranking models. RL is harder to evaluate and optimize, but it is the key path to GenPage’s full vision of page-level optimization, with the flexibility to incorporate test-time reasoning and multi-token entity representations.</p><h4>Pretraining via next-token prediction</h4><p>We pretrain the model with a standard next-token prediction objective: given the context tokens and a prefix of page tokens, the model learns to predict the next page token. This stage focuses on representation learning, teaching the model the relationship between user contexts and successful homepages. Note that our context-page training examples resemble the prompt-response pairs used in LLM supervised fine-tuning (SFT) more than the raw text used in LLM pretraining. We nonetheless call this stage <em>pretraining</em> because we train the model from scratch rather than fine-tuning from an existing checkpoint.</p><p>Unlike LLMs, which often face a scarcity of high-quality labeled data, recommender systems have an abundance of user feedback. For pretraining, we use homepage impressions that received positive feedback when served in production, bootstrapping the model to generate pages similar to those produced by the existing production system.</p><p>However, pretraining mainly teaches GenPage to imitate the production system. It does not directly optimize the magnitude of the reward, and as GenPage becomes part of production, repeatedly training on pages generated by earlier versions of the model can risk <a href="https://www.nature.com/articles/s41586-024-07566-y">model degeneration</a>. To address these limitations, we explore two post-training approaches.</p><h4>Post-training via weighted binary classification</h4><p>One effective way to align the generative model with user satisfaction is weighted binary classification (WBC). At a high level, WBC turns generation into token-level value prediction: given the user context and the tokens generated so far, the model learns to estimate the value of generating each possible next row or entity token.</p><p>This objective is easier to optimize than page-level RL. By decomposing the homepage into per-token targets, WBC provides token-level credit assignment by construction, rather than requiring RL to infer how each generated decision contributed to the final page-level reward.</p><p>This training setup is enabled by our custom tokenization. Each page token corresponds directly to a specific entity or row, making it straightforward to assign a reward. For every impressed entity on the page, our reward system provides a scalar reward based on user feedback. For each impressed row, we derive a row-level reward by aggregating the rewards of the entities in that row.</p><p>From each reward, we derive a binary label from its sign, such as positive engagement versus abandonment, and a weight from its magnitude, such as binge-watching receiving a higher weight than a short play. We then optimize a weighted binary cross-entropy loss on the logit for the corresponding token. Under this setup, the logit for a token can be interpreted as the model’s value estimate for generating that token at that position.</p><p>Although the model is trained as a value predictor, it can still generate pages autoregressively. At each step, the model scores the candidate next tokens, greedily selects the token with the highest value, and appends it to the prefix. This process repeats token by token until the full homepage is generated.</p><h3>Post-training via reinforcement learning</h3><p>Our second post-training approach is reinforcement learning (RL). WBC is effective for optimizing entity-level metrics, but it does not directly optimize the homepage as a whole. RL treats page generation as a sequential decision-making problem, allowing the model to optimize a page-level reward while preserving the flexibility of autoregressive generation.</p><p>This opens the door to several important capabilities:</p><ul><li><strong>Whole-page optimization.</strong> RL directly optimizes an aggregate page-level reward, allowing the model to account for interactions across rows and entities, such as diversity, stopping power, and page-level business constraints.</li><li><strong>Test-time reasoning.</strong> Analogous to its application in LLMs, RL can optimize reasoning capabilities for generative recommendation. Reasoning outputs can also be viewed as a form of automated feature engineering.</li><li><strong>Multi-token entity support.</strong> In our current tokenization, each entity and row is represented as a single token, so rewards map cleanly to individual tokens. In more complex settings, however, an entity may require multiple tokens, such as [Show_ID] plus [Episode_#] for an episode, or a sequence of <a href="https://arxiv.org/abs/2305.05065">semantic ID</a> tokens. In that case, WBC’s per-token labeling becomes ambiguous because a single entity-level reward must be distributed across multiple tokens. RL avoids this issue by optimizing the sequence-level return, making it a more natural fit for variable-length, multi-token entities.</li></ul><p>Inspired by the<a href="https://arxiv.org/abs/1706.03741"> RLHF</a> recipe used to align large language models, we adopt a two-step approach. First, we train a reward model that predicts the page-level reward for a generated page. This reward model is distinct from the reward system described earlier. The reward system converts <em>observed</em> user feedback into a scalar reward for a page that was actually shown, whereas the reward model <em>predicts</em> the page-level reward for a generated page without showing it to the user. This prediction is what lets RL optimize against arbitrary candidate pages during training.</p><p>Training against a reward model avoids the high variance of off-policy correction on logged or predicted propensities, but introduces the risk of reward hacking. Since the reward model is trained on data generated from the production policy, it is most reliable on pages similar to those the production policy generates. We therefore use a KL penalty to keep the policy close to the pretrained checkpoint, which itself was trained to mimic the production policy. This keeps the pages within the reward model’s region of coverage and limits opportunities for reward hacking.</p><p>For the RL algorithm, we adopt<a href="https://arxiv.org/abs/2503.20783"> Dr. GRPO</a>, a variant of<a href="https://arxiv.org/abs/2501.12948"> GRPO</a> that mitigates biases in the training objective. To train the model within this framework, we need the following components:</p><ul><li><strong>Prompts:</strong> production user requests, represented by context tokens.</li><li><strong>Policy and reference models:</strong> both are initialized from the pretrained checkpoint; the reference model anchors the KL penalty discussed above.</li><li><strong>Reward model:</strong> a dedicated transformer-based reward model, also initialized from the pretrained checkpoint, predicts the page-level outcome reward, using the sum of entity-level rewards from our internal reward system as the supervision target. We also incorporate rule-based format rewards to guide the RL policy. For example, the page should resemble a list of rows, and business-critical rows or entities should not appear too low on the page.</li></ul><h3>Addressing production challenges</h3><h4>Cold start</h4><p>New entities lack the rich interaction data needed to learn robust token embeddings. We address this through two complementary strategies:</p><ul><li><strong>Context injection. </strong>We inject metadata about new or time-sensitive entities (e.g., Live Now events) directly into the context tokens, providing the model with semantic and time-sensitive information.</li><li><strong>Semantic embedding fusion.</strong> Rather than relying solely on entity ID embeddings learned from user interaction data, we represent each entity as a fusion of its ID embedding and a content-based embedding derived from semantic information such as synopses, cast, transcripts, genres, and video content. This fused embedding serves as the input embedding for the entity’s token in the transformer. During training, with small probability, we randomly replace an entity ID token with the generic fallback token (described below), so the model learns to make recommendations from the content-based embedding alone. This ensures that a new entity has a meaningful representation in the same latent space as established entities as soon as its content metadata is available — even before it has any interaction data.</li></ul><h4>Multi-cadence incremental training</h4><p>At Netflix scale, daily retraining of a large transformer from scratch is prohibitively expensive, but recommendation models must remain fresh to capture shifting trends and new catalog additions. We address this with a multi-cadence incremental training strategy (Figure 3).</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*3Ybwr7_r1reMnPQTYIsTuA.png"><figcaption><strong>Figure 3.</strong> Multi-cadence incremental training. Periodic large-scale pretraining and post-training passes run on a broad historical window. Between them, daily incremental updates combine the latest day’s data with a sampled subset of past data to keep the model fresh while avoiding catastrophic forgetting.</figcaption></figure><p>Our training pipeline operates on a cyclic schedule with two distinct rhythms. At a tunable cadence, we conduct a large-scale pretraining and post-training pass on data from a broad historical window. Between these passes, each day we perform an incremental update by continuing post-training from the previous day’s checkpoint, using a mix of the latest day’s data and a sampled subset of past data. This helps the model stay current with new trends and catalog changes while preventing overfitting and <a href="https://arxiv.org/abs/1612.00796">catastrophic forgetting</a>.</p><p>To manage the daily influx of new tokens (e.g., new entities, rows), we employ fallback tokens. New tokens are initialized using fallback tokens of their type (e.g., [Row_Fallback_Token] for new rows, [Entity_Fallback_Token] for new entities). During training, we randomly replace a small percentage of known tokens with fallback tokens, teaching the model to handle unknown tokens gracefully.</p><h4>Enforcing business rules</h4><p>A Netflix homepage must satisfy structural constraints (e.g., organized as a list of rows) as well as product logic such as deduplication, row pinning, and category consistency (e.g., entities in a Comedy row must be comedies). While training signals can encourage rule adherence, they cannot guarantee strict compliance.</p><p>We enforce these rules at inference time through <em>constrained decoding</em>. At each autoregressive generation step, we compute a mask of eligible tokens based on the applicable business rules and apply it to the output logits, allowing only rule-compliant tokens to be generated. This is greatly simplified by our custom tokenization: because each entity and row is a single token, business rules map directly to token-level masks, avoiding the multi-token bookkeeping that constrained decoding requires over a text vocabulary. For example, to pin a specific row (e.g., popular games) at a fixed position (e.g., row position 2), we simply mask out all other tokens at that position.</p><h4>Hybrid row decoding</h4><p>Autoregressive generation ensures that each newly generated token is conditioned on the full preceding context, but generating every entity token one at a time can be expensive. We leverage the structure of the homepage to balance inference efficiency with the amount of contextual information available to each generated token.</p><p>Within each row, the first few entities are especially important: they receive the most user attention and strongly shape the row’s perceived quality and theme. To reduce inference latency, we use a hybrid row decoding strategy. The model autoregressively generates only the first few entities in each row. Conditioned on this generated prefix, we obtain logits for all eligible entities in a single forward pass and select the top-scoring remaining entities, subject to the same inference-time business-rule constraints described above.</p><p>This approach preserves autoregressive conditioning where it matters most while avoiding the latency and cost of decoding long rows token by token.</p><h3>Offline experiments</h3><p>We ran a series of ablations on Netflix internal data to understand how different components of GenPage affect model quality. Because the system was developed iteratively, individual ablations span different training configurations and data snapshots, so we report only relative comparisons within each study. Unless otherwise noted, experiments use ~200M-parameter models and report results on a held-out evaluation set.</p><h4>Does pretraining help?</h4><p>We compare WBC post-training with and without a preceding next-token-prediction pretraining stage. Figure 4 shows that pretraining yields substantial improvements across all metrics.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*o-dD9HXsqX9OQIhtxW4Qhg.png"><figcaption><strong>Figure 4.</strong> Relative improvement from pretraining (versus WBC post-training without a pretraining stage), across loss reduction, row AUC lift, and entity AUC lift. Loss is the weighted binary cross-entropy; Row and Entity AUC are sample-weighted ROC-AUC over row and entity targets.</figcaption></figure><p>The gains may look small in absolute terms, but they are large in our production regime: setting aside the sample weighting, an Entity AUC lift from 0.91 to 0.92 means that for a randomly drawn pair of impressed entities, the model’s misranking rate drops from 9% to 8% — a magnitude of improvement we rarely observe from a single change on a mature production system. Pretraining the model on the “language” of the Netflix homepage provides a strong initialization for post-training, mirroring the pretrain-then-post-train recipe behind modern LLMs.</p><h4>How does performance scale with model size?</h4><p>We sweep model size from ~120M to ~900M parameters (Figure 5) and report the next-token-prediction loss from pretraining and the WBC loss from post-training. Both losses decrease in a power-law-like fashion, mirroring the <a href="https://arxiv.org/abs/2001.08361">scaling trends seen in LLMs</a>. This confirms that the generative approach scales favorably with model size, suggesting that recommendation quality can be further improved by scaling capacity.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*CUwhWQhd3dXu2hoUMLwOyw.png"><figcaption><strong><em>Figure 5. </em></strong><em>Pretraining and WBC post-training losses as model size scales from 120M to 900M parameters. Both decrease in a power-law-like fashion, mirroring LLM scaling trends.</em></figcaption></figure><h4>How does performance scale with information in the user context?</h4><p>Over the course of development, we progressively enriched the prompt, both by adding new data sources to the context and by refining how each source is tokenized. With model size held fixed, the WBC post-training loss decreases substantially as the context is enriched (Figure 6).</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*-oMxiSFxv1I2cv6tOgPljg.png"><figcaption><strong>Figure 6.</strong> WBC post-training loss as we progressively enrich the user context tokens. Loss is normalized to the first step (= 1.0).</figcaption></figure><p>The model-size sweep and the context-enrichment sweep span different axes and are not strictly comparable: the model-size study covers roughly an order of magnitude in parameters, while the context study spans the full trajectory of our prompt design. Even so, the gap between the two is striking. Scaling the model from 120M to 900M parameters reduces WBC loss by roughly 1.3%, whereas the cumulative effect of enriching the context is around 6.9%. In several cases, a single well-designed context addition delivers a larger improvement than the entire ~7.5× model-capacity scaling.</p><p>This suggests that, in our regime, enriching the prompt — both what we put in the context and how we tokenize it — yields a substantially larger improvement than scaling model capacity. Personalization quality appears to be bottlenecked first by the information and representation available to the model, and only then by capacity. We expect context enrichment to dominate until the context is saturated, at which point model capacity becomes the primary driver.</p><h4>Does RL post-training optimize at the page level?</h4><p>In offline evaluations (Figure 7), RL post-training consistently improves the page-level reward over the pretrained checkpoint, but this is largely confirmatory: the reward is computed using the same model the policy is optimizing against. More interestingly, although diversity is not part of the RL objective, homepage diversity — measured via pairwise embedding distance among entities on the page — also increases over the course of training. This suggests that the RL-trained policy is optimizing the page as a whole rather than myopically optimizing each token in isolation.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*QJoEYMljuWOh1wj2mCdI6Q.png"><figcaption><strong>Figure 7.</strong> RL post-training dynamics. Reward and diversity are shown relative to the initial checkpoint (1.0). Reward rises as expected; diversity also rises, despite not being part of the RL objective.</figcaption></figure><h3>Online evaluation</h3><p>We conducted an online A/B test against the current production homepage recommender using GenPage. In this test, GenPage decoded over the existing production row and entity candidate sets, which help handle many business rules (such as eligibility).</p><p>Figure 8 shows the result: all variants delivered statistically significant improvements on the core user engagement metric we use for launch decisions (p &lt; 0.001) against a mature, highly optimized multi-stage production baseline. The variants differed in their training-data configurations; that they all delivered comparable lifts suggests the gain is robust to these design choices rather than dependent on a particular configuration.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/806/1*q9qLegd75SRjC05X15gWww.png"><figcaption><strong><em>Figure 8. </em></strong><em>Daily core user engagement metric over a 14-day online A/B test. The figure shows the average treatment effect of several GenPage variants (differing in training-data configurations) against the production baseline. Shaded regions are 95% confidence intervals. All variants delivered statistically significant improvements over production.</em></figcaption></figure><p>Alongside the engagement wins, we observed unintended shifts in the distribution of impressed entity categories (e.g., new vs. established titles, TV shows vs. movies). These shifts are not necessarily negative, but they are not something we explicitly optimized for, and they warrant deeper investigation. We suspect these shifts reflect GenPage personalizing more precisely than the production stack — consistent with an increase in homepage impression efficiency, i.e., users engaging with what they saw using fewer impressions. This sharper personalization appears to surface production-inherited components (such as the reward system) that aren’t yet aligned with the new generative paradigm. We plan to characterize the drivers of these shifts and, where appropriate, tune these components so the resulting distributions better align with desired product behavior.</p><p>We also observed strong responsiveness to in-session signals: the latest in-session actions quickly influenced subsequent recommendations and faded back to long-term preferences after a day or two, confirming that the model effectively attends to action timestamps. This responsiveness emerges naturally from the generative formulation, without the extensive manual feature engineering used in our production stack.</p><p>Contrary to the common assumption that generative models are slower, GenPage reduced end-to-end serving latency by 20% relative to the baseline. By replacing multiple ranking stages and heavy feature computation with a single transformer operating on raw tokenized inputs, we eliminated substantial serving complexity and computational overhead. Custom tokenization and hybrid row decoding further reduced the number of decoding steps, and thus latency. The 20% reduction was achieved without exhausting the available optimizations; further reductions are possible, and this headroom can be reinvested in capacity or richer prompts.</p><h3>Conclusion</h3><p>We presented GenPage, an early step toward end-to-end generative Netflix homepage construction: representing user context as a tokenized prompt and generating the entire homepage autoregressively in real time. This collapses the traditional multi-stage recommender stack into a single transformer that can be optimized end-to-end.</p><p>In online A/B tests against a mature, highly optimized multi-stage production system, GenPage delivered statistically significant gains on the core user engagement metric we use for launch decisions, while reducing end-to-end serving latency by 20%. Achieving this required adapting the LLM training recipe — pretraining followed by WBC or RL post-training — together with a set of domain-specific techniques: custom tokenization for serving efficiency and product control, context injection and semantic embedding fusion for entity cold start, multi-cadence incremental training for model freshness, constrained decoding for business-rule enforcement, and hybrid row decoding for inference efficiency.</p><p>Two offline findings stand out. First, in our current regime, enriching the prompt yields a substantially larger improvement than scaling model capacity — a takeaway we expect to generalize to other industry-scale personalization settings, at least until the available context is fully exploited. Second, RL post-training increases homepage diversity even though diversity is not part of the objective — an indication that page-level optimization captures interactions across rows and entities.</p><p>Several pieces of the full vision are still in progress: long context still relies on handcrafted summarization, and broader LLM-style capabilities — language, multimodality, and reasoning — have not yet been incorporated. One promising direction here is a hybrid tokenization combining our domain-specific tokens with generic text tokens, retaining structured control while inheriting the strengths of general-purpose LLMs; conceptually, this introduces an additional recommendation modality into an LLM.</p><p>More broadly, we expect many advances from the LLM ecosystem to transfer naturally to this setting, and the boundary between an LLM and a recommender system may increasingly blur. Our results suggest this is a viable path toward simpler recommender systems that align more directly with user satisfaction.</p><h3>Acknowledgments</h3><p>Contributors to this work (in alphabetical order): Abhishek Agrawal, Baolin Li, Casey Stella, Daneo Zhang, Dan Zheng, Donnie DeBoer, Fengdi Che, Fernando Amat Gil, Grace Huang, Inbar Naor, Ishita Verma, Jason Uh, Jimmy Patel, Justin Basilico, Lanxi Huang, Lingyi Liu, Liping Peng, Louis Wang, Michelle Kislak, Nathan Kallus, Nicolas Hortiguera, Paran Jain, Qusai Al-Rabadi, Rein Houthooft, Ryan Lee, Santino Ramos, Scarlet Chen, Shaojing Li, Sheallika Singh, Si Cheng, Wei Wang, and ZQ Zhang.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=77146fba8a08" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/genpage-towards-end-to-end-generative-homepage-construction-at-netflix-77146fba8a08">GenPage: Towards End-to-End Generative Homepage Construction at Netflix</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/genpage-towards-end-to-end-generative-homepage-construction-at-netflix-77146fba8a08</link>
      <guid>https://netflixtechblog.com/genpage-towards-end-to-end-generative-homepage-construction-at-netflix-77146fba8a08</guid>
      <pubDate>Mon, 29 Jun 2026 15:01:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Toward More Controllable AI Video Editing: An Early Research Exploration at Netflix]]></title>
      <description><![CDATA[<p>By <a href="https://www.linkedin.com/in/zhuoning-yuan/">Zhuoning Yuan</a>, <a href="https://www.linkedin.com/in/tim-ta-ying-cheng-411857139/">Ta-Ying Cheng</a>, <a href="https://www.linkedin.com/in/benjamin-klein-usa/">Benjamin Klein</a>, <a href="https://www.linkedin.com/in/bahareh-azarnoush/">Bahareh Azarnoush</a></p><h3>Introduction</h3><p>At Netflix, we build technology to help storytellers bring their creative visions to life and to help members discover the stories they love.</p><p>To connect stories with diverse audiences around the world, we produce promotional assets, including trailers, teasers, and social short‑form videos, that build on and elevate the original footage. Through close collaboration with the teams crafting these assets, we identified a recurring gap in current tools. Transforming raw footage into a polished final asset often requires complex edits like seamlessly adding new visual elements, patching or replacing backgrounds, or removing unwanted objects without breaking the scene’s physical continuity. These tasks typically demand hours of specialized manual editing work. While recent generative video editing models show promise, they often struggle to preserve the integrity of the source footage. Many methods regenerate every pixel to make an edit, which can fail to isolate changes and inadvertently alter elements that should remain untouched. To execute these tasks effectively, artists need tools that empower them to dictate exactly what changes and how it changes.</p><p>Our research goal is to make this process easier for artists. We’re deliberate about where and how AI is applied, ensuring that the technology always serves the creative intent. That principle drives our recent work: exploring the benefits of generative AI in ways that protect and expand creative choice, and keeping artists in precise control of their final vision. Recent advancements in AI video editing have demonstrated impressive capabilities in streamlining complex manual editing workflows, but key challenges remain before they can reliably support professional use:</p><ul><li><strong>Unintended edits:</strong> When editing a specific element in a video clip, many methods regenerate the entire video, which can inadvertently alter identity, performance, and other elements like objects, backgrounds, or critical scene details.</li></ul><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*3_5ISo6KHihlsjf389wi8A.gif"><figcaption>Left: input video. Right: output from Ditto using the prompt “change the background to a winding coastal highway in California,” which completely changes the scene.</figcaption></figure><ul><li><strong>Unnatural physics</strong>: When removing objects, many methods focus only on erasing the target while ignoring the scene’s physical continuity. This can lead to inconsistent motion and implausible interactions, making the results look unnatural.</li></ul><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*_c7QKMqNtpw5WqtOf1tOLA.gif"><figcaption>Left: the green mask denotes the target to be removed. Right: output from Gen-Omnimatte where the target was removed, but the physical continuity of the scene was ignored — the pool float shouldn’t move if there’s no interaction with it.</figcaption></figure><p>Today, we’re sharing two research explorations that aim to address these challenges. We believe this work can help advance the field in a way that’s both meaningful and responsible:</p><ul><li><a href="https://vera-layered-diffusion.github.io/"><strong>Vera</strong></a><strong>: </strong>a layered video diffusion model. Vera generates only what needs to change as separate edit layers while leaving the rest of the video untouched, preserving the identities, performances, and other details from the source footage exactly as filmed.</li><li><a href="https://void-model.github.io/"><strong>VOID</strong></a>: a video inpainting model for video object and interaction deletion. VOID performs physically plausible inpainting in complex scenes: it doesn’t just remove an object, but also reconstructs the scene as if the object was never there.</li></ul><p>Along with this blog post, we’re also publicly releasing the research papers that detail the algorithmic innovations behind <a href="https://drive.google.com/file/d/1CjPHJ71NS9R_HdlVok1fcfc5JcG0UoDa/view?usp=sharing">Vera</a> and <a href="https://arxiv.org/abs/2604.02296">VOID</a>. We hope these publications will enable other researchers to experiment with these ideas, build upon our findings, and further advance the field.</p><h3>Vera: A Layered Video Diffusion Model</h3><p>Existing video editing models regenerate the entire clip, coupling the intended edit with regions that should remain unchanged. This increases the risk of altering details of the original footage. To tackle this challenge, we introduce <a href="https://vera-layered-diffusion.github.io/"><strong>Vera</strong></a>, a novel layered video diffusion framework for content-preserving video editing.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/620/1*GaxGb9nVQiIt_0tFmDVc_A.gif"><figcaption>Teaser for Vera (disclaimer: This is a research prototype, not an official product).</figcaption></figure><h4>Inference Pipeline</h4><p>Given a source video and a text editing instruction, Vera jointly generates an edit layer and an alpha matte. These layers are then seamlessly composed with the original footage to produce the final edited result. By design, Vera supports complex tasks such as object addition and background change, while ensuring that the pixels outside the edited regions from the source video remain perfectly intact.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*1eo2a8owloSpC-ER1CPFIA.png"><figcaption>Inference pipeline for Vera: object addition and background replacement.</figcaption></figure><h4>Training Data</h4><p>One of the main challenges in developing Vera was the lack of suitable training data. Since no public dataset provides the high-quality layered data we need (clean input, alpha matte, edit layer, composite video), we built our own. Using a combination of existing open-source videos and human annotation, we constructed a layered video dataset with a total of 486k frames at 832×480 resolution. We organized it into three subsets of increasing complexity:</p><ul><li><strong>Synthetic Composites</strong>: Clips with high-quality foreground alpha mattes are composited over diverse, automatically generated backgrounds. This subset provides strong and reliable supervision for alpha matting in object addition and background change tasks.</li><li><strong>Realistic Single-Object Videos</strong>: Real-world clips are processed through segmentation, matting, background inpainting/generation, and human quality filtering. This subset increases scene diversity and camera motion, improving composition quality across both tasks.</li><li><strong>Realistic Multi-Object Videos with Effects</strong>: This extends the previous subset by isolating individual objects with curated alpha mattes, including their associated effects such as shadows and reflections. This subset improves compositing and editing in more complex, dynamic scenes.</li></ul><h4>Model Architecture</h4><p>Beyond data, model design is another key challenge. The three target outputs Vera generates — an <strong>edit layer</strong> (decoupled creative edits), an <strong>alpha matte layer </strong>(a grayscale mask that depends on the edit content and scene interactions such as occlusions), and a <strong>composite layer </strong>(natural footage) — have substantially different distributions. In practice, using a single shared architecture to reconcile these differences proved data-inefficient. To address this, Vera uses a <a href="https://arxiv.org/abs/2411.04996">MoT (Mixture-of-Transformers)</a> design. Instead of a single DiT, we use three separate DiTs, one for each output:</p><ul><li>Each DiT maintains its own QKV projections and FFN weights, but we concatenate the output tokens from all three branches and then pass it to joint self-attention. This enables cross-layer interaction while allowing each branch to specialize.</li><li>All three DiTs are initialized from the same pretrained T2V base model. We add two additional patch-embedding layers for the input video and an optional mask video. Source-video tokens are added to the composite tokens, while mask tokens are added to the noisy alpha tokens.</li><li>All layers share the same <a href="https://arxiv.org/abs/2104.09864">RoPE (Rotary Positional Encoding)</a>. We also add zero-initialized learnable embeddings to the alpha and composite tokens to help the model distinguish between layers.</li></ul><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*84Mv_MvSFqFOL2GOk9WXuQ.png"><figcaption>Architecture of Vera compared to other methods. We train two Vera variants: 1.3B and 14B parameters.</figcaption></figure><h4>Evaluations and Results</h4><p>To evaluate Vera, we curated a benchmark of test video-prompt pairs: 72 for object addition and 69 for background change, using open-source videos. The test set spans a range of difficulty, including slow and fast motions, various camera motions, single and multiple objects, and both simple and complex scenes. We evaluated the performance across three complementary dimensions:</p><ul><li><strong>Content Preservation</strong>: Measures whether regions outside the targeted edit remain strictly unaltered, evaluated using pixel-level and perceptual similarity.</li><li><strong>Instruction Compliance</strong>: Measures how faithfully the edited video executes the text prompt.</li><li><strong>Video Quality</strong>: Assesses the temporal coherence and per-frame spatial quality of the final edited video.</li></ul><p>In our results, both Vera-1.3B and Vera-14B significantly outperform existing baselines on content preservation, while maintaining similar video quality and instruction compliance performance compared to strongest baselines (please see the <a href="https://drive.google.com/file/d/1CjPHJ71NS9R_HdlVok1fcfc5JcG0UoDa/view?usp=sharing">paper</a> for full results).</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*yP53Rallx__oNl8CNb0EfQ.gif"><figcaption>Qualitative comparisons between Vera and baselines (please see more examples on Vera’s <a href="https://vera-layered-diffusion.github.io/">project website</a>).</figcaption></figure><p>To complement automated metrics, we ran a human preference study comparing Vera against five baselines. We collaborated with 19 creative reviewers who evaluated 512 video trials in total. In each trial, reviewers were shown randomized side-by-side comparisons between the Vera model and a baseline model. The human consensus strongly aligned with our quantitative findings: Vera-1.3B was preferred over all baselines for content preservation and instruction compliance. Furthermore, reviewers rated Vera’s video quality as comparable to baselines on background change tasks, and noted a clear advantage for Vera on object addition tasks.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*s-8NqiwBFDJlLflkrs2Nww.png"><figcaption>User study on test set: Vera-1.3B vs. five strong baselines.</figcaption></figure><h3>VOID: Video Object and Interaction Deletion</h3><p>Existing video object removal methods excel at inpainting content “behind” the object and correcting appearance-level artifacts such as shadows and reflections. However, when the removed object has more significant interactions — such as collisions with other objects — current models fail to correct them and produce implausible results. To address this, we present <a href="https://void-model.github.io/"><strong>VOID</strong></a>, a video object removal framework designed to perform physically-plausible inpainting in these complex scenarios.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/640/1*LXUj6Wmb4svxXpzp98idHw.gif"><figcaption>Teaser for VOID (disclaimer: This is a research prototype, not an official product).</figcaption></figure><h4>A Two-Pass Inference Pipeline</h4><p>Given an input video, the user clicks on an object to remove. A <strong>VLM-based reasoning pipeline</strong> then analyzes the scene to identify other regions that will be causally affected, e.g., objects that will fall, collide, or change trajectory. This physical reasoning is encoded into a quadmask to guide the diffusion model:</p><ul><li><strong>First Pass:</strong> VOID takes the video and the quadmasks as input and generates a physically plausible counterfactual video in which the object — and its interactions — are removed.</li><li><strong>Second Pass</strong>: Smaller video diffusion models occasionally suffer from “object morphing” when generating moving objects. If VOID detects this failure mode, it triggers a second pass that re-runs inference using <a href="https://arxiv.org/abs/2501.08331">flow-warped noise</a> derived from the first pass, stabilizing the object’s shape along its newly synthesized trajectory.</li></ul><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*mENwjeYzX0Hl3I_5tZublA.jpeg"><figcaption>Overview of VOID’s two-pass inference pipeline.</figcaption></figure><h4>Training Data</h4><p>We built on top of the <a href="https://github.com/google-research/kubric">Kubric simulation engine</a> and the <a href="https://arxiv.org/abs/2504.10414">HUMOTO</a> human motion capture dataset to generate synthetic counterfactual video pairs along with their corresponding quadmasks. Specifically, the counterfactual videos are generated by re-simulating the exact scene from the original video, but with the target object(s) or human removed. This resimulation creates an alternate outcome based on strict laws of physics. For example, if a person holding a lamp is removed from the scene, the simulation ensures the lamp obeys gravity and falls to the ground. The quadmasks then capture the removed object (black), the affected regions (grey), their overlaps (dark grey), and the unchanged parts of the scene (white).</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*zTau8ttT7AN24g2kPwu55Q.jpeg"><figcaption>Overview of VOID data engine.</figcaption></figure><h4>Model Training</h4><p>During model training for VOID, we introduce two improvements over prior work: (i) quadmask conditioning, which explicitly identifies regions in each frame that may change after the object is removed, and (ii) a second-pass video appearance refiner that reduces artifacts such as unwanted object morphing. VOID is finally trained on the <a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.5-5b-InP">CogVideoX-Fun-V1.5–5b-InP</a> backbone with Gen-Omnimatte’s <a href="https://github.com/gen-omnimatte/gen-omnimatte-public?tab=readme-ov-file">checkpoint</a> and fine-tuned for video inpainting with interaction-aware quadmask conditioning.</p><h4>Evaluations and Results</h4><p>Experiments across both synthetic and real data demonstrate that VOID preserves consistent scene dynamics far better than prior video object removal methods (please see the <a href="https://arxiv.org/abs/2604.02296">paper</a> for full results). VOID successfully maintains object structure and produces plausible motion over time across a wide variety of real-world cases. By contrast, results from both open- and closed-source baselines consistently exhibit physically inaccurate artifacts. For instance, baselines generate water splashes without human impact (see top row of the figure below) or show spinning tops being disrupted without the presence of interacting hands.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*UsZjXjSnoOdVfUTv7XDfKA.gif"><figcaption>Comparison of VOID with other strong baselines (please see more examples on VOID’s <a href="https://void-model.github.io/">project website</a>).</figcaption></figure><p>To complement our quantitative evaluation, we conducted a user study with 25 creative reviewers to measure the perceptual realism and physical plausibility of our counterfactual edits. Each participant was randomly assigned 5 out of 75 real-world scenarios, resulting in 125 total comparisons. For each video, participants viewed the original input alongside the outputs of VOID and six baselines (seven models total) in a randomized order. Participants were asked to select the video that best reflected how the scene should realistically appear after the object was removed, factoring in visual quality, temporal consistency, blending, the realism of scene evolution, and the absence of artifacts. VOID was selected 64.8% of the time, substantially outperforming all baseline models.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*TqojcY66f8yQfIbt8cnvzw.png"><figcaption>User study on real-world test examples: VOID vs. six baselines.</figcaption></figure><h3>Looking Ahead</h3><p>Applying AI in ways that serve both member and creator needs is core to our research philosophy, and these projects reflect that approach. While Vera and VOID show promising early results, reaching production-ready quality will require addressing several limitations we encountered. For example, Vera struggles with some complex effects such as lightning or smoke due to the limited training data, and in some cases, it fails to keep background motion fully consistent with the input camera movement. Despite the various generalization capabilities VOID exhibits, we still observe domain gaps. For instance, it cannot handle videos with unusual camera angles or shots captured very close to the target object, and it currently has constraints on supported video length and resolution.</p><p>These limitations motivate continued investment in this line of research. Vera and VOID are important early efforts toward making complex video editing more controllable and accessible for artists. For this work, we used publicly available datasets with additional annotation efforts for experiments, and we hope that sharing our research will encourage the broader community to build on these ideas and advance them further.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=eb8160ed60a2" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/toward-more-controllable-ai-video-editing-an-early-research-exploration-at-netflix-eb8160ed60a2">Toward More Controllable AI Video Editing: An Early Research Exploration at Netflix</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/toward-more-controllable-ai-video-editing-an-early-research-exploration-at-netflix-eb8160ed60a2</link>
      <guid>https://netflixtechblog.com/toward-more-controllable-ai-video-editing-an-early-research-exploration-at-netflix-eb8160ed60a2</guid>
      <pubDate>Tue, 23 Jun 2026 02:31:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[How Netflix Simplified Batch Compute with Kueue]]></title>
      <description><![CDATA[<p>By <a href="https://www.linkedin.com/in/alvinbao/">Alvin Bao</a>, <a href="https://www.linkedin.com/in/alex-petrov-vt/">Alex Petrov</a>, <a href="https://www.linkedin.com/in/jennifernlai/">Jennifer Lai</a>, <a href="https://www.linkedin.com/in/aidan-sherr/">Aidan Sherr</a>, and <a href="https://www.linkedin.com/in/samartha/">Samartha Chandrashekar</a></p><p>As a part of the journey to transition Netflix’s compute infrastructure to be more Kubernetes-native, we have leaned into incorporating components from the Kubernetes ecosystem into our container platform <a href="https://medium.com/netflix-techblog/titus-the-netflix-container-management-platform-is-now-open-source-f868c9fb5436">Titus</a>. One example of this is our use of <a href="https://kueue.sigs.k8s.io/">Kueue</a>, a cloud-native job queueing system for batch workloads, which has largely replaced the custom queuing and scheduling logic in our homegrown managed batch solution Compute Managed Batch (CMB). In this post, we’ll give an overview of what motivated the migration, how we migrated millions of batch jobs to use Kueue, and what Kueue allows us to offer as a Compute platform.</p><h3>Brief Overview of CMB and Titus</h3><p>CMB is a managed batch solution that allows users and applications to execute and manage workloads that run to completion. Using a tenant hierarchy, workloads are managed and queued with ordered execution through priorities, and capacity is managed on a per-tenant basis. Workloads that are submitted to CMB are then run on Titus. The features of Titus relevant to CMB are workload federation across multiple cells (Kubernetes clusters) and federated capacity reservations. This means CMB can talk to a single Titus endpoint to get/submit workloads and update capacity reservations without having to worry about the underlying cell/cluster topology.</p><h4>CMB Tenant Hierarchy</h4><p>Tenants provide a grouping mechanism for jobs submitted on behalf of certain organizations, platforms, or applications. Users can create and organize tenants however best suits their organization or use case. For example, an organization may use a single tenant across several applications or a complex hierarchical structure that matches its team and application ownership structure.</p><p>Tenants are associated with a capacity configuration. The capacity configuration defines the amount of compute capacity available to the tenant and provides certain guarantees around isolation from other tenants. The capacity configuration contains weight (used for fair sharing) and resource dimensions.</p><p>There are two types of tenants in CMB:</p><ol><li><strong>Internal Tenants </strong>— meant to facilitate the creation of a tree of tenants. Internal tenants’ children can be both internal and leaf tenants. Internal tenants themselves do not accept work and thus do not have associated queues.</li><li><strong>Leaf Tenants</strong> — can accept work and have queues associated with them. Leaf tenants cannot have any children.</li></ol><p>With regards to capacity configuration, tenants can use 2 types of capacity:</p><p><strong>Reserved Capacity</strong></p><p>For internal tenants, if a user specifies reserved capacity, it is fair-shared across the subtree and usable by the leaf tenants under that internal tenant.</p><p>For leaf tenants, if a user specifies reserved capacity, it partitions capacity within the hierarchy so that other tenants cannot reserve the same resources. Those reserved resources are not shared with any other tenant, ensuring throughput for a given leaf tenant.</p><p><strong>Shared Capacity</strong></p><p>The Compute team maintains a global pool of shared capacity that any tenant can burst into, in addition to its reserved capacity. Reservations are not required to use CMB, so a tenant can run out of shared capacity entirely. The pool is fair-shared across tenants, but in CMB, this applied only at admission: CMB had no preemption, so once a job was admitted, it ran to completion regardless of shifts in fair-share demand.</p><p>Kueue changes the semantics for <strong>both</strong> types of capacity, which the fair sharing and preemption section covers.</p><p>Here is an example of what a tenant hierarchy looks like:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*MfDuB407Rq81AHZEbhARWA.png"></figure><h4>CMB User/Application Workload Submission Flow</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*c54cCrSMUHGfEvB0jIw5XQ.png"></figure><h4>CMB User/Application Tenant Management Flow</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*l1a0_nFSNK4Q3rNAFksfLA.png"></figure><h3>Why Kueue?</h3><p>CMB was created in 2018, before or alongside many of the open-source batch compute offerings available today. Over the years, as the Kubernetes ecosystem has evolved, many of the features that CMB offered or strived to offer have been included in these open source projects e.g., fair sharing, hierarchical tenants, capacity management, priority queuing. In addition, it became increasingly cumbersome to develop new features such as preemption when CMB was so far removed from the underlying Kubernetes cluster.</p><p>The team took a look at what it would take to modernize our batch abstraction and settled on Kueue for the following reasons:</p><ol><li>Unlike other options such as YuniKorn or Volcano, Kueue does not replace pod scheduling by the kube-scheduler, allowing integration with existing Titus scheduling profiles. Replacing Titus scheduler profiles can fragment job placement, potentially harming efficiency.</li><li>Adoption momentum and pace of innovation.</li><li>Kueue supports multi-tenant quota management over heterogeneous hardware.</li><li>Kueue can operate on primitives such as <em>v1.Pod</em> and <em>batch/v1.Job</em>, and also supports higher-level abstractions such as RayJob / RayCluster for future extensibility.</li><li>Kueue has native features that the team would have liked to implement in CMB, such as preemption, all-or-nothing scheduling, topology aware scheduling.</li></ol><h3>Migrating to Kueue</h3><p>This initiative of migrating CMB workloads to Kueue became known as Netflix Batch. The key tenets of our migration were the following:</p><ol><li>Migration should require zero lift for CMB end users and be completely transparent to them</li><li>No regressions in container launch rate and overall max throughput</li><li>Replace CMB queuing and scheduling with Kueue</li></ol><h4>Netflix Batch User/Application Workload Submission Flow</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*GL_sNsqautjh4lnu5XmRww.png"></figure><p>The key difference between the old and new flows is that we defer queuing and scheduling to Kueue, which is enabled in each Kueue-enabled Titus cell. Titus federation routes the job to Kueue cells using our custom Kueue router.</p><h4>Netflix Batch User/Application Tenant Management Flow</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*Z42qiAMxOqyFLmx6UPElTw.png"></figure><p>For us as operators, the migration was as simple as clicking a button on a tenant in our UI (as shown in the example above). This also allows us to easily rollback changes if there were issues.</p><p>Under the hood, this enrollment converts internal tenants to <a href="https://kueue.sigs.k8s.io/docs/concepts/cohort/">Cohorts</a> and leaf tenants to a <a href="https://kueue.sigs.k8s.io/docs/concepts/cluster_queue/">ClusterQueue </a>+ <a href="https://kueue.sigs.k8s.io/docs/concepts/local_queue/">LocalQueue</a>. The capacity configuration on a given tenant is converted into resource flavors and nominal quotas. The architecture for this looks as follows:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*a9TTKuKN1ovhDA00DWS4Zg.png"></figure><h4>Lessons Learned</h4><ol><li><a href="https://kyle.cascade.family/posts/how-to-actually-migrate-complex-systems-in-infrastructure/">Maintaining API parity</a> with the existing system (vs exposing a new API surface) and migrating the underlying components as a first step derisked the project by unstacking bets while also ensuring we didn’t disrupt the customer experience.</li><li>Don’t wait until the end to migrate the most complex use case. We decided early on to migrate our largest and most complex customer first. This allowed us to build confidence that we could later migrate other customers to Netflix Batch without issues, and resulted in the production migration lasting only 4 weeks.</li><li>We had to run Kueue with much higher QPS, Burst, and groupKindConcurrency than the default configuration to meet our throughput needs. This was derisked early on by running load tests in a development environment that mimics Titus.</li></ol><h3>Current State of Kueue at Netflix</h3><p>Kueue is fully rolled out in production, with it managing millions of batch workloads. In the future, we’re looking at options to enroll more of Titus batch workloads into this more managed experience. We have also productionized more fair sharing and preemptions to address better utilization of reserved capacity. In addition, our learnings are being leveraged by other internal teams, including those building Kubernetes-native training infrastructure, to inform their job scheduling and queuing configurations.</p><h4>Fair Sharing and Preemption</h4><p>With Kueue, <a href="https://kueue.sigs.k8s.io/docs/concepts/fair_sharing/#preemption-based-fair-sharing">Preemption-based Fair Sharing</a> allows Netflix Batch to maintain reservation semantics while lending resources to other tenants when those reservations are not in use. In addition, preemption allows Netflix Batch to preempt lower-priority workloads for higher-priority workloads. For our customers, this means that tenants can use more idle capacity from reservations, submit more jobs without the risk of starvation, and have quicker turnaround times for business-critical workloads.</p><p>An example preemption configuration on a ClusterQueue that we would be using is as follows:</p><pre>apiVersion: kueue.x-k8s.io/v1beta2<br>kind: ClusterQueue<br>metadata:<br>  name: "team-a-cq"<br>spec:<br>  preemption:<br>    reclaimWithinCohort: Any<br>    withinClusterQueue: LowerPriority</pre><p>With these features deployed, Compute has seen a significant increase in average resource utilization.</p><h3>Acknowledgement</h3><p>This work would not have been possible without the great work of the entire Compute team at Netflix.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=87860682629c" width="1" height="1" alt=""><hr><p><a href="https://medium.com/netflix-techblog/how-netflix-simplified-batch-compute-with-kueue-87860682629c">How Netflix Simplified Batch Compute with Kueue</a> was originally published in <a href="https://medium.com/netflix-techblog">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://medium.com/netflix-techblog/how-netflix-simplified-batch-compute-with-kueue-87860682629c</link>
      <guid>https://medium.com/netflix-techblog/how-netflix-simplified-batch-compute-with-kueue-87860682629c</guid>
      <pubDate>Mon, 22 Jun 2026 23:35:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[How Netflix Simplified Batch Compute with Kueue]]></title>
      <description><![CDATA[<p>By <a href="https://www.linkedin.com/in/alvinbao/">Alvin Bao</a>, <a href="https://www.linkedin.com/in/alex-petrov-vt/">Alex Petrov</a>, <a href="https://www.linkedin.com/in/jennifernlai/">Jennifer Lai</a>, <a href="https://www.linkedin.com/in/aidan-sherr/">Aidan Sherr</a>, and <a href="https://www.linkedin.com/in/samartha/">Samartha Chandrashekar</a></p><p>As a part of the journey to transition Netflix’s compute infrastructure to be more Kubernetes-native, we have leaned into incorporating components from the Kubernetes ecosystem into our container platform <a href="https://medium.com/netflix-techblog/titus-the-netflix-container-management-platform-is-now-open-source-f868c9fb5436">Titus</a>. One example of this is our use of <a href="https://kueue.sigs.k8s.io/">Kueue</a>, a cloud-native job queueing system for batch workloads, which has largely replaced the custom queuing and scheduling logic in our homegrown managed batch solution Compute Managed Batch (CMB). In this post, we’ll give an overview of what motivated the migration, how we migrated millions of batch jobs to use Kueue, and what Kueue allows us to offer as a Compute platform.</p><h3>Brief Overview of CMB and Titus</h3><p>CMB is a managed batch solution that allows users and applications to execute and manage workloads that run to completion. Using a tenant hierarchy, workloads are managed and queued with ordered execution through priorities, and capacity is managed on a per-tenant basis. Workloads that are submitted to CMB are then run on Titus. The features of Titus relevant to CMB are workload federation across multiple cells (Kubernetes clusters) and federated capacity reservations. This means CMB can talk to a single Titus endpoint to get/submit workloads and update capacity reservations without having to worry about the underlying cell/cluster topology.</p><h4>CMB Tenant Hierarchy</h4><p>Tenants provide a grouping mechanism for jobs submitted on behalf of certain organizations, platforms, or applications. Users can create and organize tenants however best suits their organization or use case. For example, an organization may use a single tenant across several applications or a complex hierarchical structure that matches its team and application ownership structure.</p><p>Tenants are associated with a capacity configuration. The capacity configuration defines the amount of compute capacity available to the tenant and provides certain guarantees around isolation from other tenants. The capacity configuration contains weight (used for fair sharing) and resource dimensions.</p><p>There are two types of tenants in CMB:</p><ol><li><strong>Internal Tenants </strong>— meant to facilitate the creation of a tree of tenants. Internal tenants’ children can be both internal and leaf tenants. Internal tenants themselves do not accept work and thus do not have associated queues.</li><li><strong>Leaf Tenants</strong> — can accept work and have queues associated with them. Leaf tenants cannot have any children.</li></ol><p>With regards to capacity configuration, tenants can use 2 types of capacity:</p><p><strong>Reserved Capacity</strong></p><p>For internal tenants, if a user specifies reserved capacity, it is fair-shared across the subtree and usable by the leaf tenants under that internal tenant.</p><p>For leaf tenants, if a user specifies reserved capacity, it partitions capacity within the hierarchy so that other tenants cannot reserve the same resources. Those reserved resources are not shared with any other tenant, ensuring throughput for a given leaf tenant.</p><p><strong>Shared Capacity</strong></p><p>The Compute team maintains a global pool of shared capacity that any tenant can burst into, in addition to its reserved capacity. Reservations are not required to use CMB, so a tenant can run out of shared capacity entirely. The pool is fair-shared across tenants, but in CMB, this applied only at admission: CMB had no preemption, so once a job was admitted, it ran to completion regardless of shifts in fair-share demand.</p><p>Kueue changes the semantics for <strong>both</strong> types of capacity, which the fair sharing and preemption section covers.</p><p>Here is an example of what a tenant hierarchy looks like:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*MfDuB407Rq81AHZEbhARWA.png"></figure><h4>CMB User/Application Workload Submission Flow</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*c54cCrSMUHGfEvB0jIw5XQ.png"></figure><h4>CMB User/Application Tenant Management Flow</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*l1a0_nFSNK4Q3rNAFksfLA.png"></figure><h3>Why Kueue?</h3><p>CMB was created in 2018, before or alongside many of the open-source batch compute offerings available today. Over the years, as the Kubernetes ecosystem has evolved, many of the features that CMB offered or strived to offer have been included in these open source projects e.g., fair sharing, hierarchical tenants, capacity management, priority queuing. In addition, it became increasingly cumbersome to develop new features such as preemption when CMB was so far removed from the underlying Kubernetes cluster.</p><p>The team took a look at what it would take to modernize our batch abstraction and settled on Kueue for the following reasons:</p><ol><li>Unlike other options such as YuniKorn or Volcano, Kueue does not replace pod scheduling by the kube-scheduler, allowing integration with existing Titus scheduling profiles. Replacing Titus scheduler profiles can fragment job placement, potentially harming efficiency.</li><li>Adoption momentum and pace of innovation.</li><li>Kueue supports multi-tenant quota management over heterogeneous hardware.</li><li>Kueue can operate on primitives such as <em>v1.Pod</em> and <em>batch/v1.Job</em>, and also supports higher-level abstractions such as RayJob / RayCluster for future extensibility.</li><li>Kueue has native features that the team would have liked to implement in CMB, such as preemption, all-or-nothing scheduling, topology aware scheduling.</li></ol><h3>Migrating to Kueue</h3><p>This initiative of migrating CMB workloads to Kueue became known as Netflix Batch. The key tenets of our migration were the following:</p><ol><li>Migration should require zero lift for CMB end users and be completely transparent to them</li><li>No regressions in container launch rate and overall max throughput</li><li>Replace CMB queuing and scheduling with Kueue</li></ol><h4>Netflix Batch User/Application Workload Submission Flow</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*GL_sNsqautjh4lnu5XmRww.png"></figure><p>The key difference between the old and new flows is that we defer queuing and scheduling to Kueue, which is enabled in each Kueue-enabled Titus cell. Titus federation routes the job to Kueue cells using our custom Kueue router.</p><h4>Netflix Batch User/Application Tenant Management Flow</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*Z42qiAMxOqyFLmx6UPElTw.png"></figure><p>For us as operators, the migration was as simple as clicking a button on a tenant in our UI (as shown in the example above). This also allows us to easily rollback changes if there were issues.</p><p>Under the hood, this enrollment converts internal tenants to <a href="https://kueue.sigs.k8s.io/docs/concepts/cohort/">Cohorts</a> and leaf tenants to a <a href="https://kueue.sigs.k8s.io/docs/concepts/cluster_queue/">ClusterQueue </a>+ <a href="https://kueue.sigs.k8s.io/docs/concepts/local_queue/">LocalQueue</a>. The capacity configuration on a given tenant is converted into resource flavors and nominal quotas. The architecture for this looks as follows:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*a9TTKuKN1ovhDA00DWS4Zg.png"></figure><h4>Lessons Learned</h4><ol><li><a href="https://kyle.cascade.family/posts/how-to-actually-migrate-complex-systems-in-infrastructure/">Maintaining API parity</a> with the existing system (vs exposing a new API surface) and migrating the underlying components as a first step derisked the project by unstacking bets while also ensuring we didn’t disrupt the customer experience.</li><li>Don’t wait until the end to migrate the most complex use case. We decided early on to migrate our largest and most complex customer first. This allowed us to build confidence that we could later migrate other customers to Netflix Batch without issues, and resulted in the production migration lasting only 4 weeks.</li><li>We had to run Kueue with much higher QPS, Burst, and groupKindConcurrency than the default configuration to meet our throughput needs. This was derisked early on by running load tests in a development environment that mimics Titus.</li></ol><h3>Current State of Kueue at Netflix</h3><p>Kueue is fully rolled out in production, with it managing millions of batch workloads. In the future, we’re looking at options to enroll more of Titus batch workloads into this more managed experience. We have also productionized more fair sharing and preemptions to address better utilization of reserved capacity. In addition, our learnings are being leveraged by other internal teams, including those building Kubernetes-native training infrastructure, to inform their job scheduling and queuing configurations.</p><h4>Fair Sharing and Preemption</h4><p>With Kueue, <a href="https://kueue.sigs.k8s.io/docs/concepts/fair_sharing/#preemption-based-fair-sharing">Preemption-based Fair Sharing</a> allows Netflix Batch to maintain reservation semantics while lending resources to other tenants when those reservations are not in use. In addition, preemption allows Netflix Batch to preempt lower-priority workloads for higher-priority workloads. For our customers, this means that tenants can use more idle capacity from reservations, submit more jobs without the risk of starvation, and have quicker turnaround times for business-critical workloads.</p><p>An example preemption configuration on a ClusterQueue that we would be using is as follows:</p><pre>apiVersion: kueue.x-k8s.io/v1beta2<br>kind: ClusterQueue<br>metadata:<br>  name: "team-a-cq"<br>spec:<br>  preemption:<br>    reclaimWithinCohort: Any<br>    withinClusterQueue: LowerPriority</pre><p>With these features deployed, Compute has seen a significant increase in average resource utilization.</p><h3>Acknowledgement</h3><p>This work would not have been possible without the great work of the entire Compute team at Netflix.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=87860682629c" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/how-netflix-simplified-batch-compute-with-kueue-87860682629c">How Netflix Simplified Batch Compute with Kueue</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/how-netflix-simplified-batch-compute-with-kueue-87860682629c</link>
      <guid>https://netflixtechblog.com/how-netflix-simplified-batch-compute-with-kueue-87860682629c</guid>
      <pubDate>Mon, 22 Jun 2026 23:35:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[The Data Canary: How Netflix Validates Catalog Metadata]]></title>
      <description><![CDATA[<p>By <a href="https://www.linkedin.com/in/celina-amados/">Celina Amados</a></p><p><em>At Netflix, our catalog metadata is crucial to our member experience, and a single corrupted data state can impact millions of viewers immediately. To protect streaming reliability, we built an automated data canary system that validates data transformations using production traffic. This canary detects issues in under 10 minutes, and blocks bad data from reaching our members.</em></p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*SGzOXwhed2NX9LxcDNFhig.jpeg"></figure><h3>Intro</h3><p>Catalog metadata is what makes Netflix functional. It defines what titles exist, where they’re available, whether they can be played, and more. This data gets transformed and distributed across our vast infrastructure near-continuously, powering everything that helps members find what they want to watch. Accurate catalog data delivers moments of joy. Corrupted catalog data breaks streaming.</p><h3>What Went Wrong</h3><p>A production incident revealed a critical gap in our resilience strategy. No code had been deployed. No configuration had changed. But, a manual mitigation action taken during a previous incident had inadvertently corrupted a data feed, rendering it empty for a subset of titles.</p><p>The impact was immediate: missing metadata prevented manifest generation, causing failures in our catalog service and playback issues.</p><p>Engineers were alerted immediately, but identifying the root cause took time. After intense triaging, responders pinpointed the corrupted data feed and pinned services back to a known-good state, restoring playback.</p><p>The problem? <strong>Our sophisticated code canary deployments had caught nothing.</strong> No code had changed — the data had.</p><p>This incident exposed a fundamental gap in our resiliency capabilities: we can validate code deployments, but we had no equivalent for our high-velocity data pipelines. Our catalog metadata, consisting of titles, artwork, availability, and more, was continuously transformed from multiple upstream sources and published at a regular cadence. Each upstream source had its own validation, but these checks didn’t catch corruption in the final transformed output.</p><p><strong>We needed to treat data deployments with the same rigor as code deployments.</strong></p><h3>The Challenge: Validating Data at Short Intervals</h3><p>Our catalog metadata service operates as a high-velocity data pipeline: it processes multiple input feeds, transforms them, and publishes the final catalog state that gets distributed across our infrastructure.</p><p>This creates unique validation challenges that our traditional canary analysis tools aren’t designed to handle:</p><p><strong>Time Constraints</strong>: Our existing canary analysis tools require 30–60 minutes to reach statistical confidence. We had a much shorter window between data cycles; we needed to detect issues, make a decision, and block publishing all within a single cycle.</p><p><strong>Emergent Issues</strong>: While each upstream data source has independent validation, problems often only manifest in the final transformed state. We needed to validate the actual output that clients would consume, not just the inputs, as close to the clients as possible.</p><p><strong>Production Traffic is Essential</strong>: We initially considered shadow traffic, but quickly realized it was insufficient. Shadow traffic can only replay requests to our catalog metadata service; it can’t simulate the entire playback lifecycle across multiple services and domains. To detect real customer impact, we needed real production traffic.</p><p><strong>Limit Blast Radius</strong>: Despite using production traffic for validation, we couldn’t allow customers to experience widespread issues during the validation process. Any regression needed to be detected and contained immediately.</p><h3>Our Solution: The Data Canary Orchestrator Pattern</h3><p>After evaluating several architectural approaches, we developed a solution built around three key innovations:</p><h4>1. Dedicated Orchestrator Pattern</h4><p>We created a dedicated cluster for the purposes of canarying new catalog metadata that separates concerns, avoids self-testing, and provides a pattern for extensibility. Here’s how it works:</p><p><strong>Orchestrator Instance:</strong> A dedicated orchestrator instance of our catalog metadata service coordinates the data canary flow. When a new catalog version is published to the canary environment, the orchestrator validates that both baseline and canary clusters are healthy and version-synchronized, then triggers a chaos experiment.</p><p><strong>Permanent Baseline &amp; Canary Clusters</strong>: Two dedicated service clusters run continuously in our canary region. The baseline cluster always serves the latest production catalog version, while the canary cluster receives new versions for validation.</p><p><strong>Generic Integration Point</strong>: Upon chaos experiment completion, the orchestrator reports results back to the transformer service via a REST endpoint. This generic interface means new data sources can implement their own orchestrator patterns without requiring transformer code changes.</p><p>This pattern can now be adopted by other teams at Netflix for validating different data sources, which is exactly the kind of extensibility we designed for.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*cmhWfzmnIH6VJdaiVIinCA.png"><figcaption>Data Canary workflow</figcaption></figure><h4>2. Utilizing and Extending our Chaos Platform</h4><p>Meeting the 10-minute constraint required not only leaning on our <a href="https://netflixtechblog.com/chap-chaos-automation-platform-53e6d528371f">chaos platform</a>, but also extending it to meet our needs:</p><p><strong>Custom Threshold Tuning</strong>: We worked with our Resilience team to customize experiment thresholds for our use case. Standard chaos experiment thresholds were too conservative for our time constraints.</p><p><strong>Multi-Tenant Testing</strong>: Our catalog service supports multiple client types with different traffic patterns and downstream dependencies. We ran separate experiments for major client types and discovered that running traffic through the tenant that handles playback requests consistently identified failures fastest.</p><p><strong>Sticky Canaries: </strong>To isolate experiment traffic, sticky canaries use session affinity to guarantee that once a user’s traffic is routed to the baseline or canary clusters, it stays there for the duration of the experiment window. This prevents cross-contamination from concurrent chaos experiments, ensuring a clean apples-to-apples comparison between data versions.</p><p><strong>Behavioral Metrics Over Technical Metrics</strong>: We focused on <a href="https://netflixtechblog.com/sps-the-pulse-of-netflix-streaming-ae4db0e05f8a">Starts Per Second (SPS)</a>, or actual customer playback attempts, as our primary signal. SPS proved more reliable than latency or error rates for detecting catalog corruption because it directly measures customer impact, and data errors may not always manifest as application errors to our catalog metadata service.</p><p><strong>Immediate Abort on Regression</strong>: Instead of collecting data for post-hoc analysis, we stream metrics in real-time and abort experiments the moment we detect regression. This trades some statistical confidence for speed, but our tight thresholds and clear signal make this not only acceptable, but necessary.</p><h4>3. Production-Hardened Edge Case Handling</h4><p>Building a system that runs in production every 10 minutes taught us that the devil is in the details:</p><p><strong>In-Flight Experiments During Redeployment</strong>: When the orchestrator restarts, it must detect and continue polling any ongoing experiments, as we can’t abandon a validation cycle mid-flight.</p><p><strong>Leader Election</strong>: During orchestrator deployments, multiple instances might be running simultaneously. We implemented safeguards to ensure only one experiment is triggered per version announcement.</p><p><strong>Version Synchronization</strong>: In a multi-tenant service where different clients consume data at different cadences, we track version state to ensure baseline and canary clusters are properly aligned before triggering experiments.</p><h3>Validating the Validator: Controlled Failure Injection</h3><p>To prove the system worked, we needed to break things on purpose. We ran a series of controlled experiments where we deliberately corrupted catalog data — denylisting high-profile titles and simulating real data corruption scenarios — to validate that the canary could detect issues and block publication.</p><p>These experiments were coordinated as <a href="https://netflixtechblog.com/empowering-netflix-engineers-with-incident-management-ebb967871de4">proactive incidents</a> during business hours, with product operations teams on standby. We routed approximately 0.2% of global traffic through the validation flow, minimizing blast radius while still generating meaningful signal.</p><p><strong>Key Results:</strong></p><ul><li><strong>Detection Speed</strong>: Issues identified in 2.5–4 minutes depending on client type</li><li><strong>Clear Signal</strong>: 10x error differential between canary and baseline</li><li><strong>Automatic Blocking</strong>: Publishing workflow blocked as designed when regressions detected</li></ul><p>The experiments validated our end-to-end workflow and revealed important operational insights: different client traffic patterns detect failures at different speeds, and threshold tuning requires careful refinement based on the magnitude of impact we want this system to detect. Most importantly, they proved that even with a 10-minute validation window, far shorter than traditional 30–60 minute canary analysis, we had sufficient signal to catch high-impact catalog corruption.</p><h3>Bringing Code Validation Principles to Data</h3><p>This effort wasn’t just about building a validation system, it was about recognizing that data deployments deserve the same rigor as code deployments. Just because something isn’t a binary doesn’t mean it can’t break production. The patterns we landed on aren’t specific to catalog metadata, and can be applied to systems with high-velocity data pipelines more broadly.</p><p>If you’re working with data that changes frequently and impacts customers directly, ask yourself:</p><ul><li>What’s your MTTD for data corruption?</li><li>Can you validate with production traffic safely?</li><li>How would you detect emergent issues in transformed data?</li><li>What behavioral metric most closely indicates customer impact in your domain?</li></ul><p>Today, the failure mode that caused the aforementioned incident would be caught and mitigated in under 10 minutes. We all know outages aren’t a question of if, but when. The next time you find yourself faced with bad data, how fast will you be able to respond?</p><p><strong>Acknowledgments</strong></p><p>This work was a collaborative effort across multiple teams at Netflix. Special thanks to <a href="https://www.linkedin.com/in/jongyoon-lee/">Jongyoon Lee</a>, <a href="https://www.linkedin.com/in/davesu/">David Su</a>, and <a href="https://www.linkedin.com/in/zubeen/">Zubeen Lalani</a> of the Catalog Foundations &amp; Distribution team for their contributions to the design, and to <a href="https://www.linkedin.com/in/plsek/">Ales Plsek</a> of the Resilience team for their support in customizing our chaos platform for our unique use case.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=18b699d58e36" width="1" height="1" alt=""><hr><p><a href="https://medium.com/netflix-techblog/the-data-canary-how-netflix-validates-catalog-metadata-18b699d58e36">The Data Canary: How Netflix Validates Catalog Metadata</a> was originally published in <a href="https://medium.com/netflix-techblog">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://medium.com/netflix-techblog/the-data-canary-how-netflix-validates-catalog-metadata-18b699d58e36</link>
      <guid>https://medium.com/netflix-techblog/the-data-canary-how-netflix-validates-catalog-metadata-18b699d58e36</guid>
      <pubDate>Sat, 20 Jun 2026 01:54:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Data Projects: Managing Data Assets at Netflix Scale]]></title>
      <description><![CDATA[<h4><em>By </em><a href="https://www.linkedin.com/in/amer-hesson-0886a5a5/"><em>Amer Hesson</em></a><em>, </em><a href="https://www.linkedin.com/in/mayworm/"><em>Marcelo Mayworm</em></a><em>, </em><a href="https://www.linkedin.com/in/james-mulcahy-10493518/"><em>James Mulcahy</em></a><em>, and </em><a href="https://www.linkedin.com/in/brittany-truong-a35b54bb/"><em>Brittany Truong</em></a></h4><h3>The Problem: Managing Assets at Netflix Scale</h3><p>Netflix’s Data Platform is vast. We have millions of tables in our data warehouse and tens of thousands of scheduled workloads running across our orchestration systems. Behind each of these assets sits an engineer, a team, or an initiative — and behind each of those sits a set of decisions about <em>who</em> can access <em>what</em>, and <em>how</em> those workloads execute day after day.</p><p>For years, the tools we used to manage access and identity for these assets operated at the granularity of the individual asset. Every table had its own Access Control List (ACL). Every workflow ran under the identity of the engineer who authored it. In a workforce that is fluid, where people change teams, change roles, and occasionally leave the company, this fine-grained model broke down in two persistent, painful ways.</p><h3>Problem 1: Permissions that can’t keep up with organizational changes</h3><p>Imagine you’re on a team that owns a few hundred tables. Your org restructures, a neighboring team merges into yours, and you inherit another few hundred. Now you have to find every ACL on every table, figure out who should still have access, and update them one by one. Multiply that by every reorg across every team across the company. The result? Two failure modes:</p><ol><li><strong>The support team gets flooded.</strong> A significant and outsized share of support threads were requests to update table permissions en masse in response to org changes. While self-service tooling and best practices are in place to manage this, adherence is inconsistent. Data Projects addresses this by promoting the solution from optional tooling to a foundational part of the data platform.</li><li><strong>Access gets granted far too broadly.</strong> Rather than maintain fine-grained ACLs, teams would often open up table access to the whole company. This defeated the purpose of having ACLs in the first place.</li></ol><h3>Problem 2: Workloads tied to human identities</h3><p>Scheduled and asynchronous workloads — <a href="https://netflixtechblog.com/maestro-netflixs-workflow-orchestrator-ee13a06f9c78">Maestro</a> workflows, data movement jobs, Spark pipelines — need an identity to run as. Historically, that was a <em>human</em>: whoever authored the workflow.</p><p>Human identities are not durable. People change teams, get new responsibilities, and leave the company. When they do, their permissions change, and the workflows running under their identity start to fail. The only fix was to swap in a colleague’s identity, which inevitably had <em>different</em> permissions, kicking off a “permissions whack-a-mole” as each fix surfaced the next missing grant. And then, eventually, that colleague would also move on, and the cycle would repeat.</p><h3>Enter Data Projects</h3><p>We introduced Data Projects to tackle both problems head-on. At its core, a Data Project is two things:</p><ol><li><strong>A container to manage and view a set of related assets in aggregate</strong>: tables, workflows, and other data assets grouped under a single logical umbrella.</li><li><strong>A synthetic, durable, and assumable identity</strong>: one that asynchronous and scheduled workloads can execute under, independent of any human’s lifecycle.</li></ol><p>You can think of it as hoisting the granularity of management up from the individual asset to a meaningful container: the <em>project</em>. Instead of managing permissions on 500 tables, you manage them on one project that contains those 500 tables.</p><p>While the initial focus has been access and identity, the abstraction has applications well beyond those concerns. That broader potential is part of what makes it worth investing in.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*Te9vjGxhHK7jMmpSO6sy2Q.png"><figcaption><strong>Figure 1a</strong>. Individual assets, each managed in isolation, with per-asset access controls and per-person ownership.</figcaption></figure><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*u6MJRtRFedXQ0w3QCIbqBA.png"><figcaption><strong>Figure 1b</strong>. These assets are logically grouped into projects for easier management.</figcaption></figure><h3>Grants and Roles</h3><p>Each Data Project has a set of grants managed by the owning team. Different identity types can be added as grants: users, groups, applications, and continuous integration (CI) jobs. Each grant has a role that determines what the grantee can do within the project. For example, a <em>Contributor</em> has read/write access to the project’s assets, while a <em>Viewer</em> has read-only access. These roles roll up neatly — instead of rewriting hundreds of ACLs when someone joins or leaves a team, you update a single project grant.</p><h3>The Identity Umbrella: Netflix and IAM</h3><p>Every Data Project is provisioned with a Netflix application identity, and optionally an AWS IAM role. This is the “identity umbrella” that makes workloads durable:</p><ul><li>The project’s <strong>Netflix identity</strong> is what executes the project’s async workloads (e.g. Maestro workflows). It belongs to the project, not to any person.</li><li>The project’s <strong>IAM role</strong> supports specialized use cases in AWS like Spark jobs on Amazon EMR. Crucially, the IAM role can be exchanged for the project’s Netflix identity in a cryptographically secure way.</li></ul><p>Members with privileged roles can also assume the project’s Netflix identity. This is enormously useful for testing and troubleshooting from a development context like a laptop or a notebook — you get to run commands <em>as the project</em>, exactly as the scheduled workload would.</p><h3>Gravity</h3><p>One of the more elegant properties of Data Projects is what we call <em>gravity</em>. When a workload running under a project’s identity creates a new asset — say a Maestro workflow creates three tables — those assets are automatically added to the project as contained assets. The project becomes the center of mass for everything produced under its identity. You get organization for free as a side effect of how the platform already works, eliminating future challenges of discovering relevant assets and gaining access to them.</p><h3>Securing Data Workflows with Data Projects</h3><p>Maestro is Netflix’s primary workflow orchestrator for batch analytics, covering scheduled ETL pipelines, data movement jobs, ML training, and much more. Because workflows can run on schedules without the original user present, Maestro is designated a Trusted Workload Manager (TWM), formally authorized to mint fresh identity tokens on behalf of the workloads it manages.</p><p>That identity matters everywhere. A single workflow execution may be checked against table ACLs in the Secure Data Warehouse, authorization policies for Netflix resources, and IAM policies for AWS — all in a single run. If the identity is fragile, the whole workflow is fragile.</p><h3>The Problem with User-Tied Identity</h3><p>The standard pattern was to run workflows under an On-Behalf-Of (OBO) credential — for example, <em>maestro</em> OBO <em>alice@netflix.com</em>. This gave the workflow the union of Maestro’s and the human’s permissions, but in doing so it also bound the workflow’s permissions to that person’s. When they changed teams or left Netflix, the workflow broke. A colleague might take over ownership, but they rarely had the same access as the previous owner, so the workflow would stay broken for days while permissions were sorted out. At Netflix’s scale, with tens of thousands of scheduled workloads, many of them business-critical, this was unsustainable.</p><h3>Data Projects: Durable Identity</h3><p>Data Projects solves this by replacing user-tied identity with a durable, team-owned Netflix application identity: one that doesn’t change teams, go on vacation, or leave the company. Each project groups related workflows, tables, secrets, and other assets under a single consistent identity, and Maestro validates the caller’s access to the project before executing any workflow under it.</p><p>The downstream improvements are as follows:</p><ul><li>Tables created during execution are automatically associated with the project’s identity through <em>gravity</em>, inheriting its access controls without additional configuration.</li><li>Secrets are scoped to project policies, so ownership transfers no longer strand credentials.</li><li>Access is managed once at the project level, replacing fragmented per-user grants across every asset the workflow touches.</li></ul><p>The result is a workflow identity model that is stable, auditable, and built to survive the organizational changes inevitable at any company operating at this scale.</p><h3>Success Stories</h3><p>Many Data Projects have already grown to contain tens of thousands of assets in production. A couple examples are highlighted below:</p><ul><li><strong>Streaming Quality of Experience</strong>: A core observability pipeline tracking quality of experience (QoE) metrics whose continuity used to depend on whichever engineer happened to own the underlying workflows. Now it runs under the project’s identity, stable regardless of team membership changes.</li><li><strong>Member Analytics</strong>: Analytical models and ETL workflows for member data products. A concentrated set of business-critical analytics whose access is managed at the project level rather than across hundreds of individual tables and workflows.</li></ul><p>More broadly, we’ve seen Data Projects adopted as the organizing principle for entire analytics domains. Where teams previously maintained their own access policies, ad-hoc grant lists, and tribal knowledge about “who should have access to what,” the project is now the single answer.</p><h3>Using Data Projects</h3><p>Onboarding workflows onto Data Projects is a matter of:</p><ol><li>Creating a project for the logical grouping of assets (or using an existing suitable one).</li><li>Granting the right people and groups the appropriate roles.</li><li>Configuring the workflow to run with the project’s identity.</li></ol><p>Thanks to gravity, new assets produced by project workflows land in the project automatically. Migrating existing workflows can be a challenge as it requires setting up the Data Project with the appropriate permissions before changing its execution identity. We are actively working on infrastructure to track the access patterns of existing workflows so that we can recommend precise permission updates for the destination project. Our goal is to make the Data Project the de facto option for executing any kind of asynchronous workload.</p><h3>What’s Next</h3><p>Data Projects started as an Analytics Platform initiative, a response to specific pains in the data warehouse, but the underlying ideas are not unique to data. We see a potential future where <strong>Projects</strong> (not just <em>Data</em> Projects) are a first-class platform concept spanning data assets, software assets (GitHub repositories, Spinnaker applications, Docker images), and even studio assets (production content, pipelines, and transformations).</p><p>We’re also investing in:</p><ul><li><strong>Rightsizing</strong>: we’re integrating a layer on top of our authorization policies that automatically rightsizes permissions based on actual usage patterns, proactively eliminating unnecessary access and preventing “permission creep”.</li><li><strong>Hoisting beyond access and identity</strong>: the project is a natural unit for surfacing other concerns at the aggregate level — cost attribution, health indicators, and more.</li><li><strong>Ad-hoc use case integrations</strong>: extending project identities beyond scheduled workloads to cover interactive, on-demand actions like running a query through the Data Portal.</li><li><strong>Activity logs and audits</strong>: a unified timeline of grant changes, asset changes, and workflow versions at the project level.</li></ul><h3>Conclusion</h3><p>Data Projects is an answer to a simple observation: at Netflix’s scale, the unit of identity and access management can’t be the individual asset or the individual human. It has to be something larger, something durable, something that matches the way teams actually think about the work they own.</p><p>A project is that unit. And as we continue to generalize the concept beyond the data warehouse, we expect it to become one of the foundational primitives of how engineering at Netflix is organized, not just how data is organized.</p><h3>Acknowledgments</h3><p>We would like to express our gratitude to the following individuals for their contributions to this effort: Ryan Bordo, Doug Clark, Luke Fernandez, Sarrah Figueroa, Ankit Gupta, Brian Hoying, Ye Ji, Abhishek Kapatkar, Anmol Khurana, Matheus Leão, Hechao Li, Raymond Liu, Alice Naghshineh, David Noor, Anjali Norwood, Javier Garcia Palacios, Kunaal Parekh, Brandon Quan, Andrew Seier, Jason Seo, and Ethan Zhang.</p><p>If you are interested in helping us solve these types of problems and helping entertain the world, please take a look at some of our open positions on the <a href="https://jobs.netflix.com/">Netflix jobs page</a>.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=7ca25888591e" width="1" height="1" alt=""><hr><p><a href="https://medium.com/netflix-techblog/data-projects-managing-data-assets-at-netflix-scale-7ca25888591e">Data Projects: Managing Data Assets at Netflix Scale</a> was originally published in <a href="https://medium.com/netflix-techblog">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://medium.com/netflix-techblog/data-projects-managing-data-assets-at-netflix-scale-7ca25888591e</link>
      <guid>https://medium.com/netflix-techblog/data-projects-managing-data-assets-at-netflix-scale-7ca25888591e</guid>
      <pubDate>Sat, 20 Jun 2026 01:54:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[The Data Canary: How Netflix Validates Catalog Metadata]]></title>
      <description><![CDATA[<p>By <a href="https://www.linkedin.com/in/celina-amados/">Celina Amados</a></p><p><em>At Netflix, our catalog metadata is crucial to our member experience, and a single corrupted data state can impact millions of viewers immediately. To protect streaming reliability, we built an automated data canary system that validates data transformations using production traffic. This canary detects issues in under 10 minutes, and blocks bad data from reaching our members.</em></p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*SGzOXwhed2NX9LxcDNFhig.jpeg"></figure><h3>Intro</h3><p>Catalog metadata is what makes Netflix functional. It defines what titles exist, where they’re available, whether they can be played, and more. This data gets transformed and distributed across our vast infrastructure near-continuously, powering everything that helps members find what they want to watch. Accurate catalog data delivers moments of joy. Corrupted catalog data breaks streaming.</p><h3>What Went Wrong</h3><p>A production incident revealed a critical gap in our resilience strategy. No code had been deployed. No configuration had changed. But, a manual mitigation action taken during a previous incident had inadvertently corrupted a data feed, rendering it empty for a subset of titles.</p><p>The impact was immediate: missing metadata prevented manifest generation, causing failures in our catalog service and playback issues.</p><p>Engineers were alerted immediately, but identifying the root cause took time. After intense triaging, responders pinpointed the corrupted data feed and pinned services back to a known-good state, restoring playback.</p><p>The problem? <strong>Our sophisticated code canary deployments had caught nothing.</strong> No code had changed — the data had.</p><p>This incident exposed a fundamental gap in our resiliency capabilities: we can validate code deployments, but we had no equivalent for our high-velocity data pipelines. Our catalog metadata, consisting of titles, artwork, availability, and more, was continuously transformed from multiple upstream sources and published at a regular cadence. Each upstream source had its own validation, but these checks didn’t catch corruption in the final transformed output.</p><p><strong>We needed to treat data deployments with the same rigor as code deployments.</strong></p><h3>The Challenge: Validating Data at Short Intervals</h3><p>Our catalog metadata service operates as a high-velocity data pipeline: it processes multiple input feeds, transforms them, and publishes the final catalog state that gets distributed across our infrastructure.</p><p>This creates unique validation challenges that our traditional canary analysis tools aren’t designed to handle:</p><p><strong>Time Constraints</strong>: Our existing canary analysis tools require 30–60 minutes to reach statistical confidence. We had a much shorter window between data cycles; we needed to detect issues, make a decision, and block publishing all within a single cycle.</p><p><strong>Emergent Issues</strong>: While each upstream data source has independent validation, problems often only manifest in the final transformed state. We needed to validate the actual output that clients would consume, not just the inputs, as close to the clients as possible.</p><p><strong>Production Traffic is Essential</strong>: We initially considered shadow traffic, but quickly realized it was insufficient. Shadow traffic can only replay requests to our catalog metadata service; it can’t simulate the entire playback lifecycle across multiple services and domains. To detect real customer impact, we needed real production traffic.</p><p><strong>Limit Blast Radius</strong>: Despite using production traffic for validation, we couldn’t allow customers to experience widespread issues during the validation process. Any regression needed to be detected and contained immediately.</p><h3>Our Solution: The Data Canary Orchestrator Pattern</h3><p>After evaluating several architectural approaches, we developed a solution built around three key innovations:</p><h4>1. Dedicated Orchestrator Pattern</h4><p>We created a dedicated cluster for the purposes of canarying new catalog metadata that separates concerns, avoids self-testing, and provides a pattern for extensibility. Here’s how it works:</p><p><strong>Orchestrator Instance:</strong> A dedicated orchestrator instance of our catalog metadata service coordinates the data canary flow. When a new catalog version is published to the canary environment, the orchestrator validates that both baseline and canary clusters are healthy and version-synchronized, then triggers a chaos experiment.</p><p><strong>Permanent Baseline &amp; Canary Clusters</strong>: Two dedicated service clusters run continuously in our canary region. The baseline cluster always serves the latest production catalog version, while the canary cluster receives new versions for validation.</p><p><strong>Generic Integration Point</strong>: Upon chaos experiment completion, the orchestrator reports results back to the transformer service via a REST endpoint. This generic interface means new data sources can implement their own orchestrator patterns without requiring transformer code changes.</p><p>This pattern can now be adopted by other teams at Netflix for validating different data sources, which is exactly the kind of extensibility we designed for.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*cmhWfzmnIH6VJdaiVIinCA.png"><figcaption>Data Canary workflow</figcaption></figure><h4>2. Utilizing and Extending our Chaos Platform</h4><p>Meeting the 10-minute constraint required not only leaning on our <a href="https://netflixtechblog.com/chap-chaos-automation-platform-53e6d528371f">chaos platform</a>, but also extending it to meet our needs:</p><p><strong>Custom Threshold Tuning</strong>: We worked with our Resilience team to customize experiment thresholds for our use case. Standard chaos experiment thresholds were too conservative for our time constraints.</p><p><strong>Multi-Tenant Testing</strong>: Our catalog service supports multiple client types with different traffic patterns and downstream dependencies. We ran separate experiments for major client types and discovered that running traffic through the tenant that handles playback requests consistently identified failures fastest.</p><p><strong>Sticky Canaries: </strong>To isolate experiment traffic, sticky canaries use session affinity to guarantee that once a user’s traffic is routed to the baseline or canary clusters, it stays there for the duration of the experiment window. This prevents cross-contamination from concurrent chaos experiments, ensuring a clean apples-to-apples comparison between data versions.</p><p><strong>Behavioral Metrics Over Technical Metrics</strong>: We focused on <a href="https://netflixtechblog.com/sps-the-pulse-of-netflix-streaming-ae4db0e05f8a">Starts Per Second (SPS)</a>, or actual customer playback attempts, as our primary signal. SPS proved more reliable than latency or error rates for detecting catalog corruption because it directly measures customer impact, and data errors may not always manifest as application errors to our catalog metadata service.</p><p><strong>Immediate Abort on Regression</strong>: Instead of collecting data for post-hoc analysis, we stream metrics in real-time and abort experiments the moment we detect regression. This trades some statistical confidence for speed, but our tight thresholds and clear signal make this not only acceptable, but necessary.</p><h4>3. Production-Hardened Edge Case Handling</h4><p>Building a system that runs in production every 10 minutes taught us that the devil is in the details:</p><p><strong>In-Flight Experiments During Redeployment</strong>: When the orchestrator restarts, it must detect and continue polling any ongoing experiments, as we can’t abandon a validation cycle mid-flight.</p><p><strong>Leader Election</strong>: During orchestrator deployments, multiple instances might be running simultaneously. We implemented safeguards to ensure only one experiment is triggered per version announcement.</p><p><strong>Version Synchronization</strong>: In a multi-tenant service where different clients consume data at different cadences, we track version state to ensure baseline and canary clusters are properly aligned before triggering experiments.</p><h3>Validating the Validator: Controlled Failure Injection</h3><p>To prove the system worked, we needed to break things on purpose. We ran a series of controlled experiments where we deliberately corrupted catalog data — denylisting high-profile titles and simulating real data corruption scenarios — to validate that the canary could detect issues and block publication.</p><p>These experiments were coordinated as <a href="https://netflixtechblog.com/empowering-netflix-engineers-with-incident-management-ebb967871de4">proactive incidents</a> during business hours, with product operations teams on standby. We routed approximately 0.2% of global traffic through the validation flow, minimizing blast radius while still generating meaningful signal.</p><p><strong>Key Results:</strong></p><ul><li><strong>Detection Speed</strong>: Issues identified in 2.5–4 minutes depending on client type</li><li><strong>Clear Signal</strong>: 10x error differential between canary and baseline</li><li><strong>Automatic Blocking</strong>: Publishing workflow blocked as designed when regressions detected</li></ul><p>The experiments validated our end-to-end workflow and revealed important operational insights: different client traffic patterns detect failures at different speeds, and threshold tuning requires careful refinement based on the magnitude of impact we want this system to detect. Most importantly, they proved that even with a 10-minute validation window, far shorter than traditional 30–60 minute canary analysis, we had sufficient signal to catch high-impact catalog corruption.</p><h3>Bringing Code Validation Principles to Data</h3><p>This effort wasn’t just about building a validation system, it was about recognizing that data deployments deserve the same rigor as code deployments. Just because something isn’t a binary doesn’t mean it can’t break production. The patterns we landed on aren’t specific to catalog metadata, and can be applied to systems with high-velocity data pipelines more broadly.</p><p>If you’re working with data that changes frequently and impacts customers directly, ask yourself:</p><ul><li>What’s your MTTD for data corruption?</li><li>Can you validate with production traffic safely?</li><li>How would you detect emergent issues in transformed data?</li><li>What behavioral metric most closely indicates customer impact in your domain?</li></ul><p>Today, the failure mode that caused the aforementioned incident would be caught and mitigated in under 10 minutes. We all know outages aren’t a question of if, but when. The next time you find yourself faced with bad data, how fast will you be able to respond?</p><p><strong>Acknowledgments</strong></p><p>This work was a collaborative effort across multiple teams at Netflix. Special thanks to <a href="https://www.linkedin.com/in/jongyoon-lee/">Jongyoon Lee</a>, <a href="https://www.linkedin.com/in/davesu/">David Su</a>, and <a href="https://www.linkedin.com/in/zubeen/">Zubeen Lalani</a> of the Catalog Foundations &amp; Distribution team for their contributions to the design, and to <a href="https://www.linkedin.com/in/plsek/">Ales Plsek</a> of the Resilience team for their support in customizing our chaos platform for our unique use case.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=18b699d58e36" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/the-data-canary-how-netflix-validates-catalog-metadata-18b699d58e36">The Data Canary: How Netflix Validates Catalog Metadata</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/the-data-canary-how-netflix-validates-catalog-metadata-18b699d58e36</link>
      <guid>https://netflixtechblog.com/the-data-canary-how-netflix-validates-catalog-metadata-18b699d58e36</guid>
      <pubDate>Sat, 20 Jun 2026 01:54:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Data Projects: Managing Data Assets at Netflix Scale]]></title>
      <description><![CDATA[<h4><em>By </em><a href="https://www.linkedin.com/in/amer-hesson-0886a5a5/"><em>Amer Hesson</em></a><em>, </em><a href="https://www.linkedin.com/in/mayworm/"><em>Marcelo Mayworm</em></a><em>, </em><a href="https://www.linkedin.com/in/james-mulcahy-10493518/"><em>James Mulcahy</em></a><em>, and </em><a href="https://www.linkedin.com/in/brittany-truong-a35b54bb/"><em>Brittany Truong</em></a></h4><h3>The Problem: Managing Assets at Netflix Scale</h3><p>Netflix’s Data Platform is vast. We have millions of tables in our data warehouse and tens of thousands of scheduled workloads running across our orchestration systems. Behind each of these assets sits an engineer, a team, or an initiative — and behind each of those sits a set of decisions about <em>who</em> can access <em>what</em>, and <em>how</em> those workloads execute day after day.</p><p>For years, the tools we used to manage access and identity for these assets operated at the granularity of the individual asset. Every table had its own Access Control List (ACL). Every workflow ran under the identity of the engineer who authored it. In a workforce that is fluid, where people change teams, change roles, and occasionally leave the company, this fine-grained model broke down in two persistent, painful ways.</p><h3>Problem 1: Permissions that can’t keep up with organizational changes</h3><p>Imagine you’re on a team that owns a few hundred tables. Your org restructures, a neighboring team merges into yours, and you inherit another few hundred. Now you have to find every ACL on every table, figure out who should still have access, and update them one by one. Multiply that by every reorg across every team across the company. The result? Two failure modes:</p><ol><li><strong>The support team gets flooded.</strong> A significant and outsized share of support threads were requests to update table permissions en masse in response to org changes. While self-service tooling and best practices are in place to manage this, adherence is inconsistent. Data Projects addresses this by promoting the solution from optional tooling to a foundational part of the data platform.</li><li><strong>Access gets granted far too broadly.</strong> Rather than maintain fine-grained ACLs, teams would often open up table access to the whole company. This defeated the purpose of having ACLs in the first place.</li></ol><h3>Problem 2: Workloads tied to human identities</h3><p>Scheduled and asynchronous workloads — <a href="https://netflixtechblog.com/maestro-netflixs-workflow-orchestrator-ee13a06f9c78">Maestro</a> workflows, data movement jobs, Spark pipelines — need an identity to run as. Historically, that was a <em>human</em>: whoever authored the workflow.</p><p>Human identities are not durable. People change teams, get new responsibilities, and leave the company. When they do, their permissions change, and the workflows running under their identity start to fail. The only fix was to swap in a colleague’s identity, which inevitably had <em>different</em> permissions, kicking off a “permissions whack-a-mole” as each fix surfaced the next missing grant. And then, eventually, that colleague would also move on, and the cycle would repeat.</p><h3>Enter Data Projects</h3><p>We introduced Data Projects to tackle both problems head-on. At its core, a Data Project is two things:</p><ol><li><strong>A container to manage and view a set of related assets in aggregate</strong>: tables, workflows, and other data assets grouped under a single logical umbrella.</li><li><strong>A synthetic, durable, and assumable identity</strong>: one that asynchronous and scheduled workloads can execute under, independent of any human’s lifecycle.</li></ol><p>You can think of it as hoisting the granularity of management up from the individual asset to a meaningful container: the <em>project</em>. Instead of managing permissions on 500 tables, you manage them on one project that contains those 500 tables.</p><p>While the initial focus has been access and identity, the abstraction has applications well beyond those concerns. That broader potential is part of what makes it worth investing in.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*Te9vjGxhHK7jMmpSO6sy2Q.png"><figcaption><strong>Figure 1a</strong>. Individual assets, each managed in isolation, with per-asset access controls and per-person ownership.</figcaption></figure><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*u6MJRtRFedXQ0w3QCIbqBA.png"><figcaption><strong>Figure 1b</strong>. These assets are logically grouped into projects for easier management.</figcaption></figure><h3>Grants and Roles</h3><p>Each Data Project has a set of grants managed by the owning team. Different identity types can be added as grants: users, groups, applications, and continuous integration (CI) jobs. Each grant has a role that determines what the grantee can do within the project. For example, a <em>Contributor</em> has read/write access to the project’s assets, while a <em>Viewer</em> has read-only access. These roles roll up neatly — instead of rewriting hundreds of ACLs when someone joins or leaves a team, you update a single project grant.</p><h3>The Identity Umbrella: Netflix and IAM</h3><p>Every Data Project is provisioned with a Netflix application identity, and optionally an AWS IAM role. This is the “identity umbrella” that makes workloads durable:</p><ul><li>The project’s <strong>Netflix identity</strong> is what executes the project’s async workloads (e.g. Maestro workflows). It belongs to the project, not to any person.</li><li>The project’s <strong>IAM role</strong> supports specialized use cases in AWS like Spark jobs on Amazon EMR. Crucially, the IAM role can be exchanged for the project’s Netflix identity in a cryptographically secure way.</li></ul><p>Members with privileged roles can also assume the project’s Netflix identity. This is enormously useful for testing and troubleshooting from a development context like a laptop or a notebook — you get to run commands <em>as the project</em>, exactly as the scheduled workload would.</p><h3>Gravity</h3><p>One of the more elegant properties of Data Projects is what we call <em>gravity</em>. When a workload running under a project’s identity creates a new asset — say a Maestro workflow creates three tables — those assets are automatically added to the project as contained assets. The project becomes the center of mass for everything produced under its identity. You get organization for free as a side effect of how the platform already works, eliminating future challenges of discovering relevant assets and gaining access to them.</p><h3>Securing Data Workflows with Data Projects</h3><p>Maestro is Netflix’s primary workflow orchestrator for batch analytics, covering scheduled ETL pipelines, data movement jobs, ML training, and much more. Because workflows can run on schedules without the original user present, Maestro is designated a Trusted Workload Manager (TWM), formally authorized to mint fresh identity tokens on behalf of the workloads it manages.</p><p>That identity matters everywhere. A single workflow execution may be checked against table ACLs in the Secure Data Warehouse, authorization policies for Netflix resources, and IAM policies for AWS — all in a single run. If the identity is fragile, the whole workflow is fragile.</p><h3>The Problem with User-Tied Identity</h3><p>The standard pattern was to run workflows under an On-Behalf-Of (OBO) credential — for example, <em>maestro</em> OBO <em>alice@netflix.com</em>. This gave the workflow the union of Maestro’s and the human’s permissions, but in doing so it also bound the workflow’s permissions to that person’s. When they changed teams or left Netflix, the workflow broke. A colleague might take over ownership, but they rarely had the same access as the previous owner, so the workflow would stay broken for days while permissions were sorted out. At Netflix’s scale, with tens of thousands of scheduled workloads, many of them business-critical, this was unsustainable.</p><h3>Data Projects: Durable Identity</h3><p>Data Projects solves this by replacing user-tied identity with a durable, team-owned Netflix application identity: one that doesn’t change teams, go on vacation, or leave the company. Each project groups related workflows, tables, secrets, and other assets under a single consistent identity, and Maestro validates the caller’s access to the project before executing any workflow under it.</p><p>The downstream improvements are as follows:</p><ul><li>Tables created during execution are automatically associated with the project’s identity through <em>gravity</em>, inheriting its access controls without additional configuration.</li><li>Secrets are scoped to project policies, so ownership transfers no longer strand credentials.</li><li>Access is managed once at the project level, replacing fragmented per-user grants across every asset the workflow touches.</li></ul><p>The result is a workflow identity model that is stable, auditable, and built to survive the organizational changes inevitable at any company operating at this scale.</p><h3>Success Stories</h3><p>Many Data Projects have already grown to contain tens of thousands of assets in production. A couple examples are highlighted below:</p><ul><li><strong>Streaming Quality of Experience</strong>: A core observability pipeline tracking quality of experience (QoE) metrics whose continuity used to depend on whichever engineer happened to own the underlying workflows. Now it runs under the project’s identity, stable regardless of team membership changes.</li><li><strong>Member Analytics</strong>: Analytical models and ETL workflows for member data products. A concentrated set of business-critical analytics whose access is managed at the project level rather than across hundreds of individual tables and workflows.</li></ul><p>More broadly, we’ve seen Data Projects adopted as the organizing principle for entire analytics domains. Where teams previously maintained their own access policies, ad-hoc grant lists, and tribal knowledge about “who should have access to what,” the project is now the single answer.</p><h3>Using Data Projects</h3><p>Onboarding workflows onto Data Projects is a matter of:</p><ol><li>Creating a project for the logical grouping of assets (or using an existing suitable one).</li><li>Granting the right people and groups the appropriate roles.</li><li>Configuring the workflow to run with the project’s identity.</li></ol><p>Thanks to gravity, new assets produced by project workflows land in the project automatically. Migrating existing workflows can be a challenge as it requires setting up the Data Project with the appropriate permissions before changing its execution identity. We are actively working on infrastructure to track the access patterns of existing workflows so that we can recommend precise permission updates for the destination project. Our goal is to make the Data Project the de facto option for executing any kind of asynchronous workload.</p><h3>What’s Next</h3><p>Data Projects started as an Analytics Platform initiative, a response to specific pains in the data warehouse, but the underlying ideas are not unique to data. We see a potential future where <strong>Projects</strong> (not just <em>Data</em> Projects) are a first-class platform concept spanning data assets, software assets (GitHub repositories, Spinnaker applications, Docker images), and even studio assets (production content, pipelines, and transformations).</p><p>We’re also investing in:</p><ul><li><strong>Rightsizing</strong>: we’re integrating a layer on top of our authorization policies that automatically rightsizes permissions based on actual usage patterns, proactively eliminating unnecessary access and preventing “permission creep”.</li><li><strong>Hoisting beyond access and identity</strong>: the project is a natural unit for surfacing other concerns at the aggregate level — cost attribution, health indicators, and more.</li><li><strong>Ad-hoc use case integrations</strong>: extending project identities beyond scheduled workloads to cover interactive, on-demand actions like running a query through the Data Portal.</li><li><strong>Activity logs and audits</strong>: a unified timeline of grant changes, asset changes, and workflow versions at the project level.</li></ul><h3>Conclusion</h3><p>Data Projects is an answer to a simple observation: at Netflix’s scale, the unit of identity and access management can’t be the individual asset or the individual human. It has to be something larger, something durable, something that matches the way teams actually think about the work they own.</p><p>A project is that unit. And as we continue to generalize the concept beyond the data warehouse, we expect it to become one of the foundational primitives of how engineering at Netflix is organized, not just how data is organized.</p><h3>Acknowledgments</h3><p>We would like to express our gratitude to the following individuals for their contributions to this effort: Ryan Bordo, Doug Clark, Luke Fernandez, Sarrah Figueroa, Ankit Gupta, Brian Hoying, Ye Ji, Abhishek Kapatkar, Anmol Khurana, Matheus Leão, Hechao Li, Raymond Liu, Alice Naghshineh, David Noor, Anjali Norwood, Javier Garcia Palacios, Kunaal Parekh, Brandon Quan, Andrew Seier, Jason Seo, and Ethan Zhang.</p><p>If you are interested in helping us solve these types of problems and helping entertain the world, please take a look at some of our open positions on the <a href="https://jobs.netflix.com/">Netflix jobs page</a>.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=7ca25888591e" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/data-projects-managing-data-assets-at-netflix-scale-7ca25888591e">Data Projects: Managing Data Assets at Netflix Scale</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/data-projects-managing-data-assets-at-netflix-scale-7ca25888591e</link>
      <guid>https://netflixtechblog.com/data-projects-managing-data-assets-at-netflix-scale-7ca25888591e</guid>
      <pubDate>Sat, 20 Jun 2026 01:54:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Predicting Risk in Content Launches: How Data-Driven Insights can Transform Launch Planning]]></title>
      <description><![CDATA[<p>by <a href="https://www.linkedin.com/in/ecgill/">Emily Gill</a></p><p><em>Each year, we bring the Analytics Engineering community together for an Analytics Summit — a multi-day internal conference to share analytical deliverables across Netflix, discuss analytic practice, and build relationships within the community. This post is one of several topics presented at the Summit highlighting the breadth and impact of Analytics work across different areas of the business.</em></p><h3><strong>Understanding Risk in Content Launches</strong></h3><p>Every title you see on Netflix goes through several key phases: Development, Pre-Production, Production/Principal Photography, Post-Production, and finally, Launch Preparation, all leading up to the Title Launch. Once Principal Photography wraps, the focus shifts in Post-Production from content creation to quality assurance and visual effects (if needed).</p><p>At the end of Post Production, Netflix receives the final audio and video files — often delivered as an IMF (Interoperable Master Format) — which triggers a flurry of Launch Preparation activities, focused on tasks such as the development of artwork and trailers, creation of subtitles, maturity ratings &amp; quality control, that happen within a tight window and rely on having the finalized media assets in hand.</p><p>Some of this work can be kicked off earlier using a non-final version of the media called the Locked Cut, but since it’s not the absolute final deliverable, this presents a tradeoff: should our teams who prepare content for service wait for the more finalized IMF to begin their work, or start sooner with the unfinal Locked Cut? Waiting for the IMF risks a compressed timeline if it arrives late, while starting with the Locked Cut means teams may need to do additional conformance work if there are significant changes between the Locked Cut and the final IMF.</p><h4><strong>Identifying Gaps in Schedule Accuracy</strong></h4><p>To help navigate the decision of when to start launch preparation, our teams rely on estimated delivery dates for both the Locked Cut and IMF media assets, which are manually provided by content partners in production schedules. However, these schedules often have gaps in coverage and lack accuracy for both asset types (see Figure 1).</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*4KfGZI5-3PaxEU_tMjp9NQ.png"><figcaption>Figure 1. At an asset-level we generally see that scheduled date accuracy and coverage are lower at horizons further from asset delivery. As we approach delivery (moving towards the right on this plot) schedules become more accurate (errors decrease) adn coverage improves.</figcaption></figure><p>This isn’t unexpected — productions are dynamic, facing frequent changes, scheduling conflicts, and unforeseen obstacles that can shift timelines without warning. As a result, there’s a clear opportunity to leverage the wealth of production data we collect to predict the risk of schedule slips. By developing a predictive model, we aim to both fill in ETA gaps (providing asset delivery estimates when none exist) and improve the accuracy of existing ETAs compared to traditional manual schedules.</p><h4><strong>Correlation between Schedule Accuracy and Launch Misses</strong></h4><p>Our analysis reveals a strong correlation between scheduled inaccuracies and launch misses — instances where a title experiences delays. To quantify schedule inaccuracy, we created a metric called Accumulated Error Days (AED), which measures the cumulative deviation between estimated (scheduled or predicted) delivery dates and actual delivery dates over time. AED is calculated retrospectively as the area between the scheduled (grey line) or predicted (blue line) delivery dates and the actual delivery date (green line).</p><p>When we compare titles with at least one launch miss to those without, we find that mean AED is significantly higher in the group with launch misses. Notably, this effect is even more pronounced when we focus on the period closer to delivery — indicating that high AED (i.e., inaccurate schedules) in the final stretch before launch is especially correlated with launch misses, more so than AED accumulated over a longer timeline. These findings further motivate our efforts to improve schedule accuracy and reduce AED by leveraging rich production data and predictive modeling.</p><h3><strong>Modeling Time-to-Delivery</strong></h3><p>Our predictive models are designed as boosted tree regression models that predict the “days until” either media asset delivery for in-progress productions.</p><p>To power these models, we leverage a range of upstream data sources including production-level signals of progress, title metadata, and seasonal signals. We are able to predict the days until media asset delivery using daily update snapshots, allowing us to generate up-to-date predictions that reflect the latest state of each in-progress production. This means that we have each feature and what its value was as of each day of a production. Modeling with this snapshotted data enables us to generate up-to-date predictions as new information becomes available, build a flexible model that works across all production phases, and seamlessly incorporate dynamic features that evolve over time (Figure 2).</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*Yv6mwy7QZ8ghiI4S1jGv6g.png"><figcaption>Figure 2. Hypothetical illustration of the evolving nature of production-related signals used in our models. Some signals are present throughout but dynamic, others are present at single moments in time during specific production phases. By capturing data in a snapshotted form, we’re able to build a flexible phase-agnostic model that leverages many different types of progress signals. This figure is illustrative only and does not depict actual Netflix financial or production data.</figcaption></figure><h3><strong>Evaluating Our Approach</strong></h3><h4>Building a Comprehensive Metrics Suite</h4><p>When evaluating the performance of the predictive models, we look across a suite of metrics to try to understand where and when predicted dates outperform scheduled dates. Among these are mean and median absolute error, relative to actual delivery, to understand the accuracy of our estimated dates. We also consider bias metrics, such as mean and median error, to understand if we are consistently over- or under-predicting the actual delivery. We calculate the standard deviation of our errors to understand if there are large shifts in the bulk of the distribution of errors. For the tails of our error distributions, we calculate the percentage of our absolute errors that are greater than x days to delivery.</p><p>For scheduled dates, we calculate coverage across various horizons to delivery. This is a value prop of the model; we’ve built the model in such a way that we can always provide a predicted date and recoup any coverage gaps that exist from scheduled dates alone.</p><h4>Benchmarking Against Manual Scheduling</h4><p>In a backtest, we observed significant improvements across all of our metrics and across most horizons from delivery. As an example, see Figure 3 which plots global mean absolute error (MAE) and shows large reductions in errors (greater accuracy) in predicted IMF and Locked dates as compared to scheduled dates. Additionally, we see large reductions in outliers from scheduled to predicted dates as well.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*sqKvClh5t_Y1dp4ZQU3Xkw.png"><figcaption>Figure 3. This plot compares accuracy (measured as Mean Absolute Error) between predicted and scheduled dates. The horizontal axis plots time prior to delivery, which decreases from left to right until you reach the moment of delivery at the bottom right. For this particular asset, the predicted delivery dates on average are much more accurate than manually scheduled delivery dates throughout the full horizon to delivery.</figcaption></figure><p>Since our teams use these dates over a period of time and not at a single point in time, there is an additional benefit that we’re describing as an Earlier Accuracy Signal. By leveraging predictive dates, our teams benefit from a level of accuracy that they would otherwise have to wait x amount of time for if using scheduled dates. As an example, 6 months out from Locked Cut delivery the predicted dates are better than scheduled dates on 76% of titles and have a level of accuracy (6.1 wks MAE) that scheduled dates don’t reach until 11 weeks later.</p><p>Circling back to AED, which we mentioned earlier is correlated to launch misses, we find that in our backtested titles globally, and across most buying orgs and content types (i.e., series versus standalones), predicted IMF and Locked Cut dates reduce AED from scheduled dates when calculated across the 6 months leading up to delivery. We see similar patterns when we repeat this for shorter horizons to delivery as well.</p><h3>Streamlining Workflows with Improved Scheduling</h3><p>A key advantage of this predictive model is that estimated delivery dates are already integral to our stakeholders’ workflows — meaning we can introduce predictive dates without overhauling existing processes. However, this creates a new challenge: with both scheduled and predicted dates available, teams need to determine which is more reliable. While predictive dates are often more accurate on average, there are situations where scheduled dates perform better. To address this, we’ve built serving logic that defaults to scheduled dates in buying orgs where the model underperforms. Elsewhere, teams can view both dates side by side in dashboards, allowing them to apply their own judgment. Additionally, our predictive models leverage features that are tied to scheduled dates, which has emphasized the need and impact of ensuring our upstream teams continue to input and update scheduled dates even in the presence of our predictions. We’re piloting these predictive signals in multiple ways, tailoring the approach to fit the diverse needs and tools of our various launch prep functions.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=587b1f2de928" width="1" height="1" alt=""><hr><p><a href="https://medium.com/netflix-techblog/predicting-risk-in-content-launches-how-data-driven-insights-can-transform-launch-planning-587b1f2de928">Predicting Risk in Content Launches: How Data-Driven Insights can Transform Launch Planning</a> was originally published in <a href="https://medium.com/netflix-techblog">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://medium.com/netflix-techblog/predicting-risk-in-content-launches-how-data-driven-insights-can-transform-launch-planning-587b1f2de928</link>
      <guid>https://medium.com/netflix-techblog/predicting-risk-in-content-launches-how-data-driven-insights-can-transform-launch-planning-587b1f2de928</guid>
      <pubDate>Sat, 20 Jun 2026 01:53:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[The Evolution of Cassandra Data Movement at Netflix]]></title>
      <description><![CDATA[<p>By <a href="https://www.linkedin.com/in/guilhermesmi/">Guil Pires</a>, <a href="https://www.linkedin.com/in/jenjprince/">Jennifer Prince</a>, <a href="https://www.linkedin.com/in/josecamachof/">Jose Camacho</a>, <a href="https://www.linkedin.com/in/kenkurzweil/">Ken Kurzweil</a>, <a href="https://www.linkedin.com/in/phanindra-chunduru/">Phanindra Chunduru</a></p><h3>Background</h3><p>In a previous post, we introduced <a href="https://netflixtechblog.medium.com/data-bridge-how-netflix-simplifies-data-movement-36d10d91c313">Data Bridge</a>, a unified management plane for batch Data Movement at Netflix. Historically, several bespoke Data Movement connectors were developed across different engineering organizations to fulfill their specific requirements. Over the last few years, the Data Movement team has started centralizing these offerings through an abstraction that provides a catalog of connectors, along with simple UI and APIs to initiate Data Movement jobs.</p><p>One such case is the Cassandra to Iceberg connector. Apache Cassandra powers mission critical applications at Netflix, including Member, Billing, Recommendations, Subscriptions and many more. These use cases heavily leverage Data Movement to Apache Iceberg for many analytics and operational tasks, and central to this movement was a connector for Cassandra to Iceberg built in-house named Casspactor. As many Cassandra based Data Abstractions emerged, such as <a href="https://netflixtechblog.com/introducing-netflixs-key-value-data-abstraction-layer-1ea8a0a11b30">Key Value</a>, <a href="https://netflixtechblog.com/introducing-netflix-timeseries-data-abstraction-layer-31552f6326f8">Time Series</a> and <a href="https://netflixtechblog.medium.com/high-throughput-graph-abstraction-at-netflix-part-i-e88063e6f6d5">Graph</a> — the need for larger and more complex Data Movement with transformations became more critical to the business.</p><p>Data movements are fundamentally fulfilled by leveraging the existing Cassandra backup infrastructure. Regularly scheduled backups are performed directly on the Apache Cassandra nodes, via a sidecar process managing the upload of all necessary SSTables and associated Metadata files directly into Amazon S3. When a Data Movement job is initiated, the job constructs the specific backup structure it needs by referencing the S3 based metadata, allowing it to precisely locate the SSTable files. The engine then downloads these files, performs the required mutation compaction and processing, and finally writes the fully transformed, compacted data directly into the target Apache Iceberg tables.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*DPkj0pP-y0TntplMdm224Q.png"><figcaption>Image 1: Cassandra Cluster Backups to S3</figcaption></figure><h3>Casspactor: The Engine We Outgrew</h3><p>Casspactor processed roughly 1,200 data movements per day, transferring approximately 3 PB of data from Apache Cassandra into Apache Iceberg tables. It served some of the most critical workloads at Netflix. For years, it worked. Then, two compounding challenges made it clear we needed a fundamentally different architecture.</p><h3>Fragile Metadata Dependencies</h3><p>Before Casspactor could move a single record, it needed to answer a deceptively simple question: <em>which backup exists, is it complete, and what does it contain?</em></p><p>Casspactor assembled this answer from multiple independent systems:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*ppVlrVcA2a3zOu0R6PtuUA.png"><figcaption>Image 2: Casspactor’s Composite View of a Backup</figcaption></figure><p>Each system had its own failure modes, update cadences, and accuracy guarantees. Casspactor’s view of the world was a composite, and composites diverge from reality.</p><p>Metadata fell out of sync with actual backups, causing Casspactor to read stale or incorrect data silently. Routine maintenance on the Cassandra Clusters triggered uncoordinated snapshots, and because Casspactor required all nodes in a region to snapshot at the same clock second, a single node replacement could break data movement for an entire region.</p><p>The fix was hiding in plain sight. The answer to “which backup exists and is it complete?” already lived in the backup storage layer (Amazon S3) itself. By reading metadata directly from the backup files, we could replace the entire dependency chain with a single source of truth.</p><h3>Every Connector Inherited Casspactor’s Limitations</h3><p>Cassandra at Netflix does not just store raw tables. It backs higher level data abstractions, such as <a href="https://netflixtechblog.com/introducing-netflixs-key-value-data-abstraction-layer-1ea8a0a11b30">Key Value</a>, <a href="https://netflixtechblog.com/introducing-netflix-timeseries-data-abstraction-layer-31552f6326f8">Time Series</a>, and others, each with its own data model, access patterns, and semantics. When any of these abstractions needed to move data to Iceberg, they all funneled through Casspactor.</p><p>Every abstraction inherited Casspactor’s constraints:</p><ul><li><strong>Skewed partition failures:</strong> Casspactor could not handle tables with large partitions, a common pattern in Key Value and Time Series workloads. Jobs crashed with out-of-memory errors on some of Netflix’s largest datasets.</li><li><strong>No data model awareness</strong>: Casspactor moved raw Cassandra tables as is. Connectors for Key Value and other abstractions had to bolt on post processing to reconstruct their data models from the raw output — extra cost, extra complexity, and an extra surface for failures.</li><li><strong>Intermediate table bloat</strong>: Casspactor wrote to an intermediate Iceberg table before producing the final output. The Key Value connector added another intermediate table and a snapshots table. Connectors for abstractions on top of Key Value added even more. This compounded into significant storage cost overhead.</li><li><strong>Inability to Time Travel</strong>: by relying on multiple services to compose a backup unit, Casspactor was unable to restore prior backups in the event of cluster Topology or Keyspace schema changes.</li><li><strong>Monolithic design</strong>: Casspactor was built as a single connector, not as an engine. There was no way to build a family of purpose built connectors on a shared foundation.</li></ul><p>We needed something fundamentally different: an engine that reads directly from backups in S3, produces standard Spark DataFrames, and lets each data abstraction build its own connector with full awareness of its data model. One foundation, many connectors.</p><h3>The New Stack: A Layered Architecture</h3><p>The new architecture, built upon the foundation of Apache Cassandra Analytics and the in-house Move Data framework, represents a fundamental shift toward a layered, purpose-built stack designed for reuse and maintainability. This new engine was conceived with clear separation of concerns, moving away from Casspactor’s monolithic design. The architecture is intentionally layered with the foundation being a core S3 reading capability: the Cassandra Analytics Wrapper, which is built on top of the Open Source Cassandra Analytics with Netflix’s internal backup representation and an S3 Client.</p><p>This layer handles the raw data retrieval from backups, translating it into standard Spark DataFrames. Sitting atop this foundation is a “<strong>Connector Factory</strong>” model, via both Java UDFs and transforms which allows individual data abstractions (Key Value, Time Series, others) to build highly optimized, data model aware connectors that process the generic Spark DataFrames, avoiding the need for complex, expensive, and failure-prone post-processing steps. This layered approach ensures that improvements to the core reading engine benefit all connectors, while the connectors themselves are focused solely on data transformation.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*EwqgicjfpASkWIf3nz4q6A.png"><figcaption>Image 3: The new Connector layered stack</figcaption></figure><ul><li><strong>Handles Skewed Partitions:</strong> By moving the mutation compaction and processing to the Executor level within Spark, the new engine can efficiently handle tables with highly skewed or wide partitions, a major pain point for Casspactor. Crucially, this processing occurs without excessive data shuffling, preventing out-of-memory errors and enabling reliable movement of Netflix’s largest datasets.</li><li><strong>Operates at Spark DataFrames (No Intermediary Tables)</strong>: The new architecture directly generates standard Spark DataFrames from the Cassandra backups. This eliminates the need for Casspactor’s costly, multi-stage intermediate Iceberg tables, which led to storage bloat and operational complexity. This native DataFrame operation enables the “Connector Factory” by providing a universal, easily consumable interface for building diverse, model specific connectors.</li><li><strong>Jobs Auto Size:</strong> The engine integrates intelligent auto-sizing capabilities, allowing jobs to dynamically adjust resource consumption based on the source table’s characteristics. This removes the burden of manual tuning from engineering teams, ensuring optimal performance and cost efficiency without sacrificing reliability.</li><li><strong>Reduced Dependencies</strong>: By reading metadata directly from the backup files stored in S3, the new stack removes the fragile, multi-service dependency chain that plagued Casspactor. S3 becomes the single, authoritative source of truth for backup existence and completeness, vastly improving data movement reliability and consistency.</li><li><strong>Time Travel</strong>: A critical feature of the new stack is the ability to process the <strong>schema, cluster topology, and data as a cohesive unit </strong>at a specific point in time. This capability provides robust time travel functionality, essential for auditing, debugging, disaster recovery and reproducing past data states.</li><li><strong>Performance</strong>: Collectively, these architectural improvements, including native DataFrame processing, optimized partition handling, and streamlined metadata retrieval have resulted in notable performance gains, reducing overall data movement execution runtime and cost compared to the legacy Casspactor system.</li><li><strong>Cost</strong>: by eliminating intermediary Iceberg tables and efficient SSTable compaction on Executors, the new stack needs a significantly smaller storage and compute footprint leading to significant cost savings in the order of USD millions.</li></ul><h3>The Journey Towards a Safe Migration</h3><p>The successful validation of the new stack was the critical first step, but it only marked the beginning of the most challenging phase: the migration. Large scale data migrations are inherently complex, high-risk undertakings that can be time consuming and often result in customer frustration and service disruption. To navigate the high stakes of decommissioning a mission-critical system like Casspactor and seamlessly replacing it, we needed a strategy that prioritized reliability and transparency above all else.</p><p>The migration was fundamentally enabled by a <strong>Like-for-Like</strong> strategy, which served as the cornerstone of our Platform Engineering philosophy, abstracting complexity. The core tenet was to maintain absolute consistency across the user-facing interface, the output contract, and the final data artifact. This meant ensuring that the data movement parameters defined via the Data Bridge abstraction remained unchanged, and, critically, the schema, metadata, and data within the destination Iceberg tables were identical to the legacy output. By preserving these external contracts, we eliminated the need for complex, time-consuming coordination with dozens of internal teams who relied on these data pipelines. This approach transformed the migration from a distributed, high-risk, multi-team effort into an internal platform implementation detail, allowing us to achieve a transparent, zero-impact transition and accelerate the retirement of the legacy system without requiring any code changes or validation from downstream users.</p><p>To navigate this migration, we developed a strategy anchored by three core pillars that serve as a blueprint for successful, large-scale data migrations:</p><ol><li><strong>Validation</strong>: Establishing and maintaining absolute confidence in data consistency through rigorous, ongoing validation.</li><li><strong>Visibility</strong>: Instrumenting every part of the system to provide a clear, real-time understanding of migration progress and system health.</li><li><strong>Safety</strong>: Ensuring user impact is minimized or eliminated, despite the inevitable system failures, by leveraging abstractions and robust fallbacks.</li></ol><p>The next section will provide a detailed exploration of these key pillars.</p><h3>Pillar 1: Validation</h3><p>Trust is earned, and in data migration, it is earned one row at a time. The first pillar is the most critical: providing a measurable guarantee to users and partners that the data produced by the new system is an exact, row-by-row replica of the data produced by the old one.</p><p>Our foundational tactic was deploying the new Move Data connector in a “shadow” testing that ran in parallel with the production Casspactor jobs. This allowed us to validate the new system with real-world, production workloads without any customer impact.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*L8yxnLU-7YqzHeEAmVRv6g.png"><figcaption>Image 4: Shadow job structure leveraged for data validation</figcaption></figure><ul><li>Let <strong>C</strong> be the set of rows in the legacy Casspactor output (Iceberg table).</li><li>Let <strong>M</strong> be the set of rows in the new Move Data output (Iceberg table).</li></ul><p>The test for trust: prove that <strong>C = M</strong>. This required continuously checking for two conditions:</p><ol><li>Rows in <strong>C</strong> but not in <strong>M</strong> (<em>C-M</em>): The new system missed data.</li><li>Rows in <strong>M</strong> but not in <strong>C</strong> (<em>M-C</em>): The new system introduced phantom or erroneous data.</li></ol><p>Any result where the cardinality of these difference sets (the number of differing rows) was greater than zero triggered an immediate, high-priority investigation. The target was 100% similarity.</p><h3>Uncovering and Resolving Disparities</h3><p>The shadow mode quickly became a powerful forensic tool, exposing “unknown unknowns”, subtle discrepancies that were not bugs in the new system but rather differences in behavior between the new and old systems. Resolving these was the core work of building trust. For each problem we initiated an investigation log where we captured the details, logs, queries that allowed us to diagnose. Based on the assessment the issues were categorized so that similar differences on other datasets were later resolved affecting many of the shadow pipelines.</p><p>Maintaining an investigation log was critical to organize the outstanding issues and effectively communicate to stakeholders the progress and confidence of the new connector so that we effectively measure the appropriate level of “confidence” to initiate the migration.</p><p>We observed differences in how connectors leverage reference timestamps for Time-to-Live, Consistency Levels, backup selection, and various internal business logic. This continuous, data-driven cycle of discovery and resolution was the mechanism by which we built confidence in the new architecture.</p><h3>Pillar 2: Visibility</h3><p>Trust is built in the background, but an active migration requires real-time insight: Visibility. The second pillar involves instrumenting the system to provide an unambiguous, clear understanding of operational health and migration progress.</p><p>We extended our instrumentation to the overall migration workflow and its dependencies:</p><ul><li><strong>Dashboards</strong>: We created centralized dashboards to track migration status, visualizing the total number of data movements migrated versus those remaining. The dashboards tracked execution status, average runtime, and cost comparisons between the two connectors.</li><li><strong>Dependency Tracking</strong>: Since the new system relied on a new set of APIs to fetch backup metadata, we implemented detailed metrics for failures to keep track of the APIs or dependencies failed.</li><li><strong>Alerting</strong>: Proactive alerts were set up for job failures (Move Data or Casspactor), failures on Move Data that triggered a fallback to Casspactor or any data discrepancy being detected.</li></ul><p>This comprehensive instrumentation allowed the team to be proactive, fix issues as they emerged during the migration, and gain the necessary confidence to accelerate the migration timeline.</p><h3>Pillar 3: Safety</h3><p>Even with perfect data correctness and enhanced visibility, the third pillar, Safety is required for a zero-impact migration. The challenge is ensuring that when a system inevitably fails, the user experience is uninterrupted. Our strategy centered on decoupling the user’s workflow from the underlying connector implementation.</p><h3>Leveraging Abstraction: The Decider Pattern</h3><p>To achieve a transparent swap, we leveraged the <a href="https://github.com/Netflix/maestro?tab=readme-ov-file">Maestro</a> workflow orchestration platform to implement the Decider pattern:</p><ol><li><strong>Data Movement Abstraction:</strong> From a user’s perspective, their Data Movement job definition remained the same.</li><li><strong>The Decider Step</strong>: Internally the workflow responsible to execute the job was modified to include a Decider step. This step took the data movement parameters (source cluster, table name, destination) and invoked a control plane: Connector Controller.</li><li><strong>Connector Controller as the Registry</strong>: The control plane served as the dynamic registry. Based on the migration cohort and the data movement attributes, it determined and reported the appropriate connector to use either Casspactor (legacy) or Move Data (new).</li></ol><p>This abstraction gave our team complete control. We could upgrade or rollback any connector for any data movement instantly by simply updating a configuration in the controller, with zero modification required to the thousands of downstream customer workflows. Crucially, this abstraction guaranteed the critical safety net: a conditional step in the Maestro workflow logic ensured that if the Move Data step fails, it would immediately execute the Casspactor step.</p><p>This pattern would increase the chances that the user’s data movement completes successfully, even if the new connector encountered a bug or transient failure during the initial rollout phases. User impact was completely eliminated; they might see a slightly longer runtime in the event of a failure and fallback, but they would never see a migration failure or suffer from stale data.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*1Qq5sgvtKuLfpP2byfggMw.png"><figcaption>Image 5: The Decider Pattern Implementation via Maestro</figcaption></figure><p>Beyond the workflow, the new system architecture itself was inherently more resilient. By building the new data movement connector on Cassandra Analytics and reading backups directly from S3, we removed fragile dependencies on deprecated internal services.</p><h3>Conclusion</h3><p>The migration from Casspactor to the new, layered architecture built on Cassandra Analytics and the Move Data connector was more than a typical “tech debt” project; it was a fundamental shift in our approach to data movement reliability and scalability at Netflix.</p><p>The legacy system, while serving us well for years, was ultimately constrained by monolithic design, fragile metadata dependencies, and an inability to handle the complexity of modern data abstractions. The new stack resolves these issues by delivering a robust, cost-efficient, and inherently more resilient solution that reads directly from S3, handles wide partitions gracefully, and eliminates costly intermediate tables.</p><p>Our blueprint for the migration, anchored by the three pillars of Validation, Visibility, and Safety, ensured a transparent and high-confidence transition. Through rigorous shadow testing and a data-driven audit framework, we achieved the desired data consistency. Enhanced dashboards and alerting provided the real-time operational insight necessary to manage risk. Most critically, the implementation of the Decider pattern within our workflow abstraction minimized the impact for all downstream users.</p><p>This successful migration validates a core philosophy: by abstracting complexity at the platform level, we can perform large system migrations without burdening our product engineering partners. The new foundation is now ready to support the next generation of Netflix’s data abstractions.</p><h3>Looking ahead</h3><p>This foundational work on the Cassandra Data Movement stack has done more than just replace a legacy system: it has become an accelerator for innovation across the entire Data Movement organization. By providing a reliable, performant engine that standardizes data retrieval into Spark DataFrames, we’ve enabled the rapid development of new, highly optimized connectors. This new “Connector Factory” approach has already delivered a dedicated Key-Value to Iceberg and Time Series connectors, both of which are fully aware of their respective data models, eliminating costly post-processing. This architecture is also paving the way for ambitious new initiatives, including the development of a solution for bulk loading data into Cassandra itself, effectively completing the data movement cycle, and enabling safer fleetwide connector rollout with canaries inspired by the Decider Pattern.</p><p>We are incredibly grateful for the extensive collaboration among the Data Movement, Data Bridge, Online Data Stores, Membership, Billing, Subscriber and Ads platform teams at Netflix; this work simply couldn’t have been accomplished without their partnership!</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=6e13329c80a1" width="1" height="1" alt=""><hr><p><a href="https://medium.com/netflix-techblog/the-evolution-of-cassandra-data-movement-at-netflix-6e13329c80a1">The Evolution of Cassandra Data Movement at Netflix</a> was originally published in <a href="https://medium.com/netflix-techblog">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://medium.com/netflix-techblog/the-evolution-of-cassandra-data-movement-at-netflix-6e13329c80a1</link>
      <guid>https://medium.com/netflix-techblog/the-evolution-of-cassandra-data-movement-at-netflix-6e13329c80a1</guid>
      <pubDate>Sat, 20 Jun 2026 01:53:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Thinking Fast & Slow for a Personalized Notification System]]></title>
      <description><![CDATA[<p>by <a href="https://www.linkedin.com/in/matthew-wood-bbb32584/">Matthew Wood</a>, <a href="https://www.linkedin.com/in/ishangupta1993/">Ishan Gupta</a>, Kevin Mercurio, <a href="https://www.linkedin.com/in/devon-bryant-9910ab5/">Devon Bryant</a>, and <a href="https://www.linkedin.com/in/clairedorman/">Claire Dorman</a></p><p>In his seminal book “Thinking, Fast and Slow,” Daniel Kahneman describes two systems that drive human cognition: System 1, which operates automatically and quickly with little effort, and System 2, which allocates attention to more challenging mental activities requiring deliberate focus. This dual-process theory has profound implications not just for understanding human behavior, but for designing intelligent systems that must balance immediate responsiveness with strategic foresight. Similar “plan vs. act” decompositions show up in other domains too — for example, robotics and autonomous driving often separate a slower planning layer (setting goals and constraints over longer horizons) from faster control and execution loops, and modern LLM agents frequently pair deliberate planning with rapid, step-by-step tool use and reaction.</p><p>At Netflix, our messaging platform faces a similar challenge every day. We send hundreds of millions of personalized notifications — push messages, emails, and in-app alerts — to help members discover content they’ll love. This creates a central tension: optimizing each notification for near-term engagement can conflict with what is best for the member over the long term. Higher message frequency can increase fatigue and opt-out risk, while lower frequency can reduce awareness of relevant titles and features the member would value.</p><p>This blog post introduces our framework for personalized notifications — a hierarchical system where a “slow” policy makes strategic, personalized decisions about a member’s weekly messaging plan (e.g., the intended frequency per channel and the resulting pacing over the week), while a “fast” policy handles the tactical, real-time decisions about which specific message to send when a send opportunity occurs. Together, they balance near-term engagement with longer-term member experience.</p><h3>The Problem:</h3><p>Before introducing our new framework, it is helpful to ground the discussion in a representative baseline for a personalized notification system. In our previous production system, we used a causal model to make send decisions by predicting the causal effect of a single message over a short time horizon. While this approach is effective as a baseline, it suffers from two fundamental limitations:</p><h3>Short-Term Reward Horizons</h3><p>The single-message outcome model is trained to optimize short-horizon metrics, such as immediate user actions occurring shortly after a notification is sent. While this is excellent for driving near-term engagement, it misses the cumulative, long-term effects of a messaging strategy. A message that drives an interaction today might also contribute to notification fatigue, reducing responsiveness in the weeks to follow. Because critical indicators of member satisfaction — like sustained viewing habits or gradual opt-out risk — only surface over extended timeframes, a short-term model will always miss the bigger picture.</p><h3>Coupled Ranking and Pacing Decisions</h3><p>When a single system evaluates daily incrementality to decide both whether to send something and, if so, which item to send, an individual member’s weekly message frequency becomes a by-product of those daily decisions rather than an explicit control variable. In our previous single-policy system, frequency was controlled implicitly through a relevance threshold on the model score calibrated to achieve a target aggregate send rate. While effective for managing overall frequency, this mechanism limited the system’s ability to personalize frequency based on individual engagement patterns. Moreover, because send eligibility and message selection were coupled in the same decision rule, adjusting the threshold to control frequency also changed the distribution and quality of selected messages, and vice versa.</p><p>To solve these challenges, we needed a system that could separate longer-term strategy from shorter-term decisions. What if we could determine an optimal, personalized message plan for each member, and then focus on selecting the most relevant content within those bounds? In the following sections, we detail how we realized this vision by decoupling our notification engine into a hierarchical ‘System 1’ and ‘System 2’ framework.</p><h3>The Proposed Method: A Hierarchical Slow-Fast Architecture</h3><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*BffuQUAFrEhYzynBKsVPLA.png"></figure><p>The Slow policy’s primary role is to define a <strong>personalized pacing of messages over a defined time horizon</strong>. The decisions made by slow policy are consumed by the Fast Policy whose role is to maximize immediate relevance and select the optimal message for the member at any given moment.</p><p>To illustrate the Slow Policy in practice: For example, if optimized at a weekly cadence, the policy evaluates a member’s long-term engagement patterns to select a “Pacing Plan Action.” To keep the action space manageable yet expressive, we discretize the decision space into a set of actions that independently specify push and email frequencies. This provides approximately O(100) distinct combinations of cross-channel pacing strategies.</p><p><strong>The Utility Function</strong></p><p>The Slow policy selects actions by maximizing a personalized utility function. This function explicitly trades off positive engagement signals against the long-term “cost” of messaging.</p><p><em>U(member, action) = Σ wₖ·Reward_k(member,action) — Cost(action)</em></p><p>To capture a holistic view of member health, this utility is composed of:</p><ul><li><strong>Positive Signals:</strong> Capturing the likelihood that a member will find value in and engage with the platform.</li><li><strong>Negative Signals:</strong> Capturing the likelihood of member fatigue or a propensity to opt out of a messaging channel.</li></ul><p>Ideally, negative signals alone would naturally penalize over-messaging. In practice, however, explicit negative feedback is extremely sparse. Without an additional constraint, the predicted ‘cost’ of an incremental message appears negligible, causing the model to gravitate toward maximum frequency.</p><p>To address this, we introduce a <strong>universal message cost</strong> that is added to the personalized negative‑feedback prediction for every send. This additional cost term keeps the reward function concave and well‑behaved, preventing degenerate “always send” policies. The message cost parameter is empirically tuned using a combination of online experiments and offline evaluation metrics.</p><p><strong>Pacing Strategy</strong></p><p>The two-stage design naturally allows for optimizing both the average frequency as well as pacing of messages over time. The simplest pacing strategy is uniform random: we translate the frequency target into a per-opportunity send probability and, at each eligible opportunity, effectively flip a weighted coin to decide whether to send. This produces an organically randomized pattern whose expected send rate matches the target.</p><p>While uniform pacing provides a clean and robust baseline, the framework readily extends to richer, non-uniform pacing profiles (for example, day-of-week patterns, conditioning on user activity, or launch-aligned bursts) whenever product or user-experience considerations call for more structured temporal distributions.</p><p><strong>Policy-to-Policy Communication</strong></p><p>The true power of this hierarchy lies in decoupling. By splitting into “Slow” and “Fast” policies, we allow each to focus on what it does best.</p><p>To bridge these two worlds asynchronously, decisions are events and state is managed through a low-latency feature store:</p><ul><li><strong>The Planner (Slow):</strong> The Slow policy calculates a member’s ideal pacing plan. It writes this strategic intent to a feature store</li><li><strong>The Executor (Fast):</strong> Every day, when a notification opportunity arises, the Fast Policy simply pulls that stored “plan” as a feature. It then executes the tactical send decision within those strategic guardrails.</li></ul><p>This architecture provides two critical advantages:</p><ol><li><strong>“Stickiness”:</strong> It ensures a member receives a consistent experience. The Slow policy will be executed once at a defined cadence; the plan is stored and honored.</li><li><strong>Independent Evolution:</strong> We can retrain, optimize, or A/B test our weekly pacing strategies (the “Slow” layer) without ever touching the real-time ranking logic (the “Fast” layer).</li></ol><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*gZiJ6HiK_BG3rDCRupx5vA.png"><figcaption>Figure 1: Schematic of the two-layer message personalization system composed of a slow planning policy (top) and a fast execution policy (bottom). A feature store serves as the communication bridge between the two policies.</figcaption></figure><h3>Key Results &amp; Takeaways</h3><p>The transition to a hierarchical architecture resulted in one of our <strong>largest production metric lifts to date</strong>. We observed several key breakthroughs:</p><ul><li><strong>Empowering the “Casual Viewer”</strong>: Gains were most significant among members who watch less frequently — a critical cohort where timely, high-relevance awareness of new content is vital.</li><li><strong>The Power of Decoupling</strong>: Separating <em>frequency planning</em> from <em>message selection</em> was as transformative as the modeling itself. This new architecture unlocks incredible flexibility, allowing us to iterate on content ranking models and pacing strategies as two independent, clean variables.</li><li><strong>Respecting the Horizon</strong>: The impact of messaging is rarely an isolated event; its effects build up cumulatively based on ongoing interactions between our system and the member. By isolating pacing into a dedicated strategic layer, we now have the mechanism to explicitly manage long-term fatigue and opt-out risk.</li></ul><h3>Acknowledgments</h3><p>We could not have delivered this project without the help of our outstanding colleagues, and we sincerely thank them for their contributions.</p><p><strong>Feature Store Team</strong>: <a href="https://www.linkedin.com/in/aaronlewism/">Aaron Lewis</a>, <a href="https://www.linkedin.com/in/tom-switzer-59824356/">Tom Switzer</a>, <a href="https://www.linkedin.com/in/abbywh/">Abby Whittier</a>, <a href="https://www.linkedin.com/in/ray-zhang-a7168a32/">Ray Zhang</a><br><strong>Product: </strong><a href="https://www.linkedin.com/in/fenglinli/">Fiona Li</a><br><strong>AI for Member Systems (supporting contributor): </strong><a href="https://www.linkedin.com/in/sergipv/">Sergi Perez</a></p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=4d89b26525cd" width="1" height="1" alt=""><hr><p><a href="https://medium.com/netflix-techblog/thinking-fast-slow-for-a-personalized-notification-system-4d89b26525cd">Thinking Fast &amp; Slow for a Personalized Notification System</a> was originally published in <a href="https://medium.com/netflix-techblog">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://medium.com/netflix-techblog/thinking-fast-slow-for-a-personalized-notification-system-4d89b26525cd</link>
      <guid>https://medium.com/netflix-techblog/thinking-fast-slow-for-a-personalized-notification-system-4d89b26525cd</guid>
      <pubDate>Sat, 20 Jun 2026 01:53:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[A Human-Augmenting Agentic Workflow for Causal Inference]]></title>
      <description><![CDATA[<p>By Winston Chou, Adrien Alexandre, Lars Olds, Yi Zhang, Garrett Hagemann, and Nathan Kallus</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*7midjxNL3H3BlZvQy4seSA.png"></figure><h3>Introduction</h3><p>Imagine asking a data agent to analyze the causal relationship between two variables, such as the effect of watching a popular Netflix show on long-term member retention. It queries your data, runs a regression, and confidently returns an answer. How much should you trust it? Can you be confident that the agent accounted for subtle biases — or does it treat passionate fans as if they were the average viewer? Without deep understanding and expertise, would you even be able to tell if it got the answer wrong?</p><p>Data analysis is increasingly being delegated to software agents. While this reduces human effort and toil, oversight is still needed to ensure the validity of results. This is especially true for specialized tasks like <a href="https://en.wikipedia.org/wiki/Observational_study">Observational Causal Inference (OCI)</a>, which require substantial judgment and domain expertise.</p><p>In this blog post, we share an agentic workflow for performing OCI under <a href="https://www.nber.org/system/files/working_papers/t0294/t0294.pdf">unconfoundedness</a>. Our workflow is designed for software agents to adhere to rigorous, exhaustive templates for causal inference tasks. Yet, it also seeks to be “<a href="https://pubsonline.informs.org/doi/10.1287/mnsc.2024.05684">human-augmenting</a>,” and to enable and empower human inspection and evaluation.</p><p>We designed this workflow with OCI practitioners in mind. Although OCI requires context and care to do well, aspects of it — e.g., checking and rechecking covariate balance, conducting sensitivity analyses, and keeping track of multiple iterations — can be repetitive and prone to error. Our workflow seeks to eliminate this toil so that humans can focus on more nuanced tasks, such as framing questions, scrutinizing assumptions, and evaluating results.</p><p>To this end, <strong>we are open-sourcing a standalone version of our </strong><a href="https://github.com/Netflix-Skunkworks/oci-agent"><strong>oci-agent</strong></a> so that OCI practitioners can model workflows on and suggest improvements to it. We also share evaluations of our agent on the 2016 Atlantic Causal Inference Conference (ACIC) <a href="https://files.eric.ed.gov/fulltext/ED591944.pdf">competition datasets</a>, and show that our agent systematically beats one-shot iterations under numerous data-generating processes — while achieving competitive results against hand-tuned benchmarks.</p><p>This post describes the principles behind our workflow and gives a case study of its deployment at Netflix.</p><h3>Philosophy</h3><p>Our workflow is built on top of Netflix’s pre-existing OCI toolkit. We built this toolkit — largely in a pre-AI world — to answer “point-in-time” causal questions, such as “what is the effect of playing a Netflix game on member retention?” or “what is the effect of watching a highly popular show on subsequent engagement?” Questions of this kind inform business strategy, guide metric development, and contribute to a rich understanding of what drives member satisfaction.</p><p>Our toolkit is guided by a “<a href="https://academic.oup.com/aje/article-abstract/183/8/758/1739860">target trial emulation</a>” philosophy. For any point-in-time OCI question, we ask “what is the ideal A/B test for addressing this question?” This A/B test may be expensive, slow, or even infeasible to run. However, the thought exercise helps to pin down the assumptions needed for a credible answer, such as unconfoundedness of the treatment.</p><p>To make the target trial analogy actionable, our toolkit embeds a series of <strong>design diagnostics</strong>. These diagnostics assess whether we are drawing fair comparisons between treated and untreated units — or if there are hidden differences that could imperil our conclusions:</p><ul><li><em>Covariate balance.</em> After weighting, the standardized mean difference of pre-treatment covariates between treatment and control groups should be less than 0.2.</li><li><em>Overlap.</em> The probability of receiving treatment (aka propensity score) should be bounded between 0.1 and 0.9.</li><li><em>Placebo outcome.</em> The “treatment effect” on variables measured prior to the treatment should not be significantly different from zero.</li><li><em>Sensitivity to hidden confounders.</em> Findings of treatment effects are contextualized by sensitivity to hypothetical omitted variables that explain both treatment and outcome.</li></ul><p>As we uplevel our OCI toolkit with agents, such evaluation remains paramount. The standard approach to evaluating agents is to programmatically compare their outputs to ground truth. Yet, outside of artificially simulated data, there is <a href="https://en.wikipedia.org/wiki/Rubin_causal_model#The_fundamental_problem_of_causal_inference">no ground truth in observational causal inference</a>.</p><p>Without discounting the need for evals (which our workflow also supports), one of our key principles is to augment <em>human</em> evaluation by making each analytic step as transparent as possible. For example, in our workflow, agents publish artifacts in the form of plans, specifications, plots, and notebooks that humans can inspect and re-execute if they wish. In the absence of ground truth, we rely on these “process audits” — coupled with human oversight — to build good agents.</p><h3>Principles</h3><p>Our workflow has three key personas:</p><ol><li><strong>Principal</strong> — the human user (e.g., data scientist) whose goal is to provide a thorough and correct analysis</li><li><strong>Actor</strong> — the software persona that performs the analysis, including diagnostics</li><li><strong>Critic</strong> — the software persona that synthesizes results, identifies gaps, and offers suggestions to improve the analysis</li></ol><p>Our agent orchestrates the latter two personas in an actor-critic loop: specifying and triggering the analysis as the actor, then interpreting results and diagnosing flaws as the critic.</p><p>Each persona has responsibilities:</p><p><strong>Principals</strong></p><ul><li>Provide an initial analysis plan containing its context and goals.</li><li>Provide context on the main threats to valid inference and the confounders that must be controlled.</li><li>Specify the tools that can be used for the analysis.</li><li>Specify the data model and dataset.</li></ul><p><strong>Actors</strong></p><ul><li>Refine the principal’s plan into a data analysis spec.</li><li>Use only the tools provided by the principal.</li><li>Create human- and machine-checkable artifacts.</li><li>Perform the four design diagnostics in addition to the core analysis.</li><li>Report any remediations taken in case of diagnostic failures.</li></ul><p><strong>Critics</strong></p><ul><li>Check for blind spots, such as unmentioned confounders, in the principal’s plan.</li><li>Check for alignment between the plan, spec, and executed analysis.</li><li>Specify a credibility level in the results after seeing the diagnostics.</li><li>Specify if and how the estimand differs from the Average Treatment Effect (ATE), for example due to propensity score trimming.</li><li>Contrast the executed analysis with the ideal target Randomized Controlled Trial (RCT).</li><li>Suggest at least one alternative measurement strategy (e.g., encouragement RCTs).</li></ul><p>Although our workflow is designed for OCI under unconfoundedness, the principles listed in this section are meant to be extensible to other approaches to OCI, such as panel methods with very different assumptions (e.g., parallel trends).</p><h3>Empowering Human Evaluation</h3><p>To empower human oversight of each analytic step, we provide principals with a templated notebook that uses our vetted (non-agentic) OCI toolkit, which employs <a href="https://arxiv.org/abs/2506.07462">doubly robust learning for causal effect estimation</a>.</p><p>The principal’s remaining responsibilities are to write the initial analysis plan and to evaluate the analysis artifacts (the executed notebook and the critic’s report). To enable thorough evaluation, agents version-control their reports and upload executed notebooks to a file store, where they can be downloaded and re-executed by principals (if they wish).</p><p>We diagram this workflow below:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*1cTcjXY_Gh80OB2x1xzlAQ.png"></figure><h3>Case Study — Estimating the Impact of New Entertainment Types</h3><p>In recent years, we have added a wide variety of entertainment types beyond streaming video to Netflix. A natural question is how these new entertainment types affect members’ satisfaction and their likelihood of continuing to subscribe to Netflix.</p><p>To analyze the impact of one of these new entertainment types, which we will call Type X, we wrote a simple analysis plan specifying our</p><ul><li><strong>Treatment</strong>: Days engaging with Type X (or “Type X days” for short)</li><li><strong>Outcome</strong>: Two-month retention</li><li><strong>Potential</strong> <strong>confounders</strong>, including pre-treatment Type X days</li></ul><p>To establish a baseline, we fed this analysis plan without additional scaffolding to Claude Sonnet 4.6, a powerful yet accessible general-purpose model. The model chose and executed a defensible analysis strategy: linearly regressing retention on Type X days along with controls.</p><p>While the result was polished and impressive, when we ran the same analysis through our paved path tooling and agentic workflow, also using Sonnet 4.6, our agent produced an updated estimate that was just 25% of the baseline! What explains the difference between the baseline and the paved-path estimates?</p><p>A core challenge when analyzing new entertainment types is <strong>early adopter bias</strong>. The first users of any new offering are likely to be systematically different from the general population. For example, they may be heavier users of Netflix generally, or they may be extremely strong fans of the underlying titles. Early adopter bias manifested in our analysis as poor “overlap”: the vast majority of observations had a small estimated probability of engaging with Type X, reflecting its early maturity.</p><p>This imbalance was caught by our critic agent in its writeup of the analysis. The critic also flagged the failure of the placebo test: early Type X adopters differed significantly from non-adopters in terms of important confounders measured <em>before</em> experiencing the treatment, a warning sign of potential bias.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*Ee0gswtciDVvuFr1QenlvQ.png"></figure><h3>Addressing Failed Diagnostics</h3><p>To address these diagnostic failures, our workflow provides agents with a playbook. For example, to overcome poor overlap, we instruct the agent to use <a href="https://academic.oup.com/biomet/article-abstract/96/1/187/235329?redirectedFrom=fulltext">Crump-style trimming</a>. That is, before estimating causal effects, the actor trims units with estimated propensity scores outside the range [0.1, 0.9]. This scopes the treatment effect being estimated to the ATE in the population that is not very likely or unlikely to engage in the new entertainment type — an important caveat we instruct the critic to flag in its report.</p><p>Trimming yields an estimate that is much smaller than the baseline estimate, and which only applies to the “overlapping” population (for whom engagement with the new entertainment type is non-deterministic). However, the trimmed estimate is substantially more credible, as it focuses on the members for whom the treatment could plausibly be randomly assigned, as in a target trial.</p><p>Contrastively, the baseline effect relies heavily on assumptions to extrapolate outcomes for all members, even those with a very low probability of treatment. The <a href="https://gking.harvard.edu/files/counterft.pdf?utm_source=chatgpt.com">danger</a> here is that extrapolation produces a number that is not backed by robust data and is likely confounded by early adopter bias.</p><h3>Orchestrating Followup Analyses</h3><p>There are two natural followups to this analysis:</p><ol><li>First, we need to analyze the sensitivity of estimates to the choice of trimming threshold. Practically, this requires redoing the analysis with multiple trimming thresholds.</li><li>Second, we also care about how these causal effects evolve over time. Yet, comparing causal effects across time raises subtle challenges. For example, we need to coordinate the population across all analyses: if a set of users is trimmed to make one analysis more credible, it should be trimmed in the other analyses as well.</li></ol><p>Both of these followups require conducting multiple versions of the same analysis, tweaking some parameters while keeping others the same. Managing this complexity and ensuring consistent execution is another area where agents add value.</p><p>To illustrate this, below we show a sensitivity analysis for our case study in which we asked the agent to vary the trimming bounds from [0, 1] (no trimming) to [0.15, 0.85]. As the plot shows, the estimated ATE on the overlapping population is robust to the choice of trimming threshold within bounds of [0.005, 0.995]. Although principals could (and should) execute <a href="https://psicostat.github.io/4ms-winter-school/papers/Steegen%20et%20al.%202016%20-%20Increasing%20Transparency%20Through%20a%20Multiverse%20Analysis.pdf">this and other robustness analyses</a>, delegating them to agents helps to reduce toil.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*GbQUhS5LB9gcn16VDnC69w.png"></figure><p>Another example is generating a time series by repeating the same analysis across multiple date partitions. For example, below we plot the results of using our agent to refit a different analysis on ten distinct date partitions. The plot shows evidence of seasonality: the treatment has a stronger effect on the winter dates compared to the summer dates.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*792mstFbTU-1ZVSs7KJEVw.png"></figure><h3>Public Repo and Evals</h3><p>To help OCI practitioners build on and contribute to our workflow, we are open-sourcing a standalone version of <a href="https://github.com/Netflix-Skunkworks/oci-agent">oci-agent</a>. This repo implements two evaluations on public datasets from the 2016 Atlantic Causal Inference Competition (ACIC) data analysis competition. It also includes a lightweight version of our internal causal machine learning notebook that only uses open-source software (<a href="https://github.com/py-why/EconML">EconML</a>).</p><p>Our first evaluation runs this notebook for three randomly sampled datasets generated by each of the 77 data-generating processes (DGPs) in the ACIC data. Next, it uses the critic to grade the resulting 231 estimates as either satisfactory or unsatisfactory based on the diagnostics.</p><p>Below, we plot the average RMSE and coverage of 95% confidence intervals of our ATT estimates against the 44 competitor methods in the ACIC competition. As the scatterplot shows, our statistical methodology is competitive against these benchmarks: it achieves reasonably low RMSE and well-calibrated confidence intervals that cover the truth in ~95% of DGPs.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*nvNVibkHRO8EgqNcliB_Zg.png"></figure><p>More to the point, our diagnostics and agentic workflow help to separate more reliable estimates from less reliable estimates. To illustrate this, the following chart plots our ATE estimates in terms of RMSE and coverage. Note that we separate out the RMSE and coverage of:</p><ul><li>All 231 estimates (purple dot)</li><li>The 192 satisfactory estimates (blue star)</li><li>The 39 unsatisfactory estimates (red dot)</li></ul><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*SGM_i82SxCI64iMIp-diFw.png"></figure><p>As the plot shows, when aided by our diagnostic suite, the critic agent is able to separate good estimates from bad estimates: the satisfactory estimates have <em>much </em>lower RMSE and better calibrated confidence intervals than do the unsatisfactory estimates.</p><p>Our second evaluation compares the performance of an LLM on the same analysis plan with our scaffolding and without it (i.e., one-shot prompting). Unsurprisingly, we find that our scaffolding is critical to helping the LLM return useful estimates. This can be seen in the following random sample of ten ACIC datasets. Using our scaffolding, the LLM recovers the ground truth in nine out of ten datasets. Furthermore, estimates are highly correlated with ground truth.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*k105qpwAZ3bjY_kLvAEpUA.png"></figure><p>In contrast, giving the same analysis plan to Sonnet 4.6 without any scaffolding (i.e., just prompting it) results in consistently wrong answers that are not at all correlated with ground truth.</p><p>A key limitation of our public repo is that, due to the synthetic nature of the underlying datasets, it doesn’t pressure-test our agent’s semantic understanding or performance on real-world OCI tasks. Nonetheless, the repo demonstrates the core principles underlying our workflow. These include (1) giving agents with extensive scaffolding so that they follow best practices by design, and (2) requiring inspectable artifacts so that humans can audit agents’ <em>processes,</em> not just their outcomes.</p><h3>Conclusion</h3><p>We provide a workflow for doing observational causal inference with the help of software agents. Leveraging elements of our pre-AI OCI toolkit, such as templated notebooks, our workflow is designed to ensure that agents conduct rigorous and exhaustive analyses. This helps to reduce the human toil of OCI, which can be a highly iterative and exacting process.</p><p>At the same time, motivated by the complexity and ambiguity of observational causal inference, our workflow seeks to be <strong>human-augmenting</strong> and enables human practitioners to evaluate each analytic step.</p><p>Using agents for causal inference poses a challenge: how do we evaluate agents’ performance on tasks without ground truth? To meet this challenge, our workflow combines process audits with human oversight. To enable others to learn from and critique our workflow, we have <a href="https://github.com/Netflix-Skunkworks/oci-agent">open-sourced</a> a lightweight, standalone version. We hope this work stimulates more research and development on agentic evaluation in the absence of ground truth.</p><p><em>For valuable feedback on this post and “dogfooding,” we thank Adith Swaminathan, Ayal Chen-Zion, Colin Gray, Juliet Hougland, and Simon Ejdemyr.</em></p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=4623f0a9c5af" width="1" height="1" alt=""><hr><p><a href="https://medium.com/netflix-techblog/a-human-augmenting-agentic-workflow-for-causal-inference-4623f0a9c5af">A Human-Augmenting Agentic Workflow for Causal Inference</a> was originally published in <a href="https://medium.com/netflix-techblog">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://medium.com/netflix-techblog/a-human-augmenting-agentic-workflow-for-causal-inference-4623f0a9c5af</link>
      <guid>https://medium.com/netflix-techblog/a-human-augmenting-agentic-workflow-for-causal-inference-4623f0a9c5af</guid>
      <pubDate>Sat, 20 Jun 2026 01:53:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Predicting Risk in Content Launches: How Data-Driven Insights can Transform Launch Planning]]></title>
      <description><![CDATA[<p>by <a href="https://www.linkedin.com/in/ecgill/">Emily Gill</a></p><p><em>Each year, we bring the Analytics Engineering community together for an Analytics Summit — a multi-day internal conference to share analytical deliverables across Netflix, discuss analytic practice, and build relationships within the community. This post is one of several topics presented at the Summit highlighting the breadth and impact of Analytics work across different areas of the business.</em></p><h3><strong>Understanding Risk in Content Launches</strong></h3><p>Every title you see on Netflix goes through several key phases: Development, Pre-Production, Production/Principal Photography, Post-Production, and finally, Launch Preparation, all leading up to the Title Launch. Once Principal Photography wraps, the focus shifts in Post-Production from content creation to quality assurance and visual effects (if needed).</p><p>At the end of Post Production, Netflix receives the final audio and video files — often delivered as an IMF (Interoperable Master Format) — which triggers a flurry of Launch Preparation activities, focused on tasks such as the development of artwork and trailers, creation of subtitles, maturity ratings &amp; quality control, that happen within a tight window and rely on having the finalized media assets in hand.</p><p>Some of this work can be kicked off earlier using a non-final version of the media called the Locked Cut, but since it’s not the absolute final deliverable, this presents a tradeoff: should our teams who prepare content for service wait for the more finalized IMF to begin their work, or start sooner with the unfinal Locked Cut? Waiting for the IMF risks a compressed timeline if it arrives late, while starting with the Locked Cut means teams may need to do additional conformance work if there are significant changes between the Locked Cut and the final IMF.</p><h4><strong>Identifying Gaps in Schedule Accuracy</strong></h4><p>To help navigate the decision of when to start launch preparation, our teams rely on estimated delivery dates for both the Locked Cut and IMF media assets, which are manually provided by content partners in production schedules. However, these schedules often have gaps in coverage and lack accuracy for both asset types (see Figure 1).</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*4KfGZI5-3PaxEU_tMjp9NQ.png"><figcaption>Figure 1. At an asset-level we generally see that scheduled date accuracy and coverage are lower at horizons further from asset delivery. As we approach delivery (moving towards the right on this plot) schedules become more accurate (errors decrease) adn coverage improves.</figcaption></figure><p>This isn’t unexpected — productions are dynamic, facing frequent changes, scheduling conflicts, and unforeseen obstacles that can shift timelines without warning. As a result, there’s a clear opportunity to leverage the wealth of production data we collect to predict the risk of schedule slips. By developing a predictive model, we aim to both fill in ETA gaps (providing asset delivery estimates when none exist) and improve the accuracy of existing ETAs compared to traditional manual schedules.</p><h4><strong>Correlation between Schedule Accuracy and Launch Misses</strong></h4><p>Our analysis reveals a strong correlation between scheduled inaccuracies and launch misses — instances where a title experiences delays. To quantify schedule inaccuracy, we created a metric called Accumulated Error Days (AED), which measures the cumulative deviation between estimated (scheduled or predicted) delivery dates and actual delivery dates over time. AED is calculated retrospectively as the area between the scheduled (grey line) or predicted (blue line) delivery dates and the actual delivery date (green line).</p><p>When we compare titles with at least one launch miss to those without, we find that mean AED is significantly higher in the group with launch misses. Notably, this effect is even more pronounced when we focus on the period closer to delivery — indicating that high AED (i.e., inaccurate schedules) in the final stretch before launch is especially correlated with launch misses, more so than AED accumulated over a longer timeline. These findings further motivate our efforts to improve schedule accuracy and reduce AED by leveraging rich production data and predictive modeling.</p><h3><strong>Modeling Time-to-Delivery</strong></h3><p>Our predictive models are designed as boosted tree regression models that predict the “days until” either media asset delivery for in-progress productions.</p><p>To power these models, we leverage a range of upstream data sources including production-level signals of progress, title metadata, and seasonal signals. We are able to predict the days until media asset delivery using daily update snapshots, allowing us to generate up-to-date predictions that reflect the latest state of each in-progress production. This means that we have each feature and what its value was as of each day of a production. Modeling with this snapshotted data enables us to generate up-to-date predictions as new information becomes available, build a flexible model that works across all production phases, and seamlessly incorporate dynamic features that evolve over time (Figure 2).</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*Yv6mwy7QZ8ghiI4S1jGv6g.png"><figcaption>Figure 2. Hypothetical illustration of the evolving nature of production-related signals used in our models. Some signals are present throughout but dynamic, others are present at single moments in time during specific production phases. By capturing data in a snapshotted form, we’re able to build a flexible phase-agnostic model that leverages many different types of progress signals. This figure is illustrative only and does not depict actual Netflix financial or production data.</figcaption></figure><h3><strong>Evaluating Our Approach</strong></h3><h4>Building a Comprehensive Metrics Suite</h4><p>When evaluating the performance of the predictive models, we look across a suite of metrics to try to understand where and when predicted dates outperform scheduled dates. Among these are mean and median absolute error, relative to actual delivery, to understand the accuracy of our estimated dates. We also consider bias metrics, such as mean and median error, to understand if we are consistently over- or under-predicting the actual delivery. We calculate the standard deviation of our errors to understand if there are large shifts in the bulk of the distribution of errors. For the tails of our error distributions, we calculate the percentage of our absolute errors that are greater than x days to delivery.</p><p>For scheduled dates, we calculate coverage across various horizons to delivery. This is a value prop of the model; we’ve built the model in such a way that we can always provide a predicted date and recoup any coverage gaps that exist from scheduled dates alone.</p><h4>Benchmarking Against Manual Scheduling</h4><p>In a backtest, we observed significant improvements across all of our metrics and across most horizons from delivery. As an example, see Figure 3 which plots global mean absolute error (MAE) and shows large reductions in errors (greater accuracy) in predicted IMF and Locked dates as compared to scheduled dates. Additionally, we see large reductions in outliers from scheduled to predicted dates as well.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*sqKvClh5t_Y1dp4ZQU3Xkw.png"><figcaption>Figure 3. This plot compares accuracy (measured as Mean Absolute Error) between predicted and scheduled dates. The horizontal axis plots time prior to delivery, which decreases from left to right until you reach the moment of delivery at the bottom right. For this particular asset, the predicted delivery dates on average are much more accurate than manually scheduled delivery dates throughout the full horizon to delivery.</figcaption></figure><p>Since our teams use these dates over a period of time and not at a single point in time, there is an additional benefit that we’re describing as an Earlier Accuracy Signal. By leveraging predictive dates, our teams benefit from a level of accuracy that they would otherwise have to wait x amount of time for if using scheduled dates. As an example, 6 months out from Locked Cut delivery the predicted dates are better than scheduled dates on 76% of titles and have a level of accuracy (6.1 wks MAE) that scheduled dates don’t reach until 11 weeks later.</p><p>Circling back to AED, which we mentioned earlier is correlated to launch misses, we find that in our backtested titles globally, and across most buying orgs and content types (i.e., series versus standalones), predicted IMF and Locked Cut dates reduce AED from scheduled dates when calculated across the 6 months leading up to delivery. We see similar patterns when we repeat this for shorter horizons to delivery as well.</p><h3>Streamlining Workflows with Improved Scheduling</h3><p>A key advantage of this predictive model is that estimated delivery dates are already integral to our stakeholders’ workflows — meaning we can introduce predictive dates without overhauling existing processes. However, this creates a new challenge: with both scheduled and predicted dates available, teams need to determine which is more reliable. While predictive dates are often more accurate on average, there are situations where scheduled dates perform better. To address this, we’ve built serving logic that defaults to scheduled dates in buying orgs where the model underperforms. Elsewhere, teams can view both dates side by side in dashboards, allowing them to apply their own judgment. Additionally, our predictive models leverage features that are tied to scheduled dates, which has emphasized the need and impact of ensuring our upstream teams continue to input and update scheduled dates even in the presence of our predictions. We’re piloting these predictive signals in multiple ways, tailoring the approach to fit the diverse needs and tools of our various launch prep functions.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=587b1f2de928" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/predicting-risk-in-content-launches-how-data-driven-insights-can-transform-launch-planning-587b1f2de928">Predicting Risk in Content Launches: How Data-Driven Insights can Transform Launch Planning</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/predicting-risk-in-content-launches-how-data-driven-insights-can-transform-launch-planning-587b1f2de928</link>
      <guid>https://netflixtechblog.com/predicting-risk-in-content-launches-how-data-driven-insights-can-transform-launch-planning-587b1f2de928</guid>
      <pubDate>Sat, 20 Jun 2026 01:53:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[The Evolution of Cassandra Data Movement at Netflix]]></title>
      <description><![CDATA[<p>By <a href="https://www.linkedin.com/in/guilhermesmi/">Guil Pires</a>, <a href="https://www.linkedin.com/in/jenjprince/">Jennifer Prince</a>, <a href="https://www.linkedin.com/in/josecamachof/">Jose Camacho</a>, <a href="https://www.linkedin.com/in/kenkurzweil/">Ken Kurzweil</a>, <a href="https://www.linkedin.com/in/phanindra-chunduru/">Phanindra Chunduru</a></p><h3>Background</h3><p>In a previous post, we introduced <a href="https://netflixtechblog.medium.com/data-bridge-how-netflix-simplifies-data-movement-36d10d91c313">Data Bridge</a>, a unified management plane for batch Data Movement at Netflix. Historically, several bespoke Data Movement connectors were developed across different engineering organizations to fulfill their specific requirements. Over the last few years, the Data Movement team has started centralizing these offerings through an abstraction that provides a catalog of connectors, along with simple UI and APIs to initiate Data Movement jobs.</p><p>One such case is the Cassandra to Iceberg connector. Apache Cassandra powers mission critical applications at Netflix, including Member, Billing, Recommendations, Subscriptions and many more. These use cases heavily leverage Data Movement to Apache Iceberg for many analytics and operational tasks, and central to this movement was a connector for Cassandra to Iceberg built in-house named Casspactor. As many Cassandra based Data Abstractions emerged, such as <a href="https://netflixtechblog.com/introducing-netflixs-key-value-data-abstraction-layer-1ea8a0a11b30">Key Value</a>, <a href="https://netflixtechblog.com/introducing-netflix-timeseries-data-abstraction-layer-31552f6326f8">Time Series</a> and <a href="https://netflixtechblog.medium.com/high-throughput-graph-abstraction-at-netflix-part-i-e88063e6f6d5">Graph</a> — the need for larger and more complex Data Movement with transformations became more critical to the business.</p><p>Data movements are fundamentally fulfilled by leveraging the existing Cassandra backup infrastructure. Regularly scheduled backups are performed directly on the Apache Cassandra nodes, via a sidecar process managing the upload of all necessary SSTables and associated Metadata files directly into Amazon S3. When a Data Movement job is initiated, the job constructs the specific backup structure it needs by referencing the S3 based metadata, allowing it to precisely locate the SSTable files. The engine then downloads these files, performs the required mutation compaction and processing, and finally writes the fully transformed, compacted data directly into the target Apache Iceberg tables.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*DPkj0pP-y0TntplMdm224Q.png"><figcaption>Image 1: Cassandra Cluster Backups to S3</figcaption></figure><h3>Casspactor: The Engine We Outgrew</h3><p>Casspactor processed roughly 1,200 data movements per day, transferring approximately 3 PB of data from Apache Cassandra into Apache Iceberg tables. It served some of the most critical workloads at Netflix. For years, it worked. Then, two compounding challenges made it clear we needed a fundamentally different architecture.</p><h3>Fragile Metadata Dependencies</h3><p>Before Casspactor could move a single record, it needed to answer a deceptively simple question: <em>which backup exists, is it complete, and what does it contain?</em></p><p>Casspactor assembled this answer from multiple independent systems:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*ppVlrVcA2a3zOu0R6PtuUA.png"><figcaption>Image 2: Casspactor’s Composite View of a Backup</figcaption></figure><p>Each system had its own failure modes, update cadences, and accuracy guarantees. Casspactor’s view of the world was a composite, and composites diverge from reality.</p><p>Metadata fell out of sync with actual backups, causing Casspactor to read stale or incorrect data silently. Routine maintenance on the Cassandra Clusters triggered uncoordinated snapshots, and because Casspactor required all nodes in a region to snapshot at the same clock second, a single node replacement could break data movement for an entire region.</p><p>The fix was hiding in plain sight. The answer to “which backup exists and is it complete?” already lived in the backup storage layer (Amazon S3) itself. By reading metadata directly from the backup files, we could replace the entire dependency chain with a single source of truth.</p><h3>Every Connector Inherited Casspactor’s Limitations</h3><p>Cassandra at Netflix does not just store raw tables. It backs higher level data abstractions, such as <a href="https://netflixtechblog.com/introducing-netflixs-key-value-data-abstraction-layer-1ea8a0a11b30">Key Value</a>, <a href="https://netflixtechblog.com/introducing-netflix-timeseries-data-abstraction-layer-31552f6326f8">Time Series</a>, and others, each with its own data model, access patterns, and semantics. When any of these abstractions needed to move data to Iceberg, they all funneled through Casspactor.</p><p>Every abstraction inherited Casspactor’s constraints:</p><ul><li><strong>Skewed partition failures:</strong> Casspactor could not handle tables with large partitions, a common pattern in Key Value and Time Series workloads. Jobs crashed with out-of-memory errors on some of Netflix’s largest datasets.</li><li><strong>No data model awareness</strong>: Casspactor moved raw Cassandra tables as is. Connectors for Key Value and other abstractions had to bolt on post processing to reconstruct their data models from the raw output — extra cost, extra complexity, and an extra surface for failures.</li><li><strong>Intermediate table bloat</strong>: Casspactor wrote to an intermediate Iceberg table before producing the final output. The Key Value connector added another intermediate table and a snapshots table. Connectors for abstractions on top of Key Value added even more. This compounded into significant storage cost overhead.</li><li><strong>Inability to Time Travel</strong>: by relying on multiple services to compose a backup unit, Casspactor was unable to restore prior backups in the event of cluster Topology or Keyspace schema changes.</li><li><strong>Monolithic design</strong>: Casspactor was built as a single connector, not as an engine. There was no way to build a family of purpose built connectors on a shared foundation.</li></ul><p>We needed something fundamentally different: an engine that reads directly from backups in S3, produces standard Spark DataFrames, and lets each data abstraction build its own connector with full awareness of its data model. One foundation, many connectors.</p><h3>The New Stack: A Layered Architecture</h3><p>The new architecture, built upon the foundation of Apache Cassandra Analytics and the in-house Move Data framework, represents a fundamental shift toward a layered, purpose-built stack designed for reuse and maintainability. This new engine was conceived with clear separation of concerns, moving away from Casspactor’s monolithic design. The architecture is intentionally layered with the foundation being a core S3 reading capability: the Cassandra Analytics Wrapper, which is built on top of the Open Source Cassandra Analytics with Netflix’s internal backup representation and an S3 Client.</p><p>This layer handles the raw data retrieval from backups, translating it into standard Spark DataFrames. Sitting atop this foundation is a “<strong>Connector Factory</strong>” model, via both Java UDFs and transforms which allows individual data abstractions (Key Value, Time Series, others) to build highly optimized, data model aware connectors that process the generic Spark DataFrames, avoiding the need for complex, expensive, and failure-prone post-processing steps. This layered approach ensures that improvements to the core reading engine benefit all connectors, while the connectors themselves are focused solely on data transformation.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*EwqgicjfpASkWIf3nz4q6A.png"><figcaption>Image 3: The new Connector layered stack</figcaption></figure><ul><li><strong>Handles Skewed Partitions:</strong> By moving the mutation compaction and processing to the Executor level within Spark, the new engine can efficiently handle tables with highly skewed or wide partitions, a major pain point for Casspactor. Crucially, this processing occurs without excessive data shuffling, preventing out-of-memory errors and enabling reliable movement of Netflix’s largest datasets.</li><li><strong>Operates at Spark DataFrames (No Intermediary Tables)</strong>: The new architecture directly generates standard Spark DataFrames from the Cassandra backups. This eliminates the need for Casspactor’s costly, multi-stage intermediate Iceberg tables, which led to storage bloat and operational complexity. This native DataFrame operation enables the “Connector Factory” by providing a universal, easily consumable interface for building diverse, model specific connectors.</li><li><strong>Jobs Auto Size:</strong> The engine integrates intelligent auto-sizing capabilities, allowing jobs to dynamically adjust resource consumption based on the source table’s characteristics. This removes the burden of manual tuning from engineering teams, ensuring optimal performance and cost efficiency without sacrificing reliability.</li><li><strong>Reduced Dependencies</strong>: By reading metadata directly from the backup files stored in S3, the new stack removes the fragile, multi-service dependency chain that plagued Casspactor. S3 becomes the single, authoritative source of truth for backup existence and completeness, vastly improving data movement reliability and consistency.</li><li><strong>Time Travel</strong>: A critical feature of the new stack is the ability to process the <strong>schema, cluster topology, and data as a cohesive unit </strong>at a specific point in time. This capability provides robust time travel functionality, essential for auditing, debugging, disaster recovery and reproducing past data states.</li><li><strong>Performance</strong>: Collectively, these architectural improvements, including native DataFrame processing, optimized partition handling, and streamlined metadata retrieval have resulted in notable performance gains, reducing overall data movement execution runtime and cost compared to the legacy Casspactor system.</li><li><strong>Cost</strong>: by eliminating intermediary Iceberg tables and efficient SSTable compaction on Executors, the new stack needs a significantly smaller storage and compute footprint leading to significant cost savings in the order of USD millions.</li></ul><h3>The Journey Towards a Safe Migration</h3><p>The successful validation of the new stack was the critical first step, but it only marked the beginning of the most challenging phase: the migration. Large scale data migrations are inherently complex, high-risk undertakings that can be time consuming and often result in customer frustration and service disruption. To navigate the high stakes of decommissioning a mission-critical system like Casspactor and seamlessly replacing it, we needed a strategy that prioritized reliability and transparency above all else.</p><p>The migration was fundamentally enabled by a <strong>Like-for-Like</strong> strategy, which served as the cornerstone of our Platform Engineering philosophy, abstracting complexity. The core tenet was to maintain absolute consistency across the user-facing interface, the output contract, and the final data artifact. This meant ensuring that the data movement parameters defined via the Data Bridge abstraction remained unchanged, and, critically, the schema, metadata, and data within the destination Iceberg tables were identical to the legacy output. By preserving these external contracts, we eliminated the need for complex, time-consuming coordination with dozens of internal teams who relied on these data pipelines. This approach transformed the migration from a distributed, high-risk, multi-team effort into an internal platform implementation detail, allowing us to achieve a transparent, zero-impact transition and accelerate the retirement of the legacy system without requiring any code changes or validation from downstream users.</p><p>To navigate this migration, we developed a strategy anchored by three core pillars that serve as a blueprint for successful, large-scale data migrations:</p><ol><li><strong>Validation</strong>: Establishing and maintaining absolute confidence in data consistency through rigorous, ongoing validation.</li><li><strong>Visibility</strong>: Instrumenting every part of the system to provide a clear, real-time understanding of migration progress and system health.</li><li><strong>Safety</strong>: Ensuring user impact is minimized or eliminated, despite the inevitable system failures, by leveraging abstractions and robust fallbacks.</li></ol><p>The next section will provide a detailed exploration of these key pillars.</p><h3>Pillar 1: Validation</h3><p>Trust is earned, and in data migration, it is earned one row at a time. The first pillar is the most critical: providing a measurable guarantee to users and partners that the data produced by the new system is an exact, row-by-row replica of the data produced by the old one.</p><p>Our foundational tactic was deploying the new Move Data connector in a “shadow” testing that ran in parallel with the production Casspactor jobs. This allowed us to validate the new system with real-world, production workloads without any customer impact.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*L8yxnLU-7YqzHeEAmVRv6g.png"><figcaption>Image 4: Shadow job structure leveraged for data validation</figcaption></figure><ul><li>Let <strong>C</strong> be the set of rows in the legacy Casspactor output (Iceberg table).</li><li>Let <strong>M</strong> be the set of rows in the new Move Data output (Iceberg table).</li></ul><p>The test for trust: prove that <strong>C = M</strong>. This required continuously checking for two conditions:</p><ol><li>Rows in <strong>C</strong> but not in <strong>M</strong> (<em>C-M</em>): The new system missed data.</li><li>Rows in <strong>M</strong> but not in <strong>C</strong> (<em>M-C</em>): The new system introduced phantom or erroneous data.</li></ol><p>Any result where the cardinality of these difference sets (the number of differing rows) was greater than zero triggered an immediate, high-priority investigation. The target was 100% similarity.</p><h3>Uncovering and Resolving Disparities</h3><p>The shadow mode quickly became a powerful forensic tool, exposing “unknown unknowns”, subtle discrepancies that were not bugs in the new system but rather differences in behavior between the new and old systems. Resolving these was the core work of building trust. For each problem we initiated an investigation log where we captured the details, logs, queries that allowed us to diagnose. Based on the assessment the issues were categorized so that similar differences on other datasets were later resolved affecting many of the shadow pipelines.</p><p>Maintaining an investigation log was critical to organize the outstanding issues and effectively communicate to stakeholders the progress and confidence of the new connector so that we effectively measure the appropriate level of “confidence” to initiate the migration.</p><p>We observed differences in how connectors leverage reference timestamps for Time-to-Live, Consistency Levels, backup selection, and various internal business logic. This continuous, data-driven cycle of discovery and resolution was the mechanism by which we built confidence in the new architecture.</p><h3>Pillar 2: Visibility</h3><p>Trust is built in the background, but an active migration requires real-time insight: Visibility. The second pillar involves instrumenting the system to provide an unambiguous, clear understanding of operational health and migration progress.</p><p>We extended our instrumentation to the overall migration workflow and its dependencies:</p><ul><li><strong>Dashboards</strong>: We created centralized dashboards to track migration status, visualizing the total number of data movements migrated versus those remaining. The dashboards tracked execution status, average runtime, and cost comparisons between the two connectors.</li><li><strong>Dependency Tracking</strong>: Since the new system relied on a new set of APIs to fetch backup metadata, we implemented detailed metrics for failures to keep track of the APIs or dependencies failed.</li><li><strong>Alerting</strong>: Proactive alerts were set up for job failures (Move Data or Casspactor), failures on Move Data that triggered a fallback to Casspactor or any data discrepancy being detected.</li></ul><p>This comprehensive instrumentation allowed the team to be proactive, fix issues as they emerged during the migration, and gain the necessary confidence to accelerate the migration timeline.</p><h3>Pillar 3: Safety</h3><p>Even with perfect data correctness and enhanced visibility, the third pillar, Safety is required for a zero-impact migration. The challenge is ensuring that when a system inevitably fails, the user experience is uninterrupted. Our strategy centered on decoupling the user’s workflow from the underlying connector implementation.</p><h3>Leveraging Abstraction: The Decider Pattern</h3><p>To achieve a transparent swap, we leveraged the <a href="https://github.com/Netflix/maestro?tab=readme-ov-file">Maestro</a> workflow orchestration platform to implement the Decider pattern:</p><ol><li><strong>Data Movement Abstraction:</strong> From a user’s perspective, their Data Movement job definition remained the same.</li><li><strong>The Decider Step</strong>: Internally the workflow responsible to execute the job was modified to include a Decider step. This step took the data movement parameters (source cluster, table name, destination) and invoked a control plane: Connector Controller.</li><li><strong>Connector Controller as the Registry</strong>: The control plane served as the dynamic registry. Based on the migration cohort and the data movement attributes, it determined and reported the appropriate connector to use either Casspactor (legacy) or Move Data (new).</li></ol><p>This abstraction gave our team complete control. We could upgrade or rollback any connector for any data movement instantly by simply updating a configuration in the controller, with zero modification required to the thousands of downstream customer workflows. Crucially, this abstraction guaranteed the critical safety net: a conditional step in the Maestro workflow logic ensured that if the Move Data step fails, it would immediately execute the Casspactor step.</p><p>This pattern would increase the chances that the user’s data movement completes successfully, even if the new connector encountered a bug or transient failure during the initial rollout phases. User impact was completely eliminated; they might see a slightly longer runtime in the event of a failure and fallback, but they would never see a migration failure or suffer from stale data.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*1Qq5sgvtKuLfpP2byfggMw.png"><figcaption>Image 5: The Decider Pattern Implementation via Maestro</figcaption></figure><p>Beyond the workflow, the new system architecture itself was inherently more resilient. By building the new data movement connector on Cassandra Analytics and reading backups directly from S3, we removed fragile dependencies on deprecated internal services.</p><h3>Conclusion</h3><p>The migration from Casspactor to the new, layered architecture built on Cassandra Analytics and the Move Data connector was more than a typical “tech debt” project; it was a fundamental shift in our approach to data movement reliability and scalability at Netflix.</p><p>The legacy system, while serving us well for years, was ultimately constrained by monolithic design, fragile metadata dependencies, and an inability to handle the complexity of modern data abstractions. The new stack resolves these issues by delivering a robust, cost-efficient, and inherently more resilient solution that reads directly from S3, handles wide partitions gracefully, and eliminates costly intermediate tables.</p><p>Our blueprint for the migration, anchored by the three pillars of Validation, Visibility, and Safety, ensured a transparent and high-confidence transition. Through rigorous shadow testing and a data-driven audit framework, we achieved the desired data consistency. Enhanced dashboards and alerting provided the real-time operational insight necessary to manage risk. Most critically, the implementation of the Decider pattern within our workflow abstraction minimized the impact for all downstream users.</p><p>This successful migration validates a core philosophy: by abstracting complexity at the platform level, we can perform large system migrations without burdening our product engineering partners. The new foundation is now ready to support the next generation of Netflix’s data abstractions.</p><h3>Looking ahead</h3><p>This foundational work on the Cassandra Data Movement stack has done more than just replace a legacy system: it has become an accelerator for innovation across the entire Data Movement organization. By providing a reliable, performant engine that standardizes data retrieval into Spark DataFrames, we’ve enabled the rapid development of new, highly optimized connectors. This new “Connector Factory” approach has already delivered a dedicated Key-Value to Iceberg and Time Series connectors, both of which are fully aware of their respective data models, eliminating costly post-processing. This architecture is also paving the way for ambitious new initiatives, including the development of a solution for bulk loading data into Cassandra itself, effectively completing the data movement cycle, and enabling safer fleetwide connector rollout with canaries inspired by the Decider Pattern.</p><p>We are incredibly grateful for the extensive collaboration among the Data Movement, Data Bridge, Online Data Stores, Membership, Billing, Subscriber and Ads platform teams at Netflix; this work simply couldn’t have been accomplished without their partnership!</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=6e13329c80a1" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/the-evolution-of-cassandra-data-movement-at-netflix-6e13329c80a1">The Evolution of Cassandra Data Movement at Netflix</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/the-evolution-of-cassandra-data-movement-at-netflix-6e13329c80a1</link>
      <guid>https://netflixtechblog.com/the-evolution-of-cassandra-data-movement-at-netflix-6e13329c80a1</guid>
      <pubDate>Sat, 20 Jun 2026 01:53:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Thinking Fast & Slow for a Personalized Notification System]]></title>
      <description><![CDATA[<p>by <a href="https://www.linkedin.com/in/matthew-wood-bbb32584/">Matthew Wood</a>, <a href="https://www.linkedin.com/in/ishangupta1993/">Ishan Gupta</a>, Kevin Mercurio, <a href="https://www.linkedin.com/in/devon-bryant-9910ab5/">Devon Bryant</a>, and <a href="https://www.linkedin.com/in/clairedorman/">Claire Dorman</a></p><p>In his seminal book “Thinking, Fast and Slow,” Daniel Kahneman describes two systems that drive human cognition: System 1, which operates automatically and quickly with little effort, and System 2, which allocates attention to more challenging mental activities requiring deliberate focus. This dual-process theory has profound implications not just for understanding human behavior, but for designing intelligent systems that must balance immediate responsiveness with strategic foresight. Similar “plan vs. act” decompositions show up in other domains too — for example, robotics and autonomous driving often separate a slower planning layer (setting goals and constraints over longer horizons) from faster control and execution loops, and modern LLM agents frequently pair deliberate planning with rapid, step-by-step tool use and reaction.</p><p>At Netflix, our messaging platform faces a similar challenge every day. We send hundreds of millions of personalized notifications — push messages, emails, and in-app alerts — to help members discover content they’ll love. This creates a central tension: optimizing each notification for near-term engagement can conflict with what is best for the member over the long term. Higher message frequency can increase fatigue and opt-out risk, while lower frequency can reduce awareness of relevant titles and features the member would value.</p><p>This blog post introduces our framework for personalized notifications — a hierarchical system where a “slow” policy makes strategic, personalized decisions about a member’s weekly messaging plan (e.g., the intended frequency per channel and the resulting pacing over the week), while a “fast” policy handles the tactical, real-time decisions about which specific message to send when a send opportunity occurs. Together, they balance near-term engagement with longer-term member experience.</p><h3>The Problem:</h3><p>Before introducing our new framework, it is helpful to ground the discussion in a representative baseline for a personalized notification system. In our previous production system, we used a causal model to make send decisions by predicting the causal effect of a single message over a short time horizon. While this approach is effective as a baseline, it suffers from two fundamental limitations:</p><h3>Short-Term Reward Horizons</h3><p>The single-message outcome model is trained to optimize short-horizon metrics, such as immediate user actions occurring shortly after a notification is sent. While this is excellent for driving near-term engagement, it misses the cumulative, long-term effects of a messaging strategy. A message that drives an interaction today might also contribute to notification fatigue, reducing responsiveness in the weeks to follow. Because critical indicators of member satisfaction — like sustained viewing habits or gradual opt-out risk — only surface over extended timeframes, a short-term model will always miss the bigger picture.</p><h3>Coupled Ranking and Pacing Decisions</h3><p>When a single system evaluates daily incrementality to decide both whether to send something and, if so, which item to send, an individual member’s weekly message frequency becomes a by-product of those daily decisions rather than an explicit control variable. In our previous single-policy system, frequency was controlled implicitly through a relevance threshold on the model score calibrated to achieve a target aggregate send rate. While effective for managing overall frequency, this mechanism limited the system’s ability to personalize frequency based on individual engagement patterns. Moreover, because send eligibility and message selection were coupled in the same decision rule, adjusting the threshold to control frequency also changed the distribution and quality of selected messages, and vice versa.</p><p>To solve these challenges, we needed a system that could separate longer-term strategy from shorter-term decisions. What if we could determine an optimal, personalized message plan for each member, and then focus on selecting the most relevant content within those bounds? In the following sections, we detail how we realized this vision by decoupling our notification engine into a hierarchical ‘System 1’ and ‘System 2’ framework.</p><h3>The Proposed Method: A Hierarchical Slow-Fast Architecture</h3><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*BffuQUAFrEhYzynBKsVPLA.png"></figure><p>The Slow policy’s primary role is to define a <strong>personalized pacing of messages over a defined time horizon</strong>. The decisions made by slow policy are consumed by the Fast Policy whose role is to maximize immediate relevance and select the optimal message for the member at any given moment.</p><p>To illustrate the Slow Policy in practice: For example, if optimized at a weekly cadence, the policy evaluates a member’s long-term engagement patterns to select a “Pacing Plan Action.” To keep the action space manageable yet expressive, we discretize the decision space into a set of actions that independently specify push and email frequencies. This provides approximately O(100) distinct combinations of cross-channel pacing strategies.</p><p><strong>The Utility Function</strong></p><p>The Slow policy selects actions by maximizing a personalized utility function. This function explicitly trades off positive engagement signals against the long-term “cost” of messaging.</p><p><em>U(member, action) = Σ wₖ·Reward_k(member,action) — Cost(action)</em></p><p>To capture a holistic view of member health, this utility is composed of:</p><ul><li><strong>Positive Signals:</strong> Capturing the likelihood that a member will find value in and engage with the platform.</li><li><strong>Negative Signals:</strong> Capturing the likelihood of member fatigue or a propensity to opt out of a messaging channel.</li></ul><p>Ideally, negative signals alone would naturally penalize over-messaging. In practice, however, explicit negative feedback is extremely sparse. Without an additional constraint, the predicted ‘cost’ of an incremental message appears negligible, causing the model to gravitate toward maximum frequency.</p><p>To address this, we introduce a <strong>universal message cost</strong> that is added to the personalized negative‑feedback prediction for every send. This additional cost term keeps the reward function concave and well‑behaved, preventing degenerate “always send” policies. The message cost parameter is empirically tuned using a combination of online experiments and offline evaluation metrics.</p><p><strong>Pacing Strategy</strong></p><p>The two-stage design naturally allows for optimizing both the average frequency as well as pacing of messages over time. The simplest pacing strategy is uniform random: we translate the frequency target into a per-opportunity send probability and, at each eligible opportunity, effectively flip a weighted coin to decide whether to send. This produces an organically randomized pattern whose expected send rate matches the target.</p><p>While uniform pacing provides a clean and robust baseline, the framework readily extends to richer, non-uniform pacing profiles (for example, day-of-week patterns, conditioning on user activity, or launch-aligned bursts) whenever product or user-experience considerations call for more structured temporal distributions.</p><p><strong>Policy-to-Policy Communication</strong></p><p>The true power of this hierarchy lies in decoupling. By splitting into “Slow” and “Fast” policies, we allow each to focus on what it does best.</p><p>To bridge these two worlds asynchronously, decisions are events and state is managed through a low-latency feature store:</p><ul><li><strong>The Planner (Slow):</strong> The Slow policy calculates a member’s ideal pacing plan. It writes this strategic intent to a feature store</li><li><strong>The Executor (Fast):</strong> Every day, when a notification opportunity arises, the Fast Policy simply pulls that stored “plan” as a feature. It then executes the tactical send decision within those strategic guardrails.</li></ul><p>This architecture provides two critical advantages:</p><ol><li><strong>“Stickiness”:</strong> It ensures a member receives a consistent experience. The Slow policy will be executed once at a defined cadence; the plan is stored and honored.</li><li><strong>Independent Evolution:</strong> We can retrain, optimize, or A/B test our weekly pacing strategies (the “Slow” layer) without ever touching the real-time ranking logic (the “Fast” layer).</li></ol><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*gZiJ6HiK_BG3rDCRupx5vA.png"><figcaption>Figure 1: Schematic of the two-layer message personalization system composed of a slow planning policy (top) and a fast execution policy (bottom). A feature store serves as the communication bridge between the two policies.</figcaption></figure><h3>Key Results &amp; Takeaways</h3><p>The transition to a hierarchical architecture resulted in one of our <strong>largest production metric lifts to date</strong>. We observed several key breakthroughs:</p><ul><li><strong>Empowering the “Casual Viewer”</strong>: Gains were most significant among members who watch less frequently — a critical cohort where timely, high-relevance awareness of new content is vital.</li><li><strong>The Power of Decoupling</strong>: Separating <em>frequency planning</em> from <em>message selection</em> was as transformative as the modeling itself. This new architecture unlocks incredible flexibility, allowing us to iterate on content ranking models and pacing strategies as two independent, clean variables.</li><li><strong>Respecting the Horizon</strong>: The impact of messaging is rarely an isolated event; its effects build up cumulatively based on ongoing interactions between our system and the member. By isolating pacing into a dedicated strategic layer, we now have the mechanism to explicitly manage long-term fatigue and opt-out risk.</li></ul><h3>Acknowledgments</h3><p>We could not have delivered this project without the help of our outstanding colleagues, and we sincerely thank them for their contributions.</p><p><strong>Feature Store Team</strong>: <a href="https://www.linkedin.com/in/aaronlewism/">Aaron Lewis</a>, <a href="https://www.linkedin.com/in/tom-switzer-59824356/">Tom Switzer</a>, <a href="https://www.linkedin.com/in/abbywh/">Abby Whittier</a>, <a href="https://www.linkedin.com/in/ray-zhang-a7168a32/">Ray Zhang</a><br><strong>Product: </strong><a href="https://www.linkedin.com/in/fenglinli/">Fiona Li</a><br><strong>AI for Member Systems (supporting contributor): </strong><a href="https://www.linkedin.com/in/sergipv/">Sergi Perez</a></p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=4d89b26525cd" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/thinking-fast-slow-for-a-personalized-notification-system-4d89b26525cd">Thinking Fast &amp; Slow for a Personalized Notification System</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/thinking-fast-slow-for-a-personalized-notification-system-4d89b26525cd</link>
      <guid>https://netflixtechblog.com/thinking-fast-slow-for-a-personalized-notification-system-4d89b26525cd</guid>
      <pubDate>Sat, 20 Jun 2026 01:53:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[A Human-Augmenting Agentic Workflow for Causal Inference]]></title>
      <description><![CDATA[<p>By Winston Chou, Adrien Alexandre, Lars Olds, Yi Zhang, Garrett Hagemann, and Nathan Kallus</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*7midjxNL3H3BlZvQy4seSA.png"></figure><h3>Introduction</h3><p>Imagine asking a data agent to analyze the causal relationship between two variables, such as the effect of watching a popular Netflix show on long-term member retention. It queries your data, runs a regression, and confidently returns an answer. How much should you trust it? Can you be confident that the agent accounted for subtle biases — or does it treat passionate fans as if they were the average viewer? Without deep understanding and expertise, would you even be able to tell if it got the answer wrong?</p><p>Data analysis is increasingly being delegated to software agents. While this reduces human effort and toil, oversight is still needed to ensure the validity of results. This is especially true for specialized tasks like <a href="https://en.wikipedia.org/wiki/Observational_study">Observational Causal Inference (OCI)</a>, which require substantial judgment and domain expertise.</p><p>In this blog post, we share an agentic workflow for performing OCI under <a href="https://www.nber.org/system/files/working_papers/t0294/t0294.pdf">unconfoundedness</a>. Our workflow is designed for software agents to adhere to rigorous, exhaustive templates for causal inference tasks. Yet, it also seeks to be “<a href="https://pubsonline.informs.org/doi/10.1287/mnsc.2024.05684">human-augmenting</a>,” and to enable and empower human inspection and evaluation.</p><p>We designed this workflow with OCI practitioners in mind. Although OCI requires context and care to do well, aspects of it — e.g., checking and rechecking covariate balance, conducting sensitivity analyses, and keeping track of multiple iterations — can be repetitive and prone to error. Our workflow seeks to eliminate this toil so that humans can focus on more nuanced tasks, such as framing questions, scrutinizing assumptions, and evaluating results.</p><p>To this end, <strong>we are open-sourcing a standalone version of our </strong><a href="https://github.com/Netflix-Skunkworks/oci-agent"><strong>oci-agent</strong></a> so that OCI practitioners can model workflows on and suggest improvements to it. We also share evaluations of our agent on the 2016 Atlantic Causal Inference Conference (ACIC) <a href="https://files.eric.ed.gov/fulltext/ED591944.pdf">competition datasets</a>, and show that our agent systematically beats one-shot iterations under numerous data-generating processes — while achieving competitive results against hand-tuned benchmarks.</p><p>This post describes the principles behind our workflow and gives a case study of its deployment at Netflix.</p><h3>Philosophy</h3><p>Our workflow is built on top of Netflix’s pre-existing OCI toolkit. We built this toolkit — largely in a pre-AI world — to answer “point-in-time” causal questions, such as “what is the effect of playing a Netflix game on member retention?” or “what is the effect of watching a highly popular show on subsequent engagement?” Questions of this kind inform business strategy, guide metric development, and contribute to a rich understanding of what drives member satisfaction.</p><p>Our toolkit is guided by a “<a href="https://academic.oup.com/aje/article-abstract/183/8/758/1739860">target trial emulation</a>” philosophy. For any point-in-time OCI question, we ask “what is the ideal A/B test for addressing this question?” This A/B test may be expensive, slow, or even infeasible to run. However, the thought exercise helps to pin down the assumptions needed for a credible answer, such as unconfoundedness of the treatment.</p><p>To make the target trial analogy actionable, our toolkit embeds a series of <strong>design diagnostics</strong>. These diagnostics assess whether we are drawing fair comparisons between treated and untreated units — or if there are hidden differences that could imperil our conclusions:</p><ul><li><em>Covariate balance.</em> After weighting, the standardized mean difference of pre-treatment covariates between treatment and control groups should be less than 0.2.</li><li><em>Overlap.</em> The probability of receiving treatment (aka propensity score) should be bounded between 0.1 and 0.9.</li><li><em>Placebo outcome.</em> The “treatment effect” on variables measured prior to the treatment should not be significantly different from zero.</li><li><em>Sensitivity to hidden confounders.</em> Findings of treatment effects are contextualized by sensitivity to hypothetical omitted variables that explain both treatment and outcome.</li></ul><p>As we uplevel our OCI toolkit with agents, such evaluation remains paramount. The standard approach to evaluating agents is to programmatically compare their outputs to ground truth. Yet, outside of artificially simulated data, there is <a href="https://en.wikipedia.org/wiki/Rubin_causal_model#The_fundamental_problem_of_causal_inference">no ground truth in observational causal inference</a>.</p><p>Without discounting the need for evals (which our workflow also supports), one of our key principles is to augment <em>human</em> evaluation by making each analytic step as transparent as possible. For example, in our workflow, agents publish artifacts in the form of plans, specifications, plots, and notebooks that humans can inspect and re-execute if they wish. In the absence of ground truth, we rely on these “process audits” — coupled with human oversight — to build good agents.</p><h3>Principles</h3><p>Our workflow has three key personas:</p><ol><li><strong>Principal</strong> — the human user (e.g., data scientist) whose goal is to provide a thorough and correct analysis</li><li><strong>Actor</strong> — the software persona that performs the analysis, including diagnostics</li><li><strong>Critic</strong> — the software persona that synthesizes results, identifies gaps, and offers suggestions to improve the analysis</li></ol><p>Our agent orchestrates the latter two personas in an actor-critic loop: specifying and triggering the analysis as the actor, then interpreting results and diagnosing flaws as the critic.</p><p>Each persona has responsibilities:</p><p><strong>Principals</strong></p><ul><li>Provide an initial analysis plan containing its context and goals.</li><li>Provide context on the main threats to valid inference and the confounders that must be controlled.</li><li>Specify the tools that can be used for the analysis.</li><li>Specify the data model and dataset.</li></ul><p><strong>Actors</strong></p><ul><li>Refine the principal’s plan into a data analysis spec.</li><li>Use only the tools provided by the principal.</li><li>Create human- and machine-checkable artifacts.</li><li>Perform the four design diagnostics in addition to the core analysis.</li><li>Report any remediations taken in case of diagnostic failures.</li></ul><p><strong>Critics</strong></p><ul><li>Check for blind spots, such as unmentioned confounders, in the principal’s plan.</li><li>Check for alignment between the plan, spec, and executed analysis.</li><li>Specify a credibility level in the results after seeing the diagnostics.</li><li>Specify if and how the estimand differs from the Average Treatment Effect (ATE), for example due to propensity score trimming.</li><li>Contrast the executed analysis with the ideal target Randomized Controlled Trial (RCT).</li><li>Suggest at least one alternative measurement strategy (e.g., encouragement RCTs).</li></ul><p>Although our workflow is designed for OCI under unconfoundedness, the principles listed in this section are meant to be extensible to other approaches to OCI, such as panel methods with very different assumptions (e.g., parallel trends).</p><h3>Empowering Human Evaluation</h3><p>To empower human oversight of each analytic step, we provide principals with a templated notebook that uses our vetted (non-agentic) OCI toolkit, which employs <a href="https://arxiv.org/abs/2506.07462">doubly robust learning for causal effect estimation</a>.</p><p>The principal’s remaining responsibilities are to write the initial analysis plan and to evaluate the analysis artifacts (the executed notebook and the critic’s report). To enable thorough evaluation, agents version-control their reports and upload executed notebooks to a file store, where they can be downloaded and re-executed by principals (if they wish).</p><p>We diagram this workflow below:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*1cTcjXY_Gh80OB2x1xzlAQ.png"></figure><h3>Case Study — Estimating the Impact of New Entertainment Types</h3><p>In recent years, we have added a wide variety of entertainment types beyond streaming video to Netflix. A natural question is how these new entertainment types affect members’ satisfaction and their likelihood of continuing to subscribe to Netflix.</p><p>To analyze the impact of one of these new entertainment types, which we will call Type X, we wrote a simple analysis plan specifying our</p><ul><li><strong>Treatment</strong>: Days engaging with Type X (or “Type X days” for short)</li><li><strong>Outcome</strong>: Two-month retention</li><li><strong>Potential</strong> <strong>confounders</strong>, including pre-treatment Type X days</li></ul><p>To establish a baseline, we fed this analysis plan without additional scaffolding to Claude Sonnet 4.6, a powerful yet accessible general-purpose model. The model chose and executed a defensible analysis strategy: linearly regressing retention on Type X days along with controls.</p><p>While the result was polished and impressive, when we ran the same analysis through our paved path tooling and agentic workflow, also using Sonnet 4.6, our agent produced an updated estimate that was just 25% of the baseline! What explains the difference between the baseline and the paved-path estimates?</p><p>A core challenge when analyzing new entertainment types is <strong>early adopter bias</strong>. The first users of any new offering are likely to be systematically different from the general population. For example, they may be heavier users of Netflix generally, or they may be extremely strong fans of the underlying titles. Early adopter bias manifested in our analysis as poor “overlap”: the vast majority of observations had a small estimated probability of engaging with Type X, reflecting its early maturity.</p><p>This imbalance was caught by our critic agent in its writeup of the analysis. The critic also flagged the failure of the placebo test: early Type X adopters differed significantly from non-adopters in terms of important confounders measured <em>before</em> experiencing the treatment, a warning sign of potential bias.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*Ee0gswtciDVvuFr1QenlvQ.png"></figure><h3>Addressing Failed Diagnostics</h3><p>To address these diagnostic failures, our workflow provides agents with a playbook. For example, to overcome poor overlap, we instruct the agent to use <a href="https://academic.oup.com/biomet/article-abstract/96/1/187/235329?redirectedFrom=fulltext">Crump-style trimming</a>. That is, before estimating causal effects, the actor trims units with estimated propensity scores outside the range [0.1, 0.9]. This scopes the treatment effect being estimated to the ATE in the population that is not very likely or unlikely to engage in the new entertainment type — an important caveat we instruct the critic to flag in its report.</p><p>Trimming yields an estimate that is much smaller than the baseline estimate, and which only applies to the “overlapping” population (for whom engagement with the new entertainment type is non-deterministic). However, the trimmed estimate is substantially more credible, as it focuses on the members for whom the treatment could plausibly be randomly assigned, as in a target trial.</p><p>Contrastively, the baseline effect relies heavily on assumptions to extrapolate outcomes for all members, even those with a very low probability of treatment. The <a href="https://gking.harvard.edu/files/counterft.pdf?utm_source=chatgpt.com">danger</a> here is that extrapolation produces a number that is not backed by robust data and is likely confounded by early adopter bias.</p><h3>Orchestrating Followup Analyses</h3><p>There are two natural followups to this analysis:</p><ol><li>First, we need to analyze the sensitivity of estimates to the choice of trimming threshold. Practically, this requires redoing the analysis with multiple trimming thresholds.</li><li>Second, we also care about how these causal effects evolve over time. Yet, comparing causal effects across time raises subtle challenges. For example, we need to coordinate the population across all analyses: if a set of users is trimmed to make one analysis more credible, it should be trimmed in the other analyses as well.</li></ol><p>Both of these followups require conducting multiple versions of the same analysis, tweaking some parameters while keeping others the same. Managing this complexity and ensuring consistent execution is another area where agents add value.</p><p>To illustrate this, below we show a sensitivity analysis for our case study in which we asked the agent to vary the trimming bounds from [0, 1] (no trimming) to [0.15, 0.85]. As the plot shows, the estimated ATE on the overlapping population is robust to the choice of trimming threshold within bounds of [0.005, 0.995]. Although principals could (and should) execute <a href="https://psicostat.github.io/4ms-winter-school/papers/Steegen%20et%20al.%202016%20-%20Increasing%20Transparency%20Through%20a%20Multiverse%20Analysis.pdf">this and other robustness analyses</a>, delegating them to agents helps to reduce toil.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*GbQUhS5LB9gcn16VDnC69w.png"></figure><p>Another example is generating a time series by repeating the same analysis across multiple date partitions. For example, below we plot the results of using our agent to refit a different analysis on ten distinct date partitions. The plot shows evidence of seasonality: the treatment has a stronger effect on the winter dates compared to the summer dates.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*792mstFbTU-1ZVSs7KJEVw.png"></figure><h3>Public Repo and Evals</h3><p>To help OCI practitioners build on and contribute to our workflow, we are open-sourcing a standalone version of <a href="https://github.com/Netflix-Skunkworks/oci-agent">oci-agent</a>. This repo implements two evaluations on public datasets from the 2016 Atlantic Causal Inference Competition (ACIC) data analysis competition. It also includes a lightweight version of our internal causal machine learning notebook that only uses open-source software (<a href="https://github.com/py-why/EconML">EconML</a>).</p><p>Our first evaluation runs this notebook for three randomly sampled datasets generated by each of the 77 data-generating processes (DGPs) in the ACIC data. Next, it uses the critic to grade the resulting 231 estimates as either satisfactory or unsatisfactory based on the diagnostics.</p><p>Below, we plot the average RMSE and coverage of 95% confidence intervals of our ATT estimates against the 44 competitor methods in the ACIC competition. As the scatterplot shows, our statistical methodology is competitive against these benchmarks: it achieves reasonably low RMSE and well-calibrated confidence intervals that cover the truth in ~95% of DGPs.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*nvNVibkHRO8EgqNcliB_Zg.png"></figure><p>More to the point, our diagnostics and agentic workflow help to separate more reliable estimates from less reliable estimates. To illustrate this, the following chart plots our ATE estimates in terms of RMSE and coverage. Note that we separate out the RMSE and coverage of:</p><ul><li>All 231 estimates (purple dot)</li><li>The 192 satisfactory estimates (blue star)</li><li>The 39 unsatisfactory estimates (red dot)</li></ul><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*SGM_i82SxCI64iMIp-diFw.png"></figure><p>As the plot shows, when aided by our diagnostic suite, the critic agent is able to separate good estimates from bad estimates: the satisfactory estimates have <em>much </em>lower RMSE and better calibrated confidence intervals than do the unsatisfactory estimates.</p><p>Our second evaluation compares the performance of an LLM on the same analysis plan with our scaffolding and without it (i.e., one-shot prompting). Unsurprisingly, we find that our scaffolding is critical to helping the LLM return useful estimates. This can be seen in the following random sample of ten ACIC datasets. Using our scaffolding, the LLM recovers the ground truth in nine out of ten datasets. Furthermore, estimates are highly correlated with ground truth.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*k105qpwAZ3bjY_kLvAEpUA.png"></figure><p>In contrast, giving the same analysis plan to Sonnet 4.6 without any scaffolding (i.e., just prompting it) results in consistently wrong answers that are not at all correlated with ground truth.</p><p>A key limitation of our public repo is that, due to the synthetic nature of the underlying datasets, it doesn’t pressure-test our agent’s semantic understanding or performance on real-world OCI tasks. Nonetheless, the repo demonstrates the core principles underlying our workflow. These include (1) giving agents with extensive scaffolding so that they follow best practices by design, and (2) requiring inspectable artifacts so that humans can audit agents’ <em>processes,</em> not just their outcomes.</p><h3>Conclusion</h3><p>We provide a workflow for doing observational causal inference with the help of software agents. Leveraging elements of our pre-AI OCI toolkit, such as templated notebooks, our workflow is designed to ensure that agents conduct rigorous and exhaustive analyses. This helps to reduce the human toil of OCI, which can be a highly iterative and exacting process.</p><p>At the same time, motivated by the complexity and ambiguity of observational causal inference, our workflow seeks to be <strong>human-augmenting</strong> and enables human practitioners to evaluate each analytic step.</p><p>Using agents for causal inference poses a challenge: how do we evaluate agents’ performance on tasks without ground truth? To meet this challenge, our workflow combines process audits with human oversight. To enable others to learn from and critique our workflow, we have <a href="https://github.com/Netflix-Skunkworks/oci-agent">open-sourced</a> a lightweight, standalone version. We hope this work stimulates more research and development on agentic evaluation in the absence of ground truth.</p><p><em>For valuable feedback on this post and “dogfooding,” we thank Adith Swaminathan, Ayal Chen-Zion, Colin Gray, Juliet Hougland, and Simon Ejdemyr.</em></p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=4623f0a9c5af" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/a-human-augmenting-agentic-workflow-for-causal-inference-4623f0a9c5af">A Human-Augmenting Agentic Workflow for Causal Inference</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/a-human-augmenting-agentic-workflow-for-causal-inference-4623f0a9c5af</link>
      <guid>https://netflixtechblog.com/a-human-augmenting-agentic-workflow-for-causal-inference-4623f0a9c5af</guid>
      <pubDate>Sat, 20 Jun 2026 01:53:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[VMAF v1: Good Is Not Good Enough]]></title>
      <description><![CDATA[<p><em>By </em><a href="https://www.linkedin.com/in/christosbampis/"><em>Christos G. Bampis</em></a><em>, </em><a href="https://www.linkedin.com/in/henryzhili/"><em>Zhi Li</em></a><em>, Kyle Swanson, </em><a href="https://www.linkedin.com/in/nil-fons-miret/"><em>Nil Fons Miret</em></a><em> and </em><a href="https://www.linkedin.com/in/pavan-chennagiri/"><em>Pavan Madhusudanarao</em></a></p><p><em>Will this encode look good to Netflix members? Does switching to a new codec improve quality at the same bitrate and by how much? What is the best way to encode a movie title given a target bitrate budget? For years, VMAF has reliably helped us answer those questions and deliver an optimized quality of experience to our members.</em></p><p><em>But </em><a href="https://jobs.netflix.com/culture"><em>good is not good enough</em></a><em>. If VMAF misjudges quality, that may lead to loss of detail for a suspenseful close-up or banding for a stunning wide-angle sky shot. That’s a lot of trust to put in one number, so we strive to make sure it earns it. Over time, we collected feedback from VMAF users, both internally and externally. A few years ago, we embarked on a journey to develop a new version of VMAF to address some of its known limitations. Today, we are happy to announce that we are open-sourcing a new version of VMAF, with version number v1. By using VMAF v1 we can more accurately assess visual quality and hence efficiently deliver higher quality for Netflix members worldwide. In this post we share how v1 addresses the previous version’s (called VMAF v0) limitations and some of the challenges we faced along the way.</em></p><h4>What is VMAF and why improve it?</h4><p>VMAF (Video Multimethod Assessment Fusion) is a video quality metric that Netflix developed with university partners and open-sourced on GitHub. It has become a de facto standard for encoding evaluation and optimization for the video industry. VMAF combines elementary quality-aware features and fuses them with a support-vector regressor (SVR) trained on subjective data. For background, see our <a href="https://netflixtechblog.com/toward-a-practical-perceptual-video-quality-metric-653f208b9652">first</a>, <a href="https://netflixtechblog.com/vmaf-the-journey-continues-44b51ee9ed12">second</a> and <a href="https://netflixtechblog.com/toward-a-better-quality-metric-for-the-video-community-7ed94e752a30">third</a> VMAF tech blogs.</p><p>Despite its accuracy and wide adoption, we have identified room to improve the core of the algorithm. That’s central to our mission of delivering the best possible visual quality to our members no matter where and how they watch Netflix. As new codecs, like <a href="https://av2.aomedia.org/">AV2</a>, are developed and use cases like live streaming and cloud gaming emerge, we strive to continue to improve VMAF to serve these business needs. We describe each key improvement below.</p><h4>Improving sensitivity to compression artifacts</h4><p>As discussed in our <a href="https://netflixtechblog.com/toward-a-practical-perceptual-video-quality-metric-653f208b9652">first VMAF tech blog</a> [1], a typical encoding pipeline introduces both compression and scaling artifacts. Intuitively, when more bits are available, higher resolutions are preferable. VMAF quantifies the tradeoff between compression and scaling and determines the optimal resolution to use given a bitrate budget. This can be demonstrated by a VMAF vs. bitrate curve.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*4Ymg7tqfTpT3SUSQX92smw.png"></figure><p>In practice, we observed that VMAF v0 tends to favor switching to a higher resolution at lower bitrates, preferring compression artifacts over scaling, which could be visually annoying. This can be partially attributed to the DLM (Detail Loss Metric) feature, which penalizes contrast/detail loss, but may be less sensitive to distracting artifacts, like blockiness [3]. In VMAF v1, to complement DLM, we added the AIM (additive impairments) component [3] from the original ADM formulation with minor modifications to improve accuracy. These two elementary metrics are linearly combined, similar to the original implementation in [3].</p><h4>One VMAF model to rule them all</h4><p>A first-order effect that influences quality perception is the visibility of artifacts and its relationship to viewing distance and canvas size. Put simply, the same encoded video looks better when displayed on a smaller canvas or viewed from further away.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*WG0WwDOjpT9u5ryel5HZ6g.png"></figure><p>The standard VMAF model assumes that viewers sit in front of a 1920×1080 display, in a living room-like environment, with a normalized viewing distance of approximately 3× the screen height (3H). This means that the standard 1080p@3H VMAF model corresponds to a viewing angle of approximately 60 pixels per degree.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*eeifY1bgBb6Fo1TPK89c7g.png"></figure><p>For phone viewing, given a smaller screen size and longer natural viewing distance (typical phone viewing can be approximated as 4 to 5H) relative to the screen height, we expect that artifacts become less visible. The phone model of VMAF v0 captures this by post-processing the standard (TV/laptop) VMAF score by a second-order polynomial mapping. This mapping was estimated using subjective data.</p><p>One drawback of the above mapping is that it is hard to generalize predictions for the myriad of viewing conditions that materially differ from the original subjective experiment. Further, in practice, we observed that the phone model can overpredict quality. In v1, instead of using a mapping function, we adjust the elementary feature values based on the normalized viewing distance. The same model can then be trained and reapplied for different use cases, e.g., phone viewing, 4K@3H, or a more discerning 4K@1.5H. We found that this approach improves accuracy and helps generalize VMAF better.</p><p>To achieve this, we modulate the spatial contrast sensitivity function (CSF) used in DLM based on the normalized viewing distance. The CSF defines human sensitivity to contrast across spatial frequencies and is related to distortion perceptibility. The CSF can hence be used to estimate perceived distortion for different viewing distances, display sizes, and resolutions. An example CSF curve is shown below.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*k-rD3vXE9yGfQxc-k3xUXw.png"></figure><p>As the viewing distance increases, more pixels fit into a degree of visual angle, which lowers distortion visibility. In VMAF v1, we use an adapted version of Barten’s CSF model from [4].</p><h4>Addressing banding artifacts</h4><p>Banding shows up as staircase-like edges in parts of the image that should look smooth. It can have a negative visual impact for viewers, but this impact is not captured well by VMAF v0. VMAF v1 integrates the Contrast Aware Multiscale Banding Index (CAMBI) as one of the elementary features. You can read more about CAMBI in our <a href="https://netflixtechblog.com/cambi-a-banding-artifact-detector-96777ae12fe2">previous tech blog</a> or the <a href="https://github.com/Netflix/vmaf/blob/master/resource/doc/papers/CAMBI_PCS2021.pdf">technical paper</a>.</p><h4>Addressing chroma artifacts</h4><p>VMAF v0 only extracts luma-based features, so it is unaware of chroma artifacts. In practice, encoding and scaling introduce chroma artifacts via quantization and subsampling. To capture such artifacts, we modified <a href="https://ieeexplore.ieee.org/document/7979533">SpEED-QA</a> and applied it to the chroma channels.</p><h4>Leveraging the no-enhancement gain (NEG) mode</h4><p>To reduce the effect of image enhancement operations, like sharpening, a standalone <a href="https://docs.google.com/document/d/1dJczEhXO0MZjBSNyKmd3ARiCTdFVMNPBykH4_HMPoyY/edit?tab=t.0#heading=h.oaikhnw46pw5">no-enhancement gain (NEG)</a> mode was made available for VMAF v0. We have found that NEG serves as a conservative quality metric and helps preserve creative intent. We already use VMAF-NEG as one of the quality metrics during codec development, such as for AV2. Therefore, NEG is enabled by default for VMAF v1 without a need for a separate model.</p><h4>Improving the motion feature</h4><p>VMAF v0’s motion feature does not have an upper bound. Further, the training data used then did not have enough coverage for high-motion sequences. Consequently, we observed that VMAF v0 could overpredict quality for very high-motion scenes. On the flip side, since motion differencing in v0 was performed between consecutive frames, v0 would underpredict quality for sequences with frame rates higher than 24 or 30 fps, like 60 fps. In v1, we apply an empirically derived hard threshold to the motion feature. Further, we add an option to measure motion differences over a larger temporal window. Expanding the temporal window alone does not fully capture the perceptual impact of 60 fps, but it does reduce the underprediction evident in v0.</p><h4>Overview of VMAF v1 models</h4><p>VMAF v1 supports the following models:</p><ul><li><strong>Standard 1080p Model:</strong> This model is calibrated for 1080p video viewed at a standard 3H distance. It uses an operating range of [0, 100].</li><li><strong>Phone Model:</strong> Derived by setting the normalized viewing distance to 5H (based on experimental data), this model adjusts the DLM, AIM, and chroma feature calculations to reflect reduced artifact visibility on smaller screens viewed from a greater relative distance. It retains the standard [0, 100] range.</li><li><strong>4K Model:</strong> We release two v1 4K models: a <strong>1.5H variant</strong> and a <strong>3H variant</strong>. The 1.5H variant is based on a discerning 4K@1.5H viewing condition. This variant is conceptually similar to its v0 4K counterpart and operates on a [0, 100] range. For most users, this variant is the default choice. The 3H variant is based on a consumer-like 4K@3H viewing condition. This variant operates on a [0, 110] range, which helps to quantify the additional perceptual benefit of 4K resolution over 1080p when both are viewed at 3H.</li></ul><h4>Interpreting the score</h4><p>VMAF v1’s score and interpretation are largely consistent with v0’s. To achieve this, we calibrated the VMAF v1 scale to align with v0 via a score transform, so that the new algorithm preserves the meaning of the numbers while keeping its accuracy benefits.</p><h4>Putting v1 to the test</h4><p>We evaluate VMAF v1 across several subjective datasets. These datasets cover a variety of codecs, content types, and use cases. For simplicity, we report the Spearman’s rank correlation coefficient (SRCC) for VMAF v0 and v1. SRCC values closer to 1 show higher agreement with subjective data. Full results will be available in a future technical paper.</p><p>In the table below, if a dataset is marked as “4K” then we measure VMAF at 4K using the 4K@1.5H model, otherwise we measure VMAF at 1080p, using the appropriate 1080p model. If a dataset is marked as “phone” the 1080p phone model is used, otherwise the standard 1080p@3H model is used.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*fyP8kVEUuoYlgoNXD3SKDQ.png"><figcaption>NFLX Banding + compression: Contains 83 AV1 videos encoded at different bitrates under banding-relevant conditions. NFLX How Low Phone: Contains 80 AVC videos encoded at 540p and below to understand viewer acceptability to low qualities. “train-only”: only the training split was publicly available.</figcaption></figure><p>As seen above, VMAF v1 matches or outperforms v0 in most datasets. There are notable improvements on large datasets like WATERLOO IVC 4K and the Netflix Screen Size Crowdsourcing, on datasets with chroma and banding artifacts, and those that involve phone viewing. On a few datasets we observe minor regressions, which are small relative to the gains elsewhere.</p><h4>Running v1</h4><p>Even with the addition of new features, we wanted VMAF v1 to have a reduced computational complexity when compared with VMAF v0. To achieve this, we have:</p><ol><li>Removed VIF (Visual Information Fidelity) as a core VMAF feature. VIF is computationally complex and did not meaningfully improve accuracy after updating the other features.</li><li>Introduced a few CAMBI-specific optimizations, both algorithmic and <a href="https://github.com/Netflix/vmaf/pull/1522">software</a>.</li><li>Measured the chroma feature at a lower scale, which does not hurt accuracy [11].</li></ol><p>The result of this work is not only a more accurate VMAF, but also a much faster VMAF. The table below shows the processing speed and threading performance for each VMAF model at 1080p and 4K. Additionally, our newest libvmaf release has a much improved threading performance, which is of benefit to both v0 and v1. Note that for content with significant banding, computing CAMBI adds some overhead, which can reduce this speedup.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*z5aAwF8Z8Bi4YPL3bFpYcA.png"><figcaption>The ducks_take_off sequence is available <a href="https://media.xiph.org/video/derf/">here</a>.</figcaption></figure><h4>Keeping some old (and good) habits</h4><p>While the core of the algorithm has changed, we still recommend:</p><ol><li>Computing VMAF at the right resolution by upsampling the distorted video to the source resolution, so that both compression and scaling artifacts are reflected. For example, bicubic upsampling can be used as a <a href="https://netflixtechblog.com/vmaf-the-journey-continues-44b51ee9ed12">general approximation</a>.</li><li>Interpreting scores in context by using the right model (e.g. 1080p vs. 4K vs. phone) for your scenario by taking into account viewing distance assumptions.</li></ol><h4>This is not the end of the road</h4><p>Just like any metric, VMAF v1 is not perfect. There is still room for improvement. Some areas that we are working on addressing in the future are: film-grain noise, improved handling for high frame rates, and perceptual codec optimizations such as adaptive quantization. We invite the community to try the latest VMAF models, report edge cases, contribute to the open-source code, and help improve VMAF. We plan to publish a detailed technical paper for VMAF v1. We also plan to release an HDR version enhanced by the v1 improvements, so stay tuned!</p><h4>Acknowledgments</h4><p>This was a collaborative effort propelled by our stunning colleagues. We want to thank the following individuals: Xiaoqing Zhu, Mariana Afonso, Anush Moorthy, Raymond Walsh, Omair Akhtar, Amelia Taylor, Ken Thomas, Zheng Lu, Chris Pham, Alex Chang, Prudhvi Chaganti, Ben Wallen, Craig Howland, Deepthi Arun, Andy Rhines and Lukáš Krasula.</p><h4>References</h4><p>[1] Z. Li et al., “<a href="https://netflixtechblog.com/toward-a-practical-perceptual-video-quality-metric-653f208b9652">Toward a practical perceptual video quality metric</a>,” Netflix Technology Blog, 2016.</p><p>[2] Z. Li et al., “<a href="https://netflixtechblog.com/vmaf-the-journey-continues-44b51ee9ed12">VMAF: The journey continues</a>,” Netflix Technology Blog, 2018.</p><p>[3] S. Li, F. Zhang, L. Ma, and K. Ngan, “Image Quality Assessment by Separately Evaluating Detail Losses and Additive Impairments,” IEEE Transactions on Multimedia, 2011.</p><p>[4] P. G. J. Barten, Contrast Sensitivity of the Human Eye and Its Effects on Image Quality. SPIE Press, 1999.</p><p>[5] Video Quality Experts Group, “Report on the validation of video quality models for high definition video content,” 2010.</p><p>[6] N. Barman, Y. Reznik, and M. G. Martini, “A Subjective Dataset for Multi-Screen Video Streaming Applications,” International Conference on Quality of Multimedia Experience (QoMEX), Ghent, Belgium, 2023, pp. 270–275.</p><p>[7] A. Katsenou, F. Zhang, M. Afonso, G. Dimitrov, and D. R. Bull, “BVI-CC: A Dataset for Research on Video Compression and Quality Assessment,” Frontiers in Signal Processing, vol. 2, 2022.</p><p>[8] C. G. Bampis, L. Krasula, Z. Li, and O. Akhtar, “Measuring and Predicting Perceptions of Video Quality Across Screen Sizes with Crowdsourcing,” International Conference on Quality of Multimedia Experience (QoMEX), Ghent, Belgium, 2023, pp. 13–18.</p><p>[9] J. Y. Lin, R. Song, C.-H. Wu, T. J. Liu, H. Wang, and C.-C. J. Kuo, “MCL-V: A streaming video quality assessment database,” Journal of Visual Communication and Image Representation, vol. 30, pp. 1–9, Jul. 2015.</p><p>[10] Z. Li, Z. Duanmu, W. Liu, and Z. Wang, “AVC, HEVC, VP9, AVS2 or AV1? — A Comparative Study of State-of-the-Art Video Encoders on 4K Videos,” Int. Conf. Image Analysis and Recognition (ICIAR), 2019.</p><p>[11] C. G. Bampis, P. Gupta, R. Soundararajan, and A. C. Bovik, “SpEED-QA: Spatial Efficient Entropic Differencing for Image and Video Quality,” IEEE Signal Process. Lett., vol. 24, no. 9, pp. 1333–1337, Sep. 2017.</p><p>[12] L.-H. Chen, C. G. Bampis, Z. Li, J. Sole, and A. C. Bovik, “Perceptual video quality prediction emphasizing chroma distortions,” IEEE Trans. Image Process., vol. 30, pp. 1941–1954, 2021.</p><p>[13] C. G. Bampis et al., “<a href="https://docs.google.com/document/d/1Ly-r196i6ekIMATuA_qKZCSNpGoKj4shqhyXU8gbSnQ/edit?tab=t.0">NFLX Screen Size Crowdsourcing dataset</a>,” 2023.</p><p>[14] H. Wei, P. Lebreton, Y. Chen, J. Zhu, and P. Le Callet, “<a href="https://sites.google.com/view/qomex26-vqm-gc/">Grand Challenge on Video Quality Assessment for Asymmetric Encoded Videos</a>,” Int. Conf. Quality of Multimedia Experience (QoMEX), Cardiff, U.K., 2026.</p><p>[15] Y. Chen, B. Chen, H. Wei, A. C. Bovik et al., “<a href="https://arxiv.org/abs/2506.22790">ICME 2025 Generalizable HDR and SDR Video Quality Measurement Grand Challenge</a>,” arXiv:2506.22790, 2025.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=60d7e4244ea8" width="1" height="1" alt=""><hr><p><a href="https://medium.com/netflix-techblog/vmaf-v1-good-is-not-good-enough-60d7e4244ea8">VMAF v1: Good Is Not Good Enough</a> was originally published in <a href="https://medium.com/netflix-techblog">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://medium.com/netflix-techblog/vmaf-v1-good-is-not-good-enough-60d7e4244ea8</link>
      <guid>https://medium.com/netflix-techblog/vmaf-v1-good-is-not-good-enough-60d7e4244ea8</guid>
      <pubDate>Sat, 20 Jun 2026 01:50:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[VMAF v1: Good Is Not Good Enough]]></title>
      <description><![CDATA[<p><em>By </em><a href="https://www.linkedin.com/in/christosbampis/"><em>Christos G. Bampis</em></a><em>, </em><a href="https://www.linkedin.com/in/henryzhili/"><em>Zhi Li</em></a><em>, Kyle Swanson, </em><a href="https://www.linkedin.com/in/nil-fons-miret/"><em>Nil Fons Miret</em></a><em> and </em><a href="https://www.linkedin.com/in/pavan-chennagiri/"><em>Pavan Madhusudanarao</em></a></p><p><em>Will this encode look good to Netflix members? Does switching to a new codec improve quality at the same bitrate and by how much? What is the best way to encode a movie title given a target bitrate budget? For years, VMAF has reliably helped us answer those questions and deliver an optimized quality of experience to our members.</em></p><p><em>But </em><a href="https://jobs.netflix.com/culture"><em>good is not good enough</em></a><em>. If VMAF misjudges quality, that may lead to loss of detail for a suspenseful close-up or banding for a stunning wide-angle sky shot. That’s a lot of trust to put in one number, so we strive to make sure it earns it. Over time, we collected feedback from VMAF users, both internally and externally. A few years ago, we embarked on a journey to develop a new version of VMAF to address some of its known limitations. Today, we are happy to announce that we are open-sourcing a new version of VMAF, with version number v1. By using VMAF v1 we can more accurately assess visual quality and hence efficiently deliver higher quality for Netflix members worldwide. In this post we share how v1 addresses the previous version’s (called VMAF v0) limitations and some of the challenges we faced along the way.</em></p><h4>What is VMAF and why improve it?</h4><p>VMAF (Video Multimethod Assessment Fusion) is a video quality metric that Netflix developed with university partners and open-sourced on GitHub. It has become a de facto standard for encoding evaluation and optimization for the video industry. VMAF combines elementary quality-aware features and fuses them with a support-vector regressor (SVR) trained on subjective data. For background, see our <a href="https://netflixtechblog.com/toward-a-practical-perceptual-video-quality-metric-653f208b9652">first</a>, <a href="https://netflixtechblog.com/vmaf-the-journey-continues-44b51ee9ed12">second</a> and <a href="https://netflixtechblog.com/toward-a-better-quality-metric-for-the-video-community-7ed94e752a30">third</a> VMAF tech blogs.</p><p>Despite its accuracy and wide adoption, we have identified room to improve the core of the algorithm. That’s central to our mission of delivering the best possible visual quality to our members no matter where and how they watch Netflix. As new codecs, like <a href="https://av2.aomedia.org/">AV2</a>, are developed and use cases like live streaming and cloud gaming emerge, we strive to continue to improve VMAF to serve these business needs. We describe each key improvement below.</p><h4>Improving sensitivity to compression artifacts</h4><p>As discussed in our <a href="https://netflixtechblog.com/toward-a-practical-perceptual-video-quality-metric-653f208b9652">first VMAF tech blog</a> [1], a typical encoding pipeline introduces both compression and scaling artifacts. Intuitively, when more bits are available, higher resolutions are preferable. VMAF quantifies the tradeoff between compression and scaling and determines the optimal resolution to use given a bitrate budget. This can be demonstrated by a VMAF vs. bitrate curve.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*4Ymg7tqfTpT3SUSQX92smw.png"></figure><p>In practice, we observed that VMAF v0 tends to favor switching to a higher resolution at lower bitrates, preferring compression artifacts over scaling, which could be visually annoying. This can be partially attributed to the DLM (Detail Loss Metric) feature, which penalizes contrast/detail loss, but may be less sensitive to distracting artifacts, like blockiness [3]. In VMAF v1, to complement DLM, we added the AIM (additive impairments) component [3] from the original ADM formulation with minor modifications to improve accuracy. These two elementary metrics are linearly combined, similar to the original implementation in [3].</p><h4>One VMAF model to rule them all</h4><p>A first-order effect that influences quality perception is the visibility of artifacts and its relationship to viewing distance and canvas size. Put simply, the same encoded video looks better when displayed on a smaller canvas or viewed from further away.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*WG0WwDOjpT9u5ryel5HZ6g.png"></figure><p>The standard VMAF model assumes that viewers sit in front of a 1920×1080 display, in a living room-like environment, with a normalized viewing distance of approximately 3× the screen height (3H). This means that the standard 1080p@3H VMAF model corresponds to a viewing angle of approximately 60 pixels per degree.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*eeifY1bgBb6Fo1TPK89c7g.png"></figure><p>For phone viewing, given a smaller screen size and longer natural viewing distance (typical phone viewing can be approximated as 4 to 5H) relative to the screen height, we expect that artifacts become less visible. The phone model of VMAF v0 captures this by post-processing the standard (TV/laptop) VMAF score by a second-order polynomial mapping. This mapping was estimated using subjective data.</p><p>One drawback of the above mapping is that it is hard to generalize predictions for the myriad of viewing conditions that materially differ from the original subjective experiment. Further, in practice, we observed that the phone model can overpredict quality. In v1, instead of using a mapping function, we adjust the elementary feature values based on the normalized viewing distance. The same model can then be trained and reapplied for different use cases, e.g., phone viewing, 4K@3H, or a more discerning 4K@1.5H. We found that this approach improves accuracy and helps generalize VMAF better.</p><p>To achieve this, we modulate the spatial contrast sensitivity function (CSF) used in DLM based on the normalized viewing distance. The CSF defines human sensitivity to contrast across spatial frequencies and is related to distortion perceptibility. The CSF can hence be used to estimate perceived distortion for different viewing distances, display sizes, and resolutions. An example CSF curve is shown below.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*k-rD3vXE9yGfQxc-k3xUXw.png"></figure><p>As the viewing distance increases, more pixels fit into a degree of visual angle, which lowers distortion visibility. In VMAF v1, we use an adapted version of Barten’s CSF model from [4].</p><h4>Addressing banding artifacts</h4><p>Banding shows up as staircase-like edges in parts of the image that should look smooth. It can have a negative visual impact for viewers, but this impact is not captured well by VMAF v0. VMAF v1 integrates the Contrast Aware Multiscale Banding Index (CAMBI) as one of the elementary features. You can read more about CAMBI in our <a href="https://netflixtechblog.com/cambi-a-banding-artifact-detector-96777ae12fe2">previous tech blog</a> or the <a href="https://github.com/Netflix/vmaf/blob/master/resource/doc/papers/CAMBI_PCS2021.pdf">technical paper</a>.</p><h4>Addressing chroma artifacts</h4><p>VMAF v0 only extracts luma-based features, so it is unaware of chroma artifacts. In practice, encoding and scaling introduce chroma artifacts via quantization and subsampling. To capture such artifacts, we modified <a href="https://ieeexplore.ieee.org/document/7979533">SpEED-QA</a> and applied it to the chroma channels.</p><h4>Leveraging the no-enhancement gain (NEG) mode</h4><p>To reduce the effect of image enhancement operations, like sharpening, a standalone <a href="https://docs.google.com/document/d/1dJczEhXO0MZjBSNyKmd3ARiCTdFVMNPBykH4_HMPoyY/edit?tab=t.0#heading=h.oaikhnw46pw5">no-enhancement gain (NEG)</a> mode was made available for VMAF v0. We have found that NEG serves as a conservative quality metric and helps preserve creative intent. We already use VMAF-NEG as one of the quality metrics during codec development, such as for AV2. Therefore, NEG is enabled by default for VMAF v1 without a need for a separate model.</p><h4>Improving the motion feature</h4><p>VMAF v0’s motion feature does not have an upper bound. Further, the training data used then did not have enough coverage for high-motion sequences. Consequently, we observed that VMAF v0 could overpredict quality for very high-motion scenes. On the flip side, since motion differencing in v0 was performed between consecutive frames, v0 would underpredict quality for sequences with frame rates higher than 24 or 30 fps, like 60 fps. In v1, we apply an empirically derived hard threshold to the motion feature. Further, we add an option to measure motion differences over a larger temporal window. Expanding the temporal window alone does not fully capture the perceptual impact of 60 fps, but it does reduce the underprediction evident in v0.</p><h4>Overview of VMAF v1 models</h4><p>VMAF v1 supports the following models:</p><ul><li><strong>Standard 1080p Model:</strong> This model is calibrated for 1080p video viewed at a standard 3H distance. It uses an operating range of [0, 100].</li><li><strong>Phone Model:</strong> Derived by setting the normalized viewing distance to 5H (based on experimental data), this model adjusts the DLM, AIM, and chroma feature calculations to reflect reduced artifact visibility on smaller screens viewed from a greater relative distance. It retains the standard [0, 100] range.</li><li><strong>4K Model:</strong> We release two v1 4K models: a <strong>1.5H variant</strong> and a <strong>3H variant</strong>. The 1.5H variant is based on a discerning 4K@1.5H viewing condition. This variant is conceptually similar to its v0 4K counterpart and operates on a [0, 100] range. For most users, this variant is the default choice. The 3H variant is based on a consumer-like 4K@3H viewing condition. This variant operates on a [0, 110] range, which helps to quantify the additional perceptual benefit of 4K resolution over 1080p when both are viewed at 3H.</li></ul><h4>Interpreting the score</h4><p>VMAF v1’s score and interpretation are largely consistent with v0’s. To achieve this, we calibrated the VMAF v1 scale to align with v0 via a score transform, so that the new algorithm preserves the meaning of the numbers while keeping its accuracy benefits.</p><h4>Putting v1 to the test</h4><p>We evaluate VMAF v1 across several subjective datasets. These datasets cover a variety of codecs, content types, and use cases. For simplicity, we report the Spearman’s rank correlation coefficient (SRCC) for VMAF v0 and v1. SRCC values closer to 1 show higher agreement with subjective data. Full results will be available in a future technical paper.</p><p>In the table below, if a dataset is marked as “4K” then we measure VMAF at 4K using the 4K@1.5H model, otherwise we measure VMAF at 1080p, using the appropriate 1080p model. If a dataset is marked as “phone” the 1080p phone model is used, otherwise the standard 1080p@3H model is used.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*fyP8kVEUuoYlgoNXD3SKDQ.png"><figcaption>NFLX Banding + compression: Contains 83 AV1 videos encoded at different bitrates under banding-relevant conditions. NFLX How Low Phone: Contains 80 AVC videos encoded at 540p and below to understand viewer acceptability to low qualities. “train-only”: only the training split was publicly available.</figcaption></figure><p>As seen above, VMAF v1 matches or outperforms v0 in most datasets. There are notable improvements on large datasets like WATERLOO IVC 4K and the Netflix Screen Size Crowdsourcing, on datasets with chroma and banding artifacts, and those that involve phone viewing. On a few datasets we observe minor regressions, which are small relative to the gains elsewhere.</p><h4>Running v1</h4><p>Even with the addition of new features, we wanted VMAF v1 to have a reduced computational complexity when compared with VMAF v0. To achieve this, we have:</p><ol><li>Removed VIF (Visual Information Fidelity) as a core VMAF feature. VIF is computationally complex and did not meaningfully improve accuracy after updating the other features.</li><li>Introduced a few CAMBI-specific optimizations, both algorithmic and <a href="https://github.com/Netflix/vmaf/pull/1522">software</a>.</li><li>Measured the chroma feature at a lower scale, which does not hurt accuracy [11].</li></ol><p>The result of this work is not only a more accurate VMAF, but also a much faster VMAF. The table below shows the processing speed and threading performance for each VMAF model at 1080p and 4K. Additionally, our newest libvmaf release has a much improved threading performance, which is of benefit to both v0 and v1. Note that for content with significant banding, computing CAMBI adds some overhead, which can reduce this speedup.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*z5aAwF8Z8Bi4YPL3bFpYcA.png"><figcaption>The ducks_take_off sequence is available <a href="https://media.xiph.org/video/derf/">here</a>.</figcaption></figure><h4>Keeping some old (and good) habits</h4><p>While the core of the algorithm has changed, we still recommend:</p><ol><li>Computing VMAF at the right resolution by upsampling the distorted video to the source resolution, so that both compression and scaling artifacts are reflected. For example, bicubic upsampling can be used as a <a href="https://netflixtechblog.com/vmaf-the-journey-continues-44b51ee9ed12">general approximation</a>.</li><li>Interpreting scores in context by using the right model (e.g. 1080p vs. 4K vs. phone) for your scenario by taking into account viewing distance assumptions.</li></ol><h4>This is not the end of the road</h4><p>Just like any metric, VMAF v1 is not perfect. There is still room for improvement. Some areas that we are working on addressing in the future are: film-grain noise, improved handling for high frame rates, and perceptual codec optimizations such as adaptive quantization. We invite the community to try the latest VMAF models, report edge cases, contribute to the open-source code, and help improve VMAF. We plan to publish a detailed technical paper for VMAF v1. We also plan to release an HDR version enhanced by the v1 improvements, so stay tuned!</p><h4>Acknowledgments</h4><p>This was a collaborative effort propelled by our stunning colleagues. We want to thank the following individuals: Xiaoqing Zhu, Mariana Afonso, Anush Moorthy, Raymond Walsh, Omair Akhtar, Amelia Taylor, Ken Thomas, Zheng Lu, Chris Pham, Alex Chang, Prudhvi Chaganti, Ben Wallen, Craig Howland, Deepthi Arun, Andy Rhines and Lukáš Krasula.</p><h4>References</h4><p>[1] Z. Li et al., “<a href="https://netflixtechblog.com/toward-a-practical-perceptual-video-quality-metric-653f208b9652">Toward a practical perceptual video quality metric</a>,” Netflix Technology Blog, 2016.</p><p>[2] Z. Li et al., “<a href="https://netflixtechblog.com/vmaf-the-journey-continues-44b51ee9ed12">VMAF: The journey continues</a>,” Netflix Technology Blog, 2018.</p><p>[3] S. Li, F. Zhang, L. Ma, and K. Ngan, “Image Quality Assessment by Separately Evaluating Detail Losses and Additive Impairments,” IEEE Transactions on Multimedia, 2011.</p><p>[4] P. G. J. Barten, Contrast Sensitivity of the Human Eye and Its Effects on Image Quality. SPIE Press, 1999.</p><p>[5] Video Quality Experts Group, “Report on the validation of video quality models for high definition video content,” 2010.</p><p>[6] N. Barman, Y. Reznik, and M. G. Martini, “A Subjective Dataset for Multi-Screen Video Streaming Applications,” International Conference on Quality of Multimedia Experience (QoMEX), Ghent, Belgium, 2023, pp. 270–275.</p><p>[7] A. Katsenou, F. Zhang, M. Afonso, G. Dimitrov, and D. R. Bull, “BVI-CC: A Dataset for Research on Video Compression and Quality Assessment,” Frontiers in Signal Processing, vol. 2, 2022.</p><p>[8] C. G. Bampis, L. Krasula, Z. Li, and O. Akhtar, “Measuring and Predicting Perceptions of Video Quality Across Screen Sizes with Crowdsourcing,” International Conference on Quality of Multimedia Experience (QoMEX), Ghent, Belgium, 2023, pp. 13–18.</p><p>[9] J. Y. Lin, R. Song, C.-H. Wu, T. J. Liu, H. Wang, and C.-C. J. Kuo, “MCL-V: A streaming video quality assessment database,” Journal of Visual Communication and Image Representation, vol. 30, pp. 1–9, Jul. 2015.</p><p>[10] Z. Li, Z. Duanmu, W. Liu, and Z. Wang, “AVC, HEVC, VP9, AVS2 or AV1? — A Comparative Study of State-of-the-Art Video Encoders on 4K Videos,” Int. Conf. Image Analysis and Recognition (ICIAR), 2019.</p><p>[11] C. G. Bampis, P. Gupta, R. Soundararajan, and A. C. Bovik, “SpEED-QA: Spatial Efficient Entropic Differencing for Image and Video Quality,” IEEE Signal Process. Lett., vol. 24, no. 9, pp. 1333–1337, Sep. 2017.</p><p>[12] L.-H. Chen, C. G. Bampis, Z. Li, J. Sole, and A. C. Bovik, “Perceptual video quality prediction emphasizing chroma distortions,” IEEE Trans. Image Process., vol. 30, pp. 1941–1954, 2021.</p><p>[13] C. G. Bampis et al., “<a href="https://docs.google.com/document/d/1Ly-r196i6ekIMATuA_qKZCSNpGoKj4shqhyXU8gbSnQ/edit?tab=t.0">NFLX Screen Size Crowdsourcing dataset</a>,” 2023.</p><p>[14] H. Wei, P. Lebreton, Y. Chen, J. Zhu, and P. Le Callet, “<a href="https://sites.google.com/view/qomex26-vqm-gc/">Grand Challenge on Video Quality Assessment for Asymmetric Encoded Videos</a>,” Int. Conf. Quality of Multimedia Experience (QoMEX), Cardiff, U.K., 2026.</p><p>[15] Y. Chen, B. Chen, H. Wei, A. C. Bovik et al., “<a href="https://arxiv.org/abs/2506.22790">ICME 2025 Generalizable HDR and SDR Video Quality Measurement Grand Challenge</a>,” arXiv:2506.22790, 2025.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=60d7e4244ea8" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/vmaf-v1-good-is-not-good-enough-60d7e4244ea8">VMAF v1: Good Is Not Good Enough</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/vmaf-v1-good-is-not-good-enough-60d7e4244ea8</link>
      <guid>https://netflixtechblog.com/vmaf-v1-good-is-not-good-enough-60d7e4244ea8</guid>
      <pubDate>Sat, 20 Jun 2026 01:50:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Dynamic Repartitioning for Time Series Workloads]]></title>
      <description><![CDATA[<p>By <a href="https://www.linkedin.com/in/rajiv-shringi/">Rajiv Shringi</a>, <a href="https://www.linkedin.com/in/kaidanfullerton/">Kaidan Fullerton</a>, <a href="https://www.linkedin.com/in/oleksii-tkachuk-98b47375/">Oleksii Tkachuk</a> and <a href="https://www.linkedin.com/in/kartik894/">Kartik Sathyanarayanan</a></p><h3>Introduction</h3><p><a href="https://netflixtechblog.com/introducing-netflix-timeseries-data-abstraction-layer-31552f6326f8">Netflix’s TimeSeries Abstraction</a> is a scalable system for ingesting and querying petabytes of temporal event data with millisecond latency. We use <a href="https://cassandra.apache.org/_/index.html">Apache Cassandra</a> 4.x as the underlying storage for these main reasons:</p><ul><li><strong>Throughput, latency, and cost</strong>: Cassandra can handle millions of low‑latency reads and writes in a cost-effective manner.</li><li><strong>Operational maturity</strong>: Our data platform team has deep operational expertise running large Cassandra clusters in production.</li></ul><p>However, using Cassandra at this scale introduces trade‑offs for TimeSeries workloads. A key challenge is <a href="https://docs.datastax.com/en/cql/hcd/data-modeling/best-practices.html#bucketing">wide partitions</a>, as TimeSeries dataset partitions can grow quite large with events accumulating over time.</p><p>This problem is further compounded by the fact that TimeSeries servers routinely deal with a very high read throughput:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*0l0prhKlcOjf-d3-JVQTJg.png"><figcaption>Reads/second for different datasets</figcaption></figure><p>This post walks through our journey to reduce the impact of wide partitions in our TimeSeries datasets, the solutions we built, and the lessons we learned.</p><blockquote>Note: Although this post walks through re-partitioning in Cassandra, the same techniques can be applied more broadly to other data stores.</blockquote><h3>Impact of Wide Partitions</h3><p>For most of our datasets, we observe an average read latency in the order of single-digit milliseconds:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*sAc4HAMIvLcKTlt9NqOmjQ.png"><figcaption>Ideal Latency for Reads (ms)</figcaption></figure><p>However, in some datasets, as partitions grow too wide, we observe high read latencies in the order of seconds, especially towards the tail end:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*uvqa9oj0P6uT-G27WPRJoA.png"><figcaption>High Tail Latency for Reads (seconds)</figcaption></figure><p>This can result in timeouts:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*gcmGxn6NdjRuy0HYQIDv9g.png"><figcaption>Read timeouts / second</figcaption></figure><p>In extreme cases, if most of the reads target wide partitions, we can see Garbage Collection pauses, high CPU utilization and thread queueing.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*lj0yVnacUHcE1RQ7iHRKVA.png"><figcaption>High CPU utilization and thread-queueing in Cassandra clusters</figcaption></figure><p>Scaling up the underlying Cassandra cluster is always an option, but we need smarter alternatives than just throwing more money at the problem.</p><h3>TimeSeries Partitioning Strategy</h3><p>The TimeSeries Abstraction was designed to solve the problem of wide partitions by dividing the data into discrete time chunks. For more in-depth information, refer to our previous <a href="https://netflixtechblog.com/introducing-netflix-timeseries-data-abstraction-layer-31552f6326f8">blog</a>.</p><p>To summarize, here is an illustration of how TimeSeries partitioning strategy helps us break up wide partitions into manageable chunks.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*ySI9hOE4_lSScqQ16KQAuw.png"><figcaption>Time Series partitioning breaking up a dataset into Time slices, time buckets and event buckets</figcaption></figure><p>This strategy further allows us to efficiently query and drop data based on time, without having to deal with <a href="https://opencredo.com/blogs/cassandra-tombstones-common-issues">tombstones</a>.</p><h3>Picking the Partitioning Strategy</h3><p>When a namespace (a.k.a. dataset) is created, users must specify their anticipated workload characteristics. This specification is then fed into our <a href="https://github.com/Netflix-Skunkworks/service-capacity-modeling/blob/main/service_capacity_modeling/models/org/netflix/time_series.py">provisioning</a> pipeline. The pipeline processes these inputs, runs <a href="https://en.wikipedia.org/wiki/Monte_Carlo_method">Monte Carlo</a> simulations, and produces an optimal infrastructure and partition configuration.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*016OPfeTjQnGXnUSGdTIKA.png"><figcaption>Provisioning picks optimal infra and configuration based on user inputs</figcaption></figure><p>You can learn more about our methodology of capacity planning in this insightful <a href="https://www.youtube.com/watch?v=Lf6B1PxIvAs">AWS re:Invent</a> talk given by one of our stunning colleagues.</p><h3>The Problem with the Current Approach</h3><p>Although this method of provisioning is effective in many situations, it proves insufficient for TimeSeries workloads under these conditions:</p><ul><li><strong>Workload is unknown or inaccurately estimated:</strong> Early on in a project, users can lack a reliable picture of production traffic or simply misestimate key parameters.</li><li><strong>Workload evolves over time:</strong> Traffic patterns, client behavior, and product requirements change. A “good” partitioning strategy on day one can become inefficient months later.</li><li><strong>Data outliers exist:</strong> Not all TimeSeries IDs behave the same. A small percentage of IDs can receive a vastly higher volume of events than the rest.</li></ul><p>Fortunately, our design with discrete Time Slices gives us a natural escape hatch for the first two scenarios; each new Time Slice can use a different partitioning strategy.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*F2gUdVv8NGAqoYbmFVMbnQ.png"><figcaption>Each Time Slice can have a unique partition strategy</figcaption></figure><p>However, manually adjusting these configurations in a fleet that has thousands of TimeSeries datasets is not sustainable. We need automation.</p><h3>Solution 1: Time Slice Re-Partitioning</h3><p>Cassandra exposes useful introspection APIs for understanding data usage and access patterns. For example, <a href="https://docs.datastax.com/en/dse/6.9/managing/tools/nodetool/table-histograms.html">nodetool tablehistograms</a> provide percentile distributions for partition sizes in a table. Using these tools, we can detect cases of both over and under partitioning.</p><p>Below is an example of over‑partitioning, where the TimeSeries provisioning pipeline selected very small <em>time_bucket</em> intervals based on user provided inputs:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/811/1*yCT6SVX5682twAp2VhLPJg.png"><figcaption>Provisioning selected 60s time buckets based on user inputs</figcaption></figure><p>causing partitions to have less than 10 KB of data, leading to high read amplification and thread queueing:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/920/1*XM-yySVKVu1YTyt9uPOk-w.png"><figcaption>Histogram of the given Cassandra table showing partition size percentiles</figcaption></figure><p>In order to tune partition strategies efficiently, we added a background worker, which monitors partition histograms of Time Slices attached to a given application, and exposes it via a Cassandra virtual table:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*ZMZ1RYb52fha32uStM7w9A.png"><figcaption>Histograms exposed through a Cassandra Virtual table</figcaption></figure><p>It then computes an adjustment factor when it detects partition sizes not meeting a configured density. This configured density is often set between 2 MiB to 10 MiB depending on the workload.</p><pre>DynamicTimeSliceConfigWorker: <br>namespace: my_dataset_1<br>Observed: TimeSlices have p99 partitions below configured target of 10MB. <br>Proposed: time_bucket interval: 60s -&gt; 604800s</pre><p>The worker can then update <em>future</em> Time Slices with the new partition strategy:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/710/1*HUAyyJeONB-5HVnMihAx8g.png"><figcaption>Partitioning adjusted for future Time Slice(s)</figcaption></figure><p>This strategy has yielded real results in reducing our read latencies, as well as reducing the number of timeouts caused by thread queueing.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*tE44f8HSDXCwvlffOq0w4A.png"><figcaption>Reduction in tail latency and thread queueing for</figcaption></figure><p>However, this strategy only works if most of the data exhibits such behavior that warrants re-partitioning of the entire table. It does not work in cases where only a percentage of IDs within the table are wide.</p><p>We have a couple of options here:</p><ul><li><strong>Do Nothing</strong>: This is sometimes the right approach if there is no observed impact to the application’s top-level metrics.</li><li><strong>Partial Returns</strong>: We implemented a ‘Partial Return’ feature, which aborts an inflight request if it has breached a configured latency SLO, while returning whatever data it has collected up until that point. This is a great option for clients who care more about latency than fetching <em>all</em> the data.</li></ul><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*t7cz4h3Ebxg-HDyaMVIp_g.png"><figcaption>Tail latency drops around the SLO cutoff as Partial Returns are enabled</figcaption></figure><ul><li><strong>Block IDs:</strong> This is an extreme step but worth mentioning, because we do deal with bad data that occasionally seeps into the system e.g. test or spam IDs that can make the system unstable.</li></ul><pre>dgwts.config.&lt;dataset&gt;.block.Ids: "&lt;tsid-1&gt;, &lt;tsid-2&gt;, &lt;tsid-3&gt;"</pre><p>Ultimately, we encounter scenarios where valid and important TimeSeries IDs accumulate a high enough volume of events, with callers needing to process all the related data. Simply tolerating elevated latencies or timeouts when querying these IDs is not a desirable outcome.</p><p>This is where dynamic partitioning comes into play.</p><h3>Solution 2: Dynamic Partitioning per ID</h3><p>Dynamic partitioning is an asynchronous pipeline that auto-detects and splits wide partitions on a TimeSeries ID level rather than at the table level.</p><p>It has three main stages:</p><ul><li><strong>Detection</strong>: Detects wide partitions for a given TimeSeries ID during the read path.</li><li><strong>Planning &amp; Splitting</strong>: Plans and executes splits of those partitions into optimal sizes asynchronously.</li><li><strong>Serving Reads</strong>: Re-routes the read queries transparently to read data from the split partitions when ready.</li></ul><p>This is how it works at a high level; we will dive into details after:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*tucaTmgbj6BqhA8YKmUn9A.png"><figcaption>Dynamic Wide Partition Split Async Pipeline</figcaption></figure><p>Here are the different stages of the pipeline:</p><h4>Detection</h4><p>Every TimeSeries read operation tracks how many bytes are read for a given partition. If the bytes read exceed a configured threshold, the server emits a detection event to Kafka:</p><pre>{<br>  "time_slice": "data_20260328", // the Cassandra table this event was detected in<br>  "time_series_id": "profileId:123", // the ID detected as wide<br>  "time_bucket": 7, // the existing time_bucket partition<br>  "event_bucket": 2, // the existing event_bucket partition<br>  "immutable": true, // TimeSeries servers can compute if this partition is no longer receiving writes<br>  "version": "0" // reserved for future use e.g. invalidate if partition is no longer immutable<br>}</pre><p>Our decision to detect wide partitions on reads, as opposed to writes, is based on our observation that the majority of the data in the wild doesn’t need this treatment. The slight downside is that some reads on these large partitions may suffer sub-optimal performance for a very short duration (typically seconds) until this process catches up.</p><h4>Immutability</h4><p>Although splitting mutable partitions is possible, it is inherently more complex. As a first step towards solving this problem, we chose to reduce the surface area of this change by focusing on immutable partitions, while still meaningfully reducing caller timeouts.</p><h4>Planning</h4><p>Detection may occur based on a partial read, so the planner must still read the entire partition <em>once</em> to compute an accurate split plan. The checkpointing becomes crucial here. For planning reads that fail to process the entire partition, the process can always continue from the last saved checkpoint.</p><h4>Checkpointing</h4><p>The <em>wide_row</em> metadata table serves as the backbone for state transitions and checkpointing of partition splits. It also stores information that is used later by TimeSeries servers to properly route Read queries.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*MNIBbz9hTPz-7ah8FRqTZQ.png"><figcaption>wide_row metadata for storing split states and checkpoints</figcaption></figure><h4>Splitting</h4><p>The Planner delegates the splitting of data to an appropriate split-strategy. For example, if <em>EventBucketPartitionSplitStrategy</em> is selected, we split the partition by assigning more event buckets to the same time bucket. If the partition is <em>ultra-wide</em>, we cap the number of event buckets we split into, in order to control the resultant read amplification. Spreading into multiple partitions in such cases is still beneficial in order to spread the read workload to multiple Cassandra replicas.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*Acbtr-hKcuopZ_7tXpWoxQ.png"><figcaption>Split by assigning more event buckets for a given time bucket</figcaption></figure><p>Further, since the Splitter has the full view of the partition, it can ensure total sort order across all the split buckets.</p><h4>Validating Splits</h4><p>The Planner stores a pre-split checksum of a given partition during the planning phase, while the Splitter computes and stores the post-split checksum. The split status is marked as completed only if the two checksums match.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/997/1*BbADzQfFTZPWUbpORlXfGQ.png"><figcaption>Ensure checksums match pre- and post-split before marking a split as COMPLETED</figcaption></figure><h4>Tracking Splits</h4><p>The pre- and post-split partition sizes across different datasets are tracked to see how effectively the partition splits are being planned and executed:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*tif_TQsk2OkBV9CCWrfDJQ.png"><figcaption>Track pre- and post-split partition sizes to ensure we are splitting optimally</figcaption></figure><h4>Serving Reads</h4><p>The TimeSeries servers load the partition-keys of completed splits periodically into in-memory Bloom filters. Every read operation checks the Bloom filter to see whether a query can be diverted to the split partitions.</p><p>Here is what the Read path looks like:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*1WeYb1fTJq8-nkkMr-RjzQ.png"><figcaption>Read path for diverting reads to existing or split partitions</figcaption></figure><p>The size of the Bloom filters is monitored to ensure we have enough memory per server. Due to the compactness of partition keys, and ratio of wide partitions in a given dataset, the filters fit comfortably in each server instance.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*73ahHVoNRUYaczuEn9KTBA.png"><figcaption>Bloom filter approximate element count per namespace and time slice</figcaption></figure><p>The Bloom filter latency to check whether a given partition key is wide for every read request is typically in single-digit microseconds or better, making this diversion practically invisible to the callers.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*t1ZASboghIXdGDyHrz0Z0Q.png"><figcaption>Latency for checking Bloom filters is extremely small for callers to notice the diversion</figcaption></figure><p>For the cases that do end up with a Bloom filter hit, the TimeSeries servers lookup the <em>wide_row</em> metadata to see how to read a specific wide partition:</p><pre>{<br>  "pre_split_data": {<br>    "time_slice": "data_20260328",<br>    "time_series_id": "6313825", → What to read<br>    "time_bucket": 0,<br>    "event_bucket": 2<br>    …<br>  },<br>  "post_split_data": {<br>    "time_slice": "wide_data_20260328_0", → Where to read it from<br>    "event_bucket_partition_strategy": { → Strategy to delegate to for reading<br>    "target_event_buckets": 2,<br>    "start_event_bucket": 32 → How should the strategy read it<br>  }<br>  …<br>}</pre><p>This metadata read is backed by a read-through cache, making it quite performant:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*CprO4zk8UQFYulId8jHdUg.png"><figcaption>Metadata fetch latency is quite low to affect read operations</figcaption></figure><p>Finally, the reads for the split partitions are delegated to our existing <em>PartitionReader</em>, which reads <em>N smaller partitions in parallel</em>, rather than 1 large partition, improving overall performance and stability!</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*CHPEUf2WlJzHyUGdNiEA3A.png"><figcaption>Read much smaller partitions in parallel and merge results</figcaption></figure><h4>Fallbacks</h4><p>The existing wide partition from the original time slice is never deleted. This helps us in creating safe fallbacks in many different scenarios of partial failures and eventual consistency. The slightly larger storage space we use as a result is worth the operational safety we gain.</p><h4>Building Additional Confidence</h4><p>Serving incorrect reads would be disastrous. To establish trust beyond checksums, we leveraged additional mechanisms such as:</p><ul><li>Using our existing <a href="https://netflixtechblog.medium.com/data-bridge-how-netflix-simplifies-data-movement-36d10d91c313">Data Bridge</a> pipelines to verify splits offline:</li></ul><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*AspDR7sF38JxFXTkaK1wWQ.png"><figcaption>Spark job to ensure that the split data is an exact match to the original data</figcaption></figure><ul><li>Implementing a phased rollout strategy to safely advance through stages as our confidence in the system grew:</li></ul><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*la6WvWA4KWUglIPWpG27qw.png"><figcaption>Advance through Read modes once previous mode passes checks</figcaption></figure><p>A critical part of this phased rollout was the <strong>Comparison</strong> phase, which compared bytes served by old read path and the new read path while in <em>shadow</em> mode:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*sELeFr6fngTGib_gcD2ejA.png"><figcaption>A chart of bytes match vs bytes differ in a given shadow period</figcaption></figure><h4>Results</h4><p>As a result of these dynamic splits, we see a huge improvement in the average read latency of most wide partitions, bringing it down from seconds:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*AZuNdkiqBRtlcLhJBC-MzA.png"><figcaption>Existing average latency for reading wide partitions</figcaption></figure><p>to <em>low double-digit milliseconds!</em></p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*y8H7tjqj-g2-ZRGPBkUl_A.png"><figcaption>Average latency for reading dynamically split partitions</figcaption></figure><p>Tail latencies of reading wide partitions dropped from several seconds:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*JbNIJYjARHjGJH77-FiV7w.png"><figcaption>Existing tail latency for reading wide partitions</figcaption></figure><p>to around 200 ms or better:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*qKneKuGrpiU0VoMRw7OOZQ.png"><figcaption>Tail latency for reading dynamically split partitions</figcaption></figure><p>resulting in a drop in read timeouts:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*UjFSHe5HeZeDBKbmayLdnw.png"></figure><p>Overall, this has resulted in a more stable Cassandra cluster with lower CPU utilization and little to no thread queuing:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*Wt_oTudRIz8XKmJ5N3FcDA.png"><figcaption>Low CPU utilization and no thread-queueing</figcaption></figure><p>Further, for extreme wide rows, where a dataset would face constant timeouts and unavailability blips, the service was able to paginate and query 500MB+ partitions while remaining available:</p><pre>grpc … com.netflix.dgw.ts.TimeSeriesService/SearchEventRecords -d<br>'{"namespace": "...",<br>    "search_query": {...},<br>    "time_interval": {<br>      "start": "2026–05–11T23:42:51.484398Z",<br>      "end": "2026–05–12T00:13:50.694205Z"<br>    },<br>    "pageSize" : 1000,<br>  }'<br># Response:<br>{<br>  "next_page_token" : ….,<br>  "records": [<br>    {<br>      …<br>    }<br>  ],<br>  "response_context": [{<br>    "namespace": "...",<br>    …<br>    # Trades elevated latency for being available<br>    "time_taken": "41.072410142s"<br>    }<br>  ]<br>}</pre><h3>Conclusion</h3><p>There is more work planned around this feature, like splitting <em>mutable</em> wide partitions, or re-processing previously failed splits, but this has been a successful start in improving service performance and reducing our support burden.</p><p>Further, we would like to highlight some key lessons that we learned at different points in this journey.</p><ul><li><strong>Reducing Surface Area: </strong>As a first step, explore simpler solutions that can still deliver meaningful impact. Also, reducing the surface area of a complex change and deploying incrementally pays off operationally.</li><li><strong>Building Confidence</strong>: Invest time and resources to build confidence in new features, especially when justified by the feature complexity, deployment blast radius, and/or potential impact.</li></ul><p><strong>Acknowledgements</strong>: Special thanks to our stunning colleagues who further contributed to this feature’s success: <a href="https://www.linkedin.com/in/tomdevoe/">Tom DeVoe</a>, <a href="https://www.linkedin.com/in/clohfink/">Chris Lohfink</a>, <a href="https://www.linkedin.com/in/sumanth-pasupuleti/">Sumanth Pasupuleti</a> and <a href="https://www.linkedin.com/in/joseph-lynch-9976a431/">Joey Lynch</a>.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=0eded064f456" width="1" height="1" alt=""><hr><p><a href="https://medium.com/netflix-techblog/dynamically-splitting-wide-partitions-in-cassandra-for-time-series-workloads-0eded064f456">Dynamic Repartitioning for Time Series Workloads</a> was originally published in <a href="https://medium.com/netflix-techblog">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://medium.com/netflix-techblog/dynamically-splitting-wide-partitions-in-cassandra-for-time-series-workloads-0eded064f456</link>
      <guid>https://medium.com/netflix-techblog/dynamically-splitting-wide-partitions-in-cassandra-for-time-series-workloads-0eded064f456</guid>
      <pubDate>Wed, 03 Jun 2026 04:05:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Dynamically Splitting Wide Partitions in Cassandra for Time Series Workloads]]></title>
      <description><![CDATA[<p>By <a href="https://www.linkedin.com/in/rajiv-shringi/">Rajiv Shringi</a>, <a href="https://www.linkedin.com/in/kaidanfullerton/">Kaidan Fullerton</a>, <a href="https://www.linkedin.com/in/oleksii-tkachuk-98b47375/">Oleksii Tkachuk</a> and <a href="https://www.linkedin.com/in/kartik894/">Kartik Sathyanarayanan</a></p><h3>Introduction</h3><p><a href="https://netflixtechblog.com/introducing-netflix-timeseries-data-abstraction-layer-31552f6326f8">Netflix’s TimeSeries Abstraction</a> is a scalable system for ingesting and querying petabytes of temporal event data with millisecond latency. We use <a href="https://cassandra.apache.org/_/index.html">Apache Cassandra</a> 4.x as the underlying storage for these main reasons:</p><ul><li><strong>Throughput, latency, and cost</strong>: Cassandra can handle millions of low‑latency reads and writes in a cost-effective manner.</li><li><strong>Operational maturity</strong>: Our data platform team has deep operational expertise running large Cassandra clusters in production.</li></ul><p>However, using Cassandra at this scale introduces trade‑offs for TimeSeries workloads. A key challenge is <a href="https://docs.datastax.com/en/cql/hcd/data-modeling/best-practices.html#bucketing">wide partitions</a>, as TimeSeries dataset partitions can grow quite large with events accumulating over time.</p><p>This problem is further compounded by the fact that TimeSeries servers routinely deal with a very high read throughput:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*0l0prhKlcOjf-d3-JVQTJg.png"><figcaption>Reads/second for different datasets</figcaption></figure><p>This post walks through our journey to reduce the impact of wide partitions in our TimeSeries datasets, the solutions we built, and the lessons we learned.</p><h3>Impact of Wide Partitions</h3><p>For most of our datasets, we observe an average read latency in the order of single-digit milliseconds:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*sAc4HAMIvLcKTlt9NqOmjQ.png"><figcaption>Ideal Latency for Reads (ms)</figcaption></figure><p>However, in some datasets, as partitions grow too wide, we observe high read latencies in the order of seconds, especially towards the tail end:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*uvqa9oj0P6uT-G27WPRJoA.png"><figcaption>High Tail Latency for Reads (seconds)</figcaption></figure><p>This can result in timeouts:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*gcmGxn6NdjRuy0HYQIDv9g.png"><figcaption>Read timeouts / second</figcaption></figure><p>In extreme cases, if most of the reads target wide partitions, we can see Garbage Collection pauses, high CPU utilization and thread queueing.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*lj0yVnacUHcE1RQ7iHRKVA.png"><figcaption>High CPU utilization and thread-queueing in Cassandra clusters</figcaption></figure><p>Scaling up the underlying Cassandra cluster is always an option, but we need smarter alternatives than just throwing more money at the problem.</p><h3>TimeSeries Partitioning Strategy</h3><p>The TimeSeries Abstraction was designed to solve the problem of wide partitions by dividing the data into discrete time chunks. For more in-depth information, refer to our previous <a href="https://netflixtechblog.com/introducing-netflix-timeseries-data-abstraction-layer-31552f6326f8">blog</a>.</p><p>To summarize, here is an illustration of how TimeSeries partitioning strategy helps us break up wide partitions into manageable chunks.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*ySI9hOE4_lSScqQ16KQAuw.png"><figcaption>Time Series partitioning breaking up a dataset into Time slices, time buckets and event buckets</figcaption></figure><p>This strategy further allows us to efficiently query and drop data based on time, without having to deal with <a href="https://opencredo.com/blogs/cassandra-tombstones-common-issues">tombstones</a>.</p><h3>Picking the Partitioning Strategy</h3><p>When a namespace (a.k.a. dataset) is created, users must specify their anticipated workload characteristics. This specification is then fed into our <a href="https://github.com/Netflix-Skunkworks/service-capacity-modeling/blob/main/service_capacity_modeling/models/org/netflix/time_series.py">provisioning</a> pipeline. The pipeline processes these inputs, runs <a href="https://en.wikipedia.org/wiki/Monte_Carlo_method">Monte Carlo</a> simulations, and produces an optimal infrastructure and partition configuration.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*016OPfeTjQnGXnUSGdTIKA.png"><figcaption>Provisioning picks optimal infra and configuration based on user inputs</figcaption></figure><p>You can learn more about our methodology of capacity planning in this insightful <a href="https://www.youtube.com/watch?v=Lf6B1PxIvAs">AWS re:Invent</a> talk given by one of our stunning colleagues.</p><h3>The Problem with the Current Approach</h3><p>Although this method of provisioning is effective in many situations, it proves insufficient for TimeSeries workloads under these conditions:</p><ul><li><strong>Workload is unknown or inaccurately estimated:</strong> Early on in a project, users can lack a reliable picture of production traffic or simply misestimate key parameters.</li><li><strong>Workload evolves over time:</strong> Traffic patterns, client behavior, and product requirements change. A “good” partitioning strategy on day one can become inefficient months later.</li><li><strong>Data outliers exist:</strong> Not all TimeSeries IDs behave the same. A small percentage of IDs can receive a vastly higher volume of events than the rest.</li></ul><p>Fortunately, our design with discrete Time Slices gives us a natural escape hatch for the first two scenarios; each new Time Slice can use a different partitioning strategy.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*F2gUdVv8NGAqoYbmFVMbnQ.png"><figcaption>Each Time Slice can have a unique partition strategy</figcaption></figure><p>However, manually adjusting these configurations in a fleet that has thousands of TimeSeries datasets is not sustainable. We need automation.</p><h3>Solution 1: Time Slice Re-Partitioning</h3><p>Cassandra exposes useful introspection APIs for understanding data usage and access patterns. For example, <a href="https://docs.datastax.com/en/dse/6.9/managing/tools/nodetool/table-histograms.html">nodetool tablehistograms</a> provide percentile distributions for partition sizes in a table. Using these tools, we can detect cases of both over and under partitioning.</p><p>Below is an example of over‑partitioning, where the TimeSeries provisioning pipeline selected very small <em>time_bucket</em> intervals based on user provided inputs:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/811/1*yCT6SVX5682twAp2VhLPJg.png"><figcaption>Provisioning selected 60s time buckets based on user inputs</figcaption></figure><p>causing partitions to have less than 10 KB of data, leading to high read amplification and thread queueing:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/920/1*XM-yySVKVu1YTyt9uPOk-w.png"><figcaption>Histogram of the given Cassandra table showing partition size percentiles</figcaption></figure><p>In order to tune partition strategies efficiently, we added a background worker, which monitors partition histograms of Time Slices attached to a given application, and exposes it via a Cassandra virtual table:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*ZMZ1RYb52fha32uStM7w9A.png"><figcaption>Histograms exposed through a Cassandra Virtual table</figcaption></figure><p>It then computes an adjustment factor when it detects partition sizes not meeting a configured density. This configured density is often set between 2 MiB to 10 MiB depending on the workload.</p><pre>DynamicTimeSliceConfigWorker: <br>namespace: my_dataset_1<br>Observed: TimeSlices have p99 partitions below configured target of 10MB. <br>Proposed: time_bucket interval: 60s -&gt; 604800s</pre><p>The worker can then update <em>future</em> Time Slices with the new partition strategy:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/710/1*HUAyyJeONB-5HVnMihAx8g.png"><figcaption>Partitioning adjusted for future Time Slice(s)</figcaption></figure><p>This strategy has yielded real results in reducing our read latencies, as well as reducing the number of timeouts caused by thread queueing.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*tE44f8HSDXCwvlffOq0w4A.png"><figcaption>Reduction in tail latency and thread queueing for</figcaption></figure><p>However, this strategy only works if most of the data exhibits such behavior that warrants re-partitioning of the entire table. It does not work in cases where only a percentage of IDs within the table are wide.</p><p>We have a couple of options here:</p><ul><li><strong>Do Nothing</strong>: This is sometimes the right approach if there is no observed impact to the application’s top-level metrics.</li><li><strong>Partial Returns</strong>: We implemented a ‘Partial Return’ feature, which aborts an inflight request if it has breached a configured latency SLO, while returning whatever data it has collected up until that point. This is a great option for clients who care more about latency than fetching <em>all</em> the data.</li></ul><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*t7cz4h3Ebxg-HDyaMVIp_g.png"><figcaption>Tail latency drops around the SLO cutoff as Partial Returns are enabled</figcaption></figure><ul><li><strong>Block IDs:</strong> This is an extreme step but worth mentioning, because we do deal with bad data that occasionally seeps into the system e.g. test or spam IDs that can make the system unstable.</li></ul><pre>dgwts.config.&lt;dataset&gt;.block.Ids: "&lt;tsid-1&gt;, &lt;tsid-2&gt;, &lt;tsid-3&gt;"</pre><p>Ultimately, we encounter scenarios where valid and important TimeSeries IDs accumulate a high enough volume of events, with callers needing to process all the related data. Simply tolerating elevated latencies or timeouts when querying these IDs is not a desirable outcome.</p><p>This is where dynamic partitioning comes into play.</p><h3>Solution 2: Dynamic Partitioning per ID</h3><p>Dynamic partitioning is an asynchronous pipeline that auto-detects and splits wide partitions on a TimeSeries ID level rather than at the table level.</p><p>It has three main stages:</p><ul><li><strong>Detection</strong>: Detects wide partitions for a given TimeSeries ID during the read path.</li><li><strong>Planning &amp; Splitting</strong>: Plans and executes splits of those partitions into optimal sizes asynchronously.</li><li><strong>Serving Reads</strong>: Re-routes the read queries transparently to read data from the split partitions when ready.</li></ul><p>This is how it works at a high level; we will dive into details after:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*tucaTmgbj6BqhA8YKmUn9A.png"><figcaption>Dynamic Wide Partition Split Async Pipeline</figcaption></figure><p>Here are the different stages of the pipeline:</p><h4>Detection</h4><p>Every TimeSeries read operation tracks how many bytes are read for a given partition. If the bytes read exceed a configured threshold, the server emits a detection event to Kafka:</p><pre>{<br>  "time_slice": "data_20260328", // the Cassandra table this event was detected in<br>  "time_series_id": "profileId:123", // the ID detected as wide<br>  "time_bucket": 7, // the existing time_bucket partition<br>  "event_bucket": 2, // the existing event_bucket partition<br>  "immutable": true, // TimeSeries servers can compute if this partition is no longer receiving writes<br>  "version": "0" // reserved for future use e.g. invalidate if partition is no longer immutable<br>}</pre><p>Our decision to detect wide partitions on reads, as opposed to writes, is based on our observation that the majority of the data in the wild doesn’t need this treatment. The slight downside is that some reads on these large partitions may suffer sub-optimal performance for a very short duration (typically seconds) until this process catches up.</p><h4>Immutability</h4><p>Although splitting mutable partitions is possible, it is inherently more complex. As a first step towards solving this problem, we chose to reduce the surface area of this change by focusing on immutable partitions, while still meaningfully reducing caller timeouts.</p><h4>Planning</h4><p>Detection may occur based on a partial read, so the planner must still read the entire partition <em>once</em> to compute an accurate split plan. The checkpointing becomes crucial here. For planning reads that fail to process the entire partition, the process can always continue from the last saved checkpoint.</p><h4>Checkpointing</h4><p>The <em>wide_row</em> metadata table serves as the backbone for state transitions and checkpointing of partition splits. It also stores information that is used later by TimeSeries servers to properly route Read queries.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*MNIBbz9hTPz-7ah8FRqTZQ.png"><figcaption>wide_row metadata for storing split states and checkpoints</figcaption></figure><h4>Splitting</h4><p>The Planner delegates the splitting of data to an appropriate split-strategy. For example, if <em>EventBucketPartitionSplitStrategy</em> is selected, we split the partition by assigning more event buckets to the same time bucket. If the partition is <em>ultra-wide</em>, we cap the number of event buckets we split into, in order to control the resultant read amplification. Spreading into multiple partitions in such cases is still beneficial in order to spread the read workload to multiple Cassandra replicas.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*Acbtr-hKcuopZ_7tXpWoxQ.png"><figcaption>Split by assigning more event buckets for a given time bucket</figcaption></figure><h4>Validating Splits</h4><p>The Planner stores a pre-split checksum of a given partition during the planning phase, while the Splitter computes and stores the post-split checksum. The split status is marked as completed only if the two checksums match.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/997/1*BbADzQfFTZPWUbpORlXfGQ.png"><figcaption>Ensure checksums match pre- and post-split before marking a split as COMPLETED</figcaption></figure><h4>Tracking Splits</h4><p>The pre- and post-split partition sizes across different datasets are tracked to see how effectively the partition splits are being planned and executed:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*tif_TQsk2OkBV9CCWrfDJQ.png"><figcaption>Track pre- and post-split partition sizes to ensure we are splitting optimally</figcaption></figure><h4>Serving Reads</h4><p>The TimeSeries servers load the partition-keys of completed splits periodically into in-memory Bloom filters. Every read operation checks the Bloom filter to see whether a query can be diverted to the split partitions.</p><p>Here is what the Read path looks like:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*aW9eB9OTQx10HqTASEmWjQ.png"><figcaption>Read path for diverting reads to existing or split partitions</figcaption></figure><p>The size of the Bloom filters is monitored to ensure we have enough memory per server. Due to the compactness of partition keys, and ratio of wide partitions in a given dataset, the filters fit comfortably in each server instance.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*73ahHVoNRUYaczuEn9KTBA.png"><figcaption>Bloom filter approximate element count per namespace and time slice</figcaption></figure><p>The Bloom filter latency to check whether a given partition key is wide for every read request is typically in single-digit microseconds or better, making this diversion practically invisible to the callers.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*t1ZASboghIXdGDyHrz0Z0Q.png"><figcaption>Latency for checking Bloom filters is extremely small for callers to notice the diversion</figcaption></figure><p>For the cases that do end up with a Bloom filter hit, the TimeSeries servers lookup the <em>wide_row</em> metadata to see how to read a specific wide partition:</p><pre>{<br>  "pre_split_data": {<br>    "time_slice": "data_20260328",<br>    "time_series_id": "6313825", → What to read<br>    "time_bucket": 0,<br>    "event_bucket": 2<br>    …<br>  },<br>  "post_split_data": {<br>    "time_slice": "wide_data_20260328_0", → Where to read it from<br>    "event_bucket_partition_strategy": { → Strategy to delegate to for reading<br>    "target_event_buckets": 2,<br>    "start_event_bucket": 32 → How should the strategy read it<br>  }<br>  …<br>}</pre><p>This metadata read is backed by a read-through cache, making it quite performant:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*CprO4zk8UQFYulId8jHdUg.png"><figcaption>Metadata fetch latency is quite low to affect read operations</figcaption></figure><p>Finally, the reads for the split partitions are delegated to our existing <em>PartitionReader</em>. Having the <em>same schema</em> for the split table allows us to reuse code and minimize changes.</p><h4>Fallbacks</h4><p>The existing wide partition from the original time slice is never deleted. This helps us in creating safe fallbacks in many different scenarios of partial failures and eventual consistency. The slightly larger storage space we use as a result is worth the operational safety we gain.</p><h4>Building Additional Confidence</h4><p>Serving incorrect reads would be disastrous. To establish trust beyond checksums, we leveraged additional mechanisms such as:</p><ul><li>Using our existing <a href="https://netflixtechblog.medium.com/data-bridge-how-netflix-simplifies-data-movement-36d10d91c313">Data Bridge</a> pipelines to verify splits offline:</li></ul><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*AspDR7sF38JxFXTkaK1wWQ.png"><figcaption>Spark job to ensure that the split data is an exact match to the original data</figcaption></figure><ul><li>Implementing a phased rollout strategy to safely advance through stages as our confidence in the system grew:</li></ul><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*la6WvWA4KWUglIPWpG27qw.png"><figcaption>Advance through Read modes once previous mode passes checks</figcaption></figure><p>A critical part of this phased rollout was the <strong>Comparison</strong> phase, which compared bytes served by old read path and the new read path while in <em>shadow</em> mode:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*sELeFr6fngTGib_gcD2ejA.png"><figcaption>A chart of bytes match vs bytes differ in a given shadow period</figcaption></figure><h4>Results</h4><p>As a result of these dynamic splits, we see a huge improvement in the average read latency of most wide partitions, bringing it down from seconds:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*AZuNdkiqBRtlcLhJBC-MzA.png"><figcaption>Existing average latency for reading wide partitions</figcaption></figure><p>to <em>low double-digit milliseconds!</em></p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*y8H7tjqj-g2-ZRGPBkUl_A.png"><figcaption>Average latency for reading dynamically split partitions</figcaption></figure><p>Tail latencies of reading wide partitions dropped from several seconds:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*JbNIJYjARHjGJH77-FiV7w.png"><figcaption>Existing tail latency for reading wide partitions</figcaption></figure><p>to around 200 ms or better:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*qKneKuGrpiU0VoMRw7OOZQ.png"><figcaption>Tail latency for reading dynamically split partitions</figcaption></figure><p>resulting in a drop in read timeouts:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*UjFSHe5HeZeDBKbmayLdnw.png"></figure><p>Further, for extreme wide rows, where a dataset would face constant timeouts and unavailability blips, the service was able to paginate and query 500MB+ partitions while remaining available:</p><pre>grpc … com.netflix.dgw.ts.TimeSeriesService/SearchEventRecords -d<br>'{"namespace": "...",<br>    "search_query": {...},<br>    "time_interval": {<br>      "start": "2026–05–11T23:42:51.484398Z",<br>      "end": "2026–05–12T00:13:50.694205Z"<br>    },<br>    "pageSize" : 1000,<br>  }'<br># Response:<br>{<br>  "next_page_token" : ….,<br>  "records": [<br>    {<br>      …<br>    }<br>  ],<br>  "response_context": [{<br>    "namespace": "...",<br>    …<br>    # Trades elevated latency for being available<br>    "time_taken": "41.072410142s"<br>    }<br>  ]<br>}</pre><h3>Conclusion</h3><p>There is more work planned around this feature, like splitting <em>mutable</em> wide partitions, or re-processing previously failed splits, but this has been a successful start in improving service performance and reducing our support burden.</p><p>Further, we would like to highlight some key lessons that we learned at different points in this journey.</p><ul><li><strong>Reducing Surface Area: </strong>As a first step, explore simpler solutions that can still deliver meaningful impact. Also, reducing the surface area of a complex change and deploying incrementally pays off operationally.</li><li><strong>Building Confidence</strong>: Invest time and resources to build confidence in new features, especially when justified by the feature complexity, deployment blast radius, and/or potential impact.</li></ul><p><strong>Acknowledgements</strong>: Special thanks to our stunning colleagues who further contributed to this feature’s success: <a href="https://www.linkedin.com/in/tomdevoe/">Tom DeVoe</a>, <a href="https://www.linkedin.com/in/clohfink/">Chris Lohfink</a>, <a href="https://www.linkedin.com/in/sumanth-pasupuleti/">Sumanth Pasupuleti</a> and <a href="https://www.linkedin.com/in/joseph-lynch-9976a431/">Joey Lynch</a>.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=0eded064f456" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/dynamically-splitting-wide-partitions-in-cassandra-for-time-series-workloads-0eded064f456">Dynamically Splitting Wide Partitions in Cassandra for Time Series Workloads</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/dynamically-splitting-wide-partitions-in-cassandra-for-time-series-workloads-0eded064f456</link>
      <guid>https://netflixtechblog.com/dynamically-splitting-wide-partitions-in-cassandra-for-time-series-workloads-0eded064f456</guid>
      <pubDate>Wed, 03 Jun 2026 04:05:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[High-Throughput Graph Abstraction at Netflix: Part I]]></title>
      <description><![CDATA[<p>By <a href="https://www.linkedin.com/in/oleksii-tkachuk-98b47375/">Oleksii Tkachuk</a>, <a href="https://www.linkedin.com/in/kartik894/">Kartik Sathyanarayanan</a>, <a href="https://www.linkedin.com/in/rajiv-shringi/">Rajiv Shringi</a></p><h3>Introduction</h3><p>Netflix has a diverse range of graph use cases, each serving specific business needs with unique functionality and performance requirements. These use cases fall into two broad categories:</p><ol><li><strong>OLAP</strong>: These use cases typically involve open-ended and algorithmic exploration of large graph datasets. They often utilize industry-standard models and languages such as RDF with SPARQL, Property Graphs with Gremlin or openCypher, and even SQL. The primary focus in these situations is in-depth analysis, rather than achieving high throughput and low latency.</li><li><strong>OLTP</strong>: These use cases require extremely high throughput — up to millions of operations per second — while delivering traversal results within milliseconds. Achieving such a level of performance often requires making trade-offs, which can include accepting eventual consistency or restricting query complexity. For example, the service can demand a specified starting point for traversals and enforce a maximum traversal depth. Such use cases are often directly tied to streaming or user experiences and demand high global availability.</li></ol><p>Netflix’s Graph Abstraction was designed specifically for this second category of use cases. As of this writing, the abstraction is handling close to 10 million operations per second across 650 TB of graph datasets with low latency and cost efficiency.</p><p>This post is the first in a multi-part series that explores the Graph Abstraction architecture in depth. We’ll cover how the abstraction indexes data for real-time and historical views, manages strongly typed graphs, performs efficient traversals, and integrates with the Netflix Big Data ecosystem.</p><h3>Usage at Netflix</h3><p>From a business standpoint, the primary driver for developing the Graph Abstraction was internal demand for supporting several key use cases:</p><ul><li><strong>Real-Time Distributed Graph (RDG)</strong>: A graph capturing dynamic relationships across entities and interactions throughout the Netflix ecosystem. You can learn more about the initial RDG implementation in this insightful <a href="https://netflixtechblog.medium.com/how-and-why-netflix-built-a-real-time-distributed-graph-part-2-building-a-scalable-storage-layer-ff4a8dbd3d1f">blog post</a>. This functionality has since been integrated into the Graph Abstraction.</li><li><strong>Social Graph</strong>: A graph of social connections within Netflix Gaming, designed to boost user engagement.</li><li><a href="https://netflixtechblog.com/from-silos-to-service-topology-why-netflix-built-a-real-time-service-map-0165ba13a7bc"><strong>Service Topology</strong></a>: A graph of all internal Netflix services, used for real-time and historical analysis to improve root cause analysis during incidents.</li></ul><p>Let’s examine the overall architecture of the Graph Abstraction and how it integrates with the Netflix Online Datastore ecosystem.</p><h3>Architecture</h3><p>Instead of building the persistence and caching layers from scratch, we chose to build taller on top of existing Netflix data abstractions.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*6XXQ786vx4AwImVynUAU6A.png"></figure><p>The <a href="https://netflixtechblog.com/introducing-netflixs-key-value-data-abstraction-layer-1ea8a0a11b30">Key-Value (KV) Abstraction</a> stores the latest view of nodes and edges, serving as the real-time index for all queries. Optionally, users can plug-in the <a href="https://netflixtechblog.com/introducing-netflix-timeseries-data-abstraction-layer-31552f6326f8">TimeSeries (TS) Abstraction</a> if they are interested in a historical view of how the graph evolves over time. Additionally, we use <a href="https://netflixtechblog.com/announcing-evcache-distributed-in-memory-datastore-for-cloud-c26a698c27f7">EVCache</a> to achieve low-millisecond latencies and are actively experimenting with more specialized caching layers to further improve performance. Finally, the Graph Abstraction integrates with the <a href="https://netflixtechblog.medium.com/data-gateway-a-platform-for-growing-and-protecting-the-data-tier-f1ed8db8f5c6">Data Gateway Control Plane</a> to manage graph schemas and automate the provisioning, deletion, and configuration of datasets in both KV and TS.</p><h3>Property Graph Model</h3><p>The Abstraction uses the <a href="https://en.wikipedia.org/wiki/Property_graph">Property Graph</a> model to store its data. The graph consists of nodes and edges of various types, each with associated properties. These properties are strongly typed to enable efficient filtering and ensure consistent data exports. For semantic reasons, edges can be either unidirectional or bidirectional.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*F9XX6m8w1csTP0erZrS8sQ.png"></figure><h3>Namespaces</h3><p>The Abstraction separates data into isolated units called “namespaces.” Each namespace is associated with a physical storage layer, as configured in the Data Gateway Control Plane, and can be deployed on either dedicated or shared hardware. The optimal, most cost-effective hardware configuration is determined by our <a href="https://github.com/Netflix-Skunkworks/service-capacity-modeling">provisioning automation</a>, based on user-provided requirements such as throughput, latency, dataset size, and workload criticality. For more details on this topic, see this <a href="https://www.youtube.com/watch?v=Lf6B1PxIvAs">talk</a> given by our stunning colleague Joey Lynch at <strong>AWS re:Invent</strong>.</p><h3>Graph Schema</h3><p>Each namespace is further associated with an explicit graph schema configured in the Control Plane. The graph schema defines node and edge types, allowed properties, permitted relationships, and directions.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*_w_kbWs-XaIS5dAjgRpxDw.png"></figure><p>The Graph schema is implemented as a collection of edge mappings that describe the nature of the relationship between given node types.</p><pre>{<br>  "edgeConfig": {<br>    "edgeMappings": [<br>      {<br>        "edgeMappingKey": {<br>          "fromNodeType": "account",<br>          "edgeType": "owns",<br>          "toNodeType": "profile"<br>        },<br>        "directionType": "UNIDIRECTIONAL"<br>      },<br>      {<br>        "edgeMappingKey": {<br>          "fromNodeType": "profile",<br>          "edgeType": "linked_to",<br>          "toNodeType": "device"<br>        },<br>        "directionType": "BIDIRECTIONAL"<br>      }<br>    ]<br>  }<br>}</pre><p>Edge mappings are further extended with specification of property schema that consists of allowed property names and their type specification:</p><pre>{<br>   "edgeMappingKey":{<br>      "fromNodeType":"profile",<br>      "edgeType":"linked_to",<br>      "toNodeType":"device"<br>   },<br>   "propertySchema":{<br>      "propertyMappings":[<br>         { "propertyKey":"registration_time", "propertyValueType":"TIMESTAMP" },<br>         { "propertyKey":"status", "propertyValueType":"STRING" }<br>      ]<br>   }<br>}</pre><p>The Abstraction servers load this schema on startup and build an<em> in-memory metadata graph</em> of possible relationships, enabling several key optimizations:</p><ul><li><strong>Data Quality:</strong> The Abstraction rejects non-conforming nodes, edges, and properties during writes, ensuring high data quality and consistent exports.</li><li><strong>Query Planning:</strong> The Abstraction uses the schema to quickly construct the possible traversal paths the service should take to answer a given user query.</li><li><strong>Deduplication of Traversed Edges:</strong> For bidirectional traversals on edges between the same node type, the schema helps avoid redundant processing by deduplicating traversed paths.</li><li><strong>Eliminating Traversal paths:</strong> For a given user query, the Abstraction removes traversal paths associated with impossible relationships, as well as those where filters or property types are incompatible.</li></ul><p>Further, the Abstraction servers periodically poll the schema from the Data Gateway Control Plane in order to keep it updated with user changes. Looking ahead, we plan to leverage the graph schema for additional improvements, such as:</p><ul><li><strong>Minimizing Query Fanout:</strong> By using edge cardinality within edge mappings, we aim to select the most efficient traversal paths and minimize query fanout.</li><li><strong>Improved Developer Experience:</strong> The schema will support generating a type-safe data access layer and enhance the Gremlin-like API with schema awareness.</li></ul><p>Next, let’s look at how this data is organized in a real-time index within the KV Abstraction.</p><h3>Real-Time Index: Key-Value Storage</h3><p>Before we discuss how the data is organized into graph indexes, let’s discuss how KV organizes data within namespaces and provides idempotency guarantees:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*MYU0GIxQFzZ_BTKCSEeThw.png"></figure><ul><li><strong>Data partitioning: </strong>A namespace is associated with a table in the underlying storage layer. Within the table, data is partitioned into records by unique IDs, with each record holding multiple sorted items as key-value pairs. This structure effectively makes each namespace a map of sorted maps, providing flexibility for diverse access patterns.</li><li><strong>Idempotency</strong>: Writes to a given ID and key are <a href="https://aws.amazon.com/builders-library/making-retries-safe-with-idempotent-APIs/">idempotent</a>, enabling <a href="https://grpc.io/docs/guides/request-hedging/">request hedging</a> and safe retries. The idempotency token contains a timestamp, which KV uses to enforce Last-Write-Wins (LWW) semantics at the storage layer.</li></ul><p>We use the KV as the underlying storage for all real-time graph indices on nodes and edges. For more on Netflix’s Key-Value Abstraction, see this excellent <a href="https://netflixtechblog.com/introducing-netflixs-key-value-data-abstraction-layer-1ea8a0a11b30">post</a> published by our KeyValue team.</p><h3>Node Storage</h3><p>The two-tiered partitioning strategy works well for node storage. Each node type is isolated within its own KV namespace, which stores all the properties for nodes of that type.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*6Fc1-LHALWVO-8G-Z8RSVA.png"></figure><p>This storage format enables several efficient access patterns for nodes:</p><ul><li><strong>Efficient reads</strong>: A given node and all its properties are fetched in a single partition lookup, achieving single-digit millisecond latency.</li><li><strong>Property selection pushdown</strong>: Target property keys are pushed down to the KV layer, reducing the amount of data fetched and further decreasing latencies and network overhead.</li><li><strong>Property filtering pushdown</strong>: Property keys and values can be efficiently filtered at the KV layer.</li><li><strong>Efficient exports</strong>: This model supports highly parallelized node exports by node type.</li></ul><h3>Edge Storage</h3><h4>Links and Property Index</h4><p>Edges utilize two distinct types of indexes: one exclusively for the edge connections (links), and one for edge properties.</p><p>The Edge links are arranged as an adjacency list mapping source nodes to their connected neighbors.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*_euxbuVK9wmitubG1UUG9g.png"></figure><p>The Edge Property index stores information about properties of every edge.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*2d3gRUTmEW7M5swL1jRKcA.png"></figure><p>Separating edge links from their properties brings several benefits, but also introduces a key trade-off:</p><p><strong>Benefits:</strong></p><ul><li><strong>Efficient property upserts:</strong> Allows individual properties to be upserted over time without needing to read the entire property set for an edge.</li><li><strong>Wide row prevention:</strong> Decoupling edge links from their properties prevents large partitions in databases like Cassandra, enabling efficient storage and low-latency reads — even for edges with millions of connections.</li></ul><p><strong>Trade-off:</strong></p><ul><li><strong>Non-atomic writes:</strong> Storing edges across multiple namespaces means that writes across these namespaces are not atomic. We’ll discuss how this is addressed in the Consistency Enforcement section.</li></ul><h4>Forward and Reverse Indexes</h4><p>Additionally, edge indexes are separated into forward and reverse indexes to support traversals in either direction. The illustration below shows an example of the reverse index counterpart for the links namespace shown above.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*vg_5v9wNhi0h9wU1y6zthg.png"></figure><p>To ensure consistent record identifiers when updating edge properties in either direction, the Abstraction lexicographically sorts and concatenates the source and destination node IDs to create a <em>direction-agnostic identifier </em>for property storage. This ensures that properties can be accessed or mutated in a single database call regardless of the direction specified in the request.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*MF15NbNtvcymomLlJlY7Nw.png"></figure><p>This storage format enables several efficient access patterns:</p><ul><li><strong>Point Reads</strong>: Given an edge id, all properties can be fetched in a single partition lookup on the properties index.</li><li><strong>Range Reads</strong>: Given a source node, a range read on a partition in the links index can efficiently return all edges. Depending on the desired direction, the Abstraction can target the forward or reverse index.</li><li><strong>Property Filtering</strong>: Properties are fetched only for the links that match the record or page limit criteria, minimizing the data exchanged over the network.</li><li><strong>Sort Orders</strong>: By default, edge links are sorted lexicographically by their target node. To support fetching the latest connections, the Abstraction retrieves target edge links in memory, sorts them by their last-write time, and returns the results. In order to ensure optimal performance without exerting too much memory pressure, we aim to limit the number of edges per source node within the system.</li></ul><p>Next, let’s explore the caching strategies used by the Abstraction.</p><h3>Caching Strategies in Graph Abstraction</h3><p>Although the Graph Abstraction already provides efficient reads and writes to durable storage, caching remains critical for the stability and performance of any graph datastore for two key reasons:</p><ul><li><strong>Write amplification</strong>: A single write on the fronting service can result in multiple writes to the backing durable storage due to the use of multiple indexes. Whenever possible, it’s best to avoid unnecessary writes — for example, by not writing an edge link that already exists.</li><li><strong>Read amplification</strong>: A single traversal request on the fronting service may translate into thousands of fetch operations on the backend, especially for highly interconnected graphs.</li></ul><p>To address these challenges, the Graph Abstraction employs two distinct caching strategies.</p><h4>Write-aside Caching of Edge Links</h4><p>An edge link contains no additional information beyond the link itself and its last-write timestamp. To reduce write amplification on durable storage, we cache edge links for short durations, helping to avoid writing a link that already exists. This mechanism is balanced with configurable TTL windows, cache invalidation on deletes, and lease acquisitions with exponential backoff. These strategies provide the necessary consistency guarantees while still allowing the last-write timestamp to be refreshed according to the predefined staleness.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*KgtOewuSORVYngTAmHOz3w.png"></figure><h4>Read-aside Caching of Properties</h4><p>To reduce read amplification on the durable store, the Graph Abstraction leverages KV’s integration with EVCache. Multiple KV namespaces can share the same caching clusters for cost efficiency. The Abstraction first fetches data from durable storage, while subsequent reads are served from the cache. Caching is applied at both the record and item levels, benefiting all graph objects.</p><p>Graph Abstraction employs two invalidation strategies, selected based on write throughput and consistency requirements:</p><ul><li><strong>Invalidation on write:</strong> Both record and item caches are invalidated with every write, ensuring consistency across regions. This strategy is ideal for graphs that change infrequently and cannot tolerate data staleness, but comes with the tradeoff of pushing a higher throughput on the cache.</li><li><strong>TTL-driven invalidation:</strong> Cache entries are invalidated only when their TTL expires. This approach works best for frequently modified objects that can tolerate some staleness.</li></ul><h4>Work In Progress: Write-Through Caching</h4><p>We are also developing a write-through caching strategy designed to store most of the data required by the Abstraction during traversals. This caching mechanism can organize indexes by different sort orders (e.g., sorting data by last-write timestamp), at the cost of increased memory consumption. Stay tuned for more details on this approach.</p><p>Next, let’s examine the consistency guarantees in Graph Abstraction and how they are enforced for both reads and writes.</p><h3>Consistency Enforcement</h3><p>Enforcing data consistency in Graph Abstraction poses several challenges. The connected nature of the data, low-latency API requirements, and the need to handle intermittent failures have led to design choices that enforce strict eventual consistency across multiple regions.</p><h4>Entropy Repair</h4><p>Each write in the Abstraction persists data for both inward and outward indices in parallel to support high throughput. Further, each write happens on multiple KV namespaces. To prevent inconsistencies or lasting entropy from failures in any operation, the Abstraction uses a robust retry mechanism using Kafka:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*b-QZ0X3ODaHCbUmVGZ1jFw.png"></figure><h4>Node Deletions</h4><p>Deleting nodes in a highly connected graph is more complex than simply removing a KV record as each node may have thousands of connected edges that must be handled to maintain graph integrity. Further, synchronously deleting all such connections would introduce unacceptable latency for the Abstraction callers.</p><p>The Abstraction employs an asynchronous deletion strategy to manage this issue. The consequence of this approach, however, is that the observed mutated state is only eventually consistent. Further, to ensure correctness of asynchronous deletes during concurrent updates, the Last-Write-Wins (LWW) conflict resolution mechanism is essential.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*AUGcrovbyUWfdNnbk8gPdw.png"></figure><h4>Global Replication</h4><p>The consistency guarantees of Graph Abstraction are shaped by its multi-region availability. As illustrated in the diagram below, both the caching layer and durable storage replicate data asynchronously across regions, resulting in an eventually consistent system.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*xJq5kKpXQD-CQZxpp58Z-w.png"></figure><p>Now that we’ve covered storing the real-time graph index, let’s see how it enables graph traversals.</p><h3>Graph Traversals</h3><p>The Abstraction provides a custom gRPC traversal API, inspired by <a href="https://tinkerpop.apache.org/gremlin.html">Gremlin</a>, which enables exploration of the distributed graph by letting users chain traversals, apply filter criteria, sort results, limit results, and more.</p><p>Let’s explore a hypothetical scenario where the Abstraction is used to recommend shows to users on a shared device, by considering the duration of the most recent viewing session for each show across all profiles and accounts associated with that device:</p><pre>TraversalRequest.newBuilder()<br>  .setNamespace("&lt;graph-namespace&gt;")<br>  .setTraversalQuery(<br>     TraversalQuery.newBuilder()<br>       // Given id of the 'device' node type.<br>       .setStartNode(node("device", "my-device-id"))<br>       .setTraversal(<br>          Traversal.newBuilder()<br>            // fetch the first 5 connections<br>            .setEdgeLimit(5)<br>            .setDirectionTraversal(<br>               DirectionTraversal.newBuilder()<br>                  // traverse in the IN direction<br>                  .setDirection(IN)<br>                  // minimize data exchange: only interested in certain properties<br>                  .addNodePropertiesSelections(propSelection("account", "created_at"))<br>                  .addNodePropertiesSelections(propSelection("profile", "last_active"))<br>                  .setDirectionFilter(<br>                     DirectionFilter.newBuilder()<br>                       // only interested in certain connected types<br>                       .setTypeMatchingStrategy(EXCLUDE_NON_TARGETED)<br>                       .addAllNodeFilters(typeFilters("account", "profile"))))<br>            // chain traversals to the intermediate result<br>            .addNextTraversals(<br>               Traversal.newBuilder()<br>                 .setOrder(LATEST)<br>                 // limit to 200 connections for the 2nd hop<br>                 .setEdgeLimit(200)<br>                 .setDirectionTraversal(<br>                    DirectionTraversal.newBuilder()<br>                      // now traverse in the OUT direction<br>                      .setDirection(OUT)<br>                      .addEdgePropertiesSelections(propSelection("watched", "view_time"))<br>                      .addEdgePropertiesSelections(propSelection("has_plan", "active"))<br>                      .setDirectionFilter(<br>                         DirectionFilter.newBuilder()<br>                           .setTypeMatchingStrategy(EXCLUDE_NON_TARGETED)<br>                           .addAllNodeFilters(typeFilters("title", "plan")))))))<br>  .build();</pre><p>And let’s visualize the intended results set produced by the request above:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*_uEbkvl_zlhcJPHTQnubgg.png"></figure><p>We’ll explore the design and implementation of traversal planning and execution, along with different traversal types, in the <strong>Part II</strong> of this blog series.</p><p>Now let’s look at the performance metrics of Graph Abstraction based on current production use cases.</p><h3>Real World Performance</h3><p>Across all applications at Netflix, Graph Abstraction ensures high availability while processing up to 10 million operations per second across all writes, individual edge / node reads and traversals at peak hours:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*71VBcBOx20n6sEaS6rYRJg.png"></figure><p>Edge and node persistence achieve single-digit millisecond latencies (p99 shown in red, p90 shown in orange, and p50 shown in green):</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*SwBXgPnu_bMeR1nJkQpbsw.png"></figure><p>Traversal performance depends on the number of hops, the edge fanout at each stage, and associated filters and sort orders. We parallelize work as much as possible to reduce latencies. Typically 1-hop traversals are executed with single-digit millisecond latency:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*CW1f9ULedUK7u_W7hsz3Qw.png"><figcaption>1-hop traversal latencies</figcaption></figure><p>We also support a Count API that performs counting traversals at a very high rate with similar latencies, which we will cover in Part II of this series:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*wUCzYRbui9BRuL42uCdL8Q.png"></figure><p>Currently, the RDG is powered by 2-hop traversals with a higher degree of fan-out. While these operations can reach upwards of 100 ms in latency, the 90th percentile (p90) latency remains under 50ms.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*fzuOdqu0RmEza6cK44v48A.png"><figcaption>2-hop traversal latencies</figcaption></figure><p>We track the average and max edge fanout at different depths to give us insights into the traversal performance for different graph datasets.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*96pqab2LqCK3-RSAe3qcEQ.png"><figcaption>Median edge fan-out</figcaption></figure><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*OHOoHxkkC2EXvTRNgyRM7w.png"><figcaption>Max edge fan-out</figcaption></figure><p>Asynchronous operations such as node deletions can be slightly latent, but typically perform with sub-second latency:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*lWev5g0OStQycd3Ebo5arw.png"></figure><p>At the moment, we are storing close to 650 TB of data globally across all our graph datasets.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*y7lKTbAqi9Lid7dSAOY_Mw.png"></figure><h3>Conclusion</h3><p>As Netflix scales further into new verticals such as live content, games, and ads, Graph Abstraction will remain crucial for uncovering and leveraging rich connections — while continuing to support a high throughput and availability at low latencies.</p><p>Stay tuned for <strong>Part II</strong> of this blog series, where we’ll explore the implementation of graph traversals, counting and constraint mechanisms.</p><p>In <strong>Part III</strong>, we’ll take a closer look at the temporal index implementation and its integration with the Time Series Abstraction.</p><h3>Acknowledgments</h3><p>Special thanks to our stunning colleagues who contributed to Graph Abstraction’s success: <a href="https://www.linkedin.com/in/kaidanfullerton/">Kaidan Fullerton</a>, <a href="https://www.linkedin.com/in/joseph-lynch-9976a431/">Joey Lynch</a>, <a href="https://www.linkedin.com/in/sudheshsuresh/">Sudhesh Suresh</a>, <a href="https://www.linkedin.com/in/vinaychella/">Vinay Chella</a>, <a href="https://www.linkedin.com/in/sumanth-pasupuleti/">Sumanth Pasupuleti</a>, <a href="https://www.linkedin.com/in/vidhya-arvind-11908723/">Vidhya Arvind</a>, <a href="https://www.linkedin.com/in/rummadis/">Raj Ummadisetty</a>, <a href="https://www.linkedin.com/in/jordan-west-8aa1731a3/">Jordan West</a>, <a href="https://www.linkedin.com/in/clohfink/">Chris Lohfink</a>, <a href="https://www.linkedin.com/in/joe-lee-a70661a2/">Joe Lee</a>, <a href="https://www.linkedin.com/in/jingxi-huang/">Jingxi Huang</a>, <a href="https://www.linkedin.com/in/jessicaswalton/">Jessica Walton</a>, <a href="https://www.linkedin.com/in/prudhviraj9/">Prudhviraj Karumanchi</a>, <a href="https://www.linkedin.com/in/akashdeepgoel/">Akashdeep Goel</a>, <a href="https://www.linkedin.com/in/sriram-rangarajan-35169715/">Sriram Rangarajan</a>, <a href="https://www.linkedin.com/in/chrisvanvlack/">Chris Van Vlack</a>, <a href="https://www.linkedin.com/in/chrisleegray/">Christopher Gray</a>, <a href="https://www.linkedin.com/in/lu4nm3/">Luis Medina</a>, <a href="https://www.linkedin.com/in/ajitkoti/">Ajit Koti</a>, <a href="https://www.linkedin.com/in/mohidul-abedin">Mohidul Abedin</a>.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=e88063e6f6d5" width="1" height="1" alt=""><hr><p><a href="https://medium.com/netflix-techblog/high-throughput-graph-abstraction-at-netflix-part-i-e88063e6f6d5">High-Throughput Graph Abstraction at Netflix: Part I</a> was originally published in <a href="https://medium.com/netflix-techblog">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://medium.com/netflix-techblog/high-throughput-graph-abstraction-at-netflix-part-i-e88063e6f6d5</link>
      <guid>https://medium.com/netflix-techblog/high-throughput-graph-abstraction-at-netflix-part-i-e88063e6f6d5</guid>
      <pubDate>Fri, 29 May 2026 20:49:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[High-Throughput Graph Abstraction at Netflix: Part I]]></title>
      <description><![CDATA[<p>By <a href="https://www.linkedin.com/in/oleksii-tkachuk-98b47375/">Oleksii Tkachuk</a>, <a href="https://www.linkedin.com/in/kartik894/">Kartik Sathyanarayanan</a>, <a href="https://www.linkedin.com/in/rajiv-shringi/">Rajiv Shringi</a></p><h3>Introduction</h3><p>Netflix has a diverse range of graph use cases, each serving specific business needs with unique functionality and performance requirements. These use cases fall into two broad categories:</p><ol><li><strong>OLAP</strong>: These use cases typically involve open-ended and algorithmic exploration of large graph datasets. They often utilize industry-standard models and languages such as RDF with SPARQL, Property Graphs with Gremlin or openCypher, and even SQL. The primary focus in these situations is in-depth analysis, rather than achieving high throughput and low latency.</li><li><strong>OLTP</strong>: These use cases require extremely high throughput — up to millions of operations per second — while delivering traversal results within milliseconds. Achieving such a level of performance often requires making trade-offs, which can include accepting eventual consistency or restricting query complexity. For example, the service can demand a specified starting point for traversals and enforce a maximum traversal depth. Such use cases are often directly tied to streaming or user experiences and demand high global availability.</li></ol><p>Netflix’s Graph Abstraction was designed specifically for this second category of use cases. As of this writing, the abstraction is handling close to 10 million operations per second across 650 TB of graph datasets with low latency and cost efficiency.</p><p>This post is the first in a multi-part series that explores the Graph Abstraction architecture in depth. We’ll cover how the abstraction indexes data for real-time and historical views, manages strongly typed graphs, performs efficient traversals, and integrates with the Netflix Big Data ecosystem.</p><h3>Usage at Netflix</h3><p>From a business standpoint, the primary driver for developing the Graph Abstraction was internal demand for supporting several key use cases:</p><ul><li><strong>Real-Time Distributed Graph (RDG)</strong>: A graph capturing dynamic relationships across entities and interactions throughout the Netflix ecosystem. You can learn more about the initial RDG implementation in this insightful <a href="https://netflixtechblog.medium.com/how-and-why-netflix-built-a-real-time-distributed-graph-part-2-building-a-scalable-storage-layer-ff4a8dbd3d1f">blog post</a>. This functionality has since been integrated into the Graph Abstraction.</li><li><strong>Social Graph</strong>: A graph of social connections within Netflix Gaming, designed to boost user engagement.</li><li><a href="https://netflixtechblog.com/from-silos-to-service-topology-why-netflix-built-a-real-time-service-map-0165ba13a7bc"><strong>Service Topology</strong></a>: A graph of all internal Netflix services, used for real-time and historical analysis to improve root cause analysis during incidents.</li></ul><p>Let’s examine the overall architecture of the Graph Abstraction and how it integrates with the Netflix Online Datastore ecosystem.</p><h3>Architecture</h3><p>Instead of building the persistence and caching layers from scratch, we chose to build taller on top of existing Netflix data abstractions.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*6XXQ786vx4AwImVynUAU6A.png"></figure><p>The <a href="https://netflixtechblog.com/introducing-netflixs-key-value-data-abstraction-layer-1ea8a0a11b30">Key-Value (KV) Abstraction</a> stores the latest view of nodes and edges, serving as the real-time index for all queries. Optionally, users can plug-in the <a href="https://netflixtechblog.com/introducing-netflix-timeseries-data-abstraction-layer-31552f6326f8">TimeSeries (TS) Abstraction</a> if they are interested in a historical view of how the graph evolves over time. Additionally, we use <a href="https://netflixtechblog.com/announcing-evcache-distributed-in-memory-datastore-for-cloud-c26a698c27f7">EVCache</a> to achieve low-millisecond latencies and are actively experimenting with more specialized caching layers to further improve performance. Finally, the Graph Abstraction integrates with the <a href="https://netflixtechblog.medium.com/data-gateway-a-platform-for-growing-and-protecting-the-data-tier-f1ed8db8f5c6">Data Gateway Control Plane</a> to manage graph schemas and automate the provisioning, deletion, and configuration of datasets in both KV and TS.</p><h3>Property Graph Model</h3><p>The Abstraction uses the <a href="https://en.wikipedia.org/wiki/Property_graph">Property Graph</a> model to store its data. The graph consists of nodes and edges of various types, each with associated properties. These properties are strongly typed to enable efficient filtering and ensure consistent data exports. For semantic reasons, edges can be either unidirectional or bidirectional.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*F9XX6m8w1csTP0erZrS8sQ.png"></figure><h3>Namespaces</h3><p>The Abstraction separates data into isolated units called “namespaces.” Each namespace is associated with a physical storage layer, as configured in the Data Gateway Control Plane, and can be deployed on either dedicated or shared hardware. The optimal, most cost-effective hardware configuration is determined by our <a href="https://github.com/Netflix-Skunkworks/service-capacity-modeling">provisioning automation</a>, based on user-provided requirements such as throughput, latency, dataset size, and workload criticality. For more details on this topic, see this <a href="https://www.youtube.com/watch?v=Lf6B1PxIvAs">talk</a> given by our stunning colleague Joey Lynch at <strong>AWS re:Invent</strong>.</p><h3>Graph Schema</h3><p>Each namespace is further associated with an explicit graph schema configured in the Control Plane. The graph schema defines node and edge types, allowed properties, permitted relationships, and directions.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*_w_kbWs-XaIS5dAjgRpxDw.png"></figure><p>The Graph schema is implemented as a collection of edge mappings that describe the nature of the relationship between given node types.</p><pre>{<br>  "edgeConfig": {<br>    "edgeMappings": [<br>      {<br>        "edgeMappingKey": {<br>          "fromNodeType": "account",<br>          "edgeType": "owns",<br>          "toNodeType": "profile"<br>        },<br>        "directionType": "UNIDIRECTIONAL"<br>      },<br>      {<br>        "edgeMappingKey": {<br>          "fromNodeType": "profile",<br>          "edgeType": "linked_to",<br>          "toNodeType": "device"<br>        },<br>        "directionType": "BIDIRECTIONAL"<br>      }<br>    ]<br>  }<br>}</pre><p>Edge mappings are further extended with specification of property schema that consists of allowed property names and their type specification:</p><pre>{<br>   "edgeMappingKey":{<br>      "fromNodeType":"profile",<br>      "edgeType":"linked_to",<br>      "toNodeType":"device"<br>   },<br>   "propertySchema":{<br>      "propertyMappings":[<br>         { "propertyKey":"registration_time", "propertyValueType":"TIMESTAMP" },<br>         { "propertyKey":"status", "propertyValueType":"STRING" }<br>      ]<br>   }<br>}</pre><p>The Abstraction servers load this schema on startup and build an<em> in-memory metadata graph</em> of possible relationships, enabling several key optimizations:</p><ul><li><strong>Data Quality:</strong> The Abstraction rejects non-conforming nodes, edges, and properties during writes, ensuring high data quality and consistent exports.</li><li><strong>Query Planning:</strong> The Abstraction uses the schema to quickly construct the possible traversal paths the service should take to answer a given user query.</li><li><strong>Deduplication of Traversed Edges:</strong> For bidirectional traversals on edges between the same node type, the schema helps avoid redundant processing by deduplicating traversed paths.</li><li><strong>Eliminating Traversal paths:</strong> For a given user query, the Abstraction removes traversal paths associated with impossible relationships, as well as those where filters or property types are incompatible.</li></ul><p>Further, the Abstraction servers periodically poll the schema from the Data Gateway Control Plane in order to keep it updated with user changes. Looking ahead, we plan to leverage the graph schema for additional improvements, such as:</p><ul><li><strong>Minimizing Query Fanout:</strong> By using edge cardinality within edge mappings, we aim to select the most efficient traversal paths and minimize query fanout.</li><li><strong>Improved Developer Experience:</strong> The schema will support generating a type-safe data access layer and enhance the Gremlin-like API with schema awareness.</li></ul><p>Next, let’s look at how this data is organized in a real-time index within the KV Abstraction.</p><h3>Real-Time Index: Key-Value Storage</h3><p>Before we discuss how the data is organized into graph indexes, let’s discuss how KV organizes data within namespaces and provides idempotency guarantees:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*MYU0GIxQFzZ_BTKCSEeThw.png"></figure><ul><li><strong>Data partitioning: </strong>A namespace is associated with a table in the underlying storage layer. Within the table, data is partitioned into records by unique IDs, with each record holding multiple sorted items as key-value pairs. This structure effectively makes each namespace a map of sorted maps, providing flexibility for diverse access patterns.</li><li><strong>Idempotency</strong>: Writes to a given ID and key are <a href="https://aws.amazon.com/builders-library/making-retries-safe-with-idempotent-APIs/">idempotent</a>, enabling <a href="https://grpc.io/docs/guides/request-hedging/">request hedging</a> and safe retries. The idempotency token contains a timestamp, which KV uses to enforce Last-Write-Wins (LWW) semantics at the storage layer.</li></ul><p>We use the KV as the underlying storage for all real-time graph indices on nodes and edges. For more on Netflix’s Key-Value Abstraction, see this excellent <a href="https://netflixtechblog.com/introducing-netflixs-key-value-data-abstraction-layer-1ea8a0a11b30">post</a> published by our KeyValue team.</p><h3>Node Storage</h3><p>The two-tiered partitioning strategy works well for node storage. Each node type is isolated within its own KV namespace, which stores all the properties for nodes of that type.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*6Fc1-LHALWVO-8G-Z8RSVA.png"></figure><p>This storage format enables several efficient access patterns for nodes:</p><ul><li><strong>Efficient reads</strong>: A given node and all its properties are fetched in a single partition lookup, achieving single-digit millisecond latency.</li><li><strong>Property selection pushdown</strong>: Target property keys are pushed down to the KV layer, reducing the amount of data fetched and further decreasing latencies and network overhead.</li><li><strong>Property filtering pushdown</strong>: Property keys and values can be efficiently filtered at the KV layer.</li><li><strong>Efficient exports</strong>: This model supports highly parallelized node exports by node type.</li></ul><h3>Edge Storage</h3><h4>Links and Property Index</h4><p>Edges utilize two distinct types of indexes: one exclusively for the edge connections (links), and one for edge properties.</p><p>The Edge links are arranged as an adjacency list mapping source nodes to their connected neighbors.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*_euxbuVK9wmitubG1UUG9g.png"></figure><p>The Edge Property index stores information about properties of every edge.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*2d3gRUTmEW7M5swL1jRKcA.png"></figure><p>Separating edge links from their properties brings several benefits, but also introduces a key trade-off:</p><p><strong>Benefits:</strong></p><ul><li><strong>Efficient property upserts:</strong> Allows individual properties to be upserted over time without needing to read the entire property set for an edge.</li><li><strong>Wide row prevention:</strong> Decoupling edge links from their properties prevents large partitions in databases like Cassandra, enabling efficient storage and low-latency reads — even for edges with millions of connections.</li></ul><p><strong>Trade-off:</strong></p><ul><li><strong>Non-atomic writes:</strong> Storing edges across multiple namespaces means that writes across these namespaces are not atomic. We’ll discuss how this is addressed in the Consistency Enforcement section.</li></ul><h4>Forward and Reverse Indexes</h4><p>Additionally, edge indexes are separated into forward and reverse indexes to support traversals in either direction. The illustration below shows an example of the reverse index counterpart for the links namespace shown above.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*vg_5v9wNhi0h9wU1y6zthg.png"></figure><p>To ensure consistent record identifiers when updating edge properties in either direction, the Abstraction lexicographically sorts and concatenates the source and destination node IDs to create a <em>direction-agnostic identifier </em>for property storage. This ensures that properties can be accessed or mutated in a single database call regardless of the direction specified in the request.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*MF15NbNtvcymomLlJlY7Nw.png"></figure><p>This storage format enables several efficient access patterns:</p><ul><li><strong>Point Reads</strong>: Given an edge id, all properties can be fetched in a single partition lookup on the properties index.</li><li><strong>Range Reads</strong>: Given a source node, a range read on a partition in the links index can efficiently return all edges. Depending on the desired direction, the Abstraction can target the forward or reverse index.</li><li><strong>Property Filtering</strong>: Properties are fetched only for the links that match the record or page limit criteria, minimizing the data exchanged over the network.</li><li><strong>Sort Orders</strong>: By default, edge links are sorted lexicographically by their target node. To support fetching the latest connections, the Abstraction retrieves target edge links in memory, sorts them by their last-write time, and returns the results. In order to ensure optimal performance without exerting too much memory pressure, we aim to limit the number of edges per source node within the system.</li></ul><p>Next, let’s explore the caching strategies used by the Abstraction.</p><h3>Caching Strategies in Graph Abstraction</h3><p>Although the Graph Abstraction already provides efficient reads and writes to durable storage, caching remains critical for the stability and performance of any graph datastore for two key reasons:</p><ul><li><strong>Write amplification</strong>: A single write on the fronting service can result in multiple writes to the backing durable storage due to the use of multiple indexes. Whenever possible, it’s best to avoid unnecessary writes — for example, by not writing an edge link that already exists.</li><li><strong>Read amplification</strong>: A single traversal request on the fronting service may translate into thousands of fetch operations on the backend, especially for highly interconnected graphs.</li></ul><p>To address these challenges, the Graph Abstraction employs two distinct caching strategies.</p><h4>Write-aside Caching of Edge Links</h4><p>An edge link contains no additional information beyond the link itself and its last-write timestamp. To reduce write amplification on durable storage, we cache edge links for short durations, helping to avoid writing a link that already exists. This mechanism is balanced with configurable TTL windows, cache invalidation on deletes, and lease acquisitions with exponential backoff. These strategies provide the necessary consistency guarantees while still allowing the last-write timestamp to be refreshed according to the predefined staleness.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*KgtOewuSORVYngTAmHOz3w.png"></figure><h4>Read-aside Caching of Properties</h4><p>To reduce read amplification on the durable store, the Graph Abstraction leverages KV’s integration with EVCache. Multiple KV namespaces can share the same caching clusters for cost efficiency. The Abstraction first fetches data from durable storage, while subsequent reads are served from the cache. Caching is applied at both the record and item levels, benefiting all graph objects.</p><p>Graph Abstraction employs two invalidation strategies, selected based on write throughput and consistency requirements:</p><ul><li><strong>Invalidation on write:</strong> Both record and item caches are invalidated with every write, ensuring consistency across regions. This strategy is ideal for graphs that change infrequently and cannot tolerate data staleness, but comes with the tradeoff of pushing a higher throughput on the cache.</li><li><strong>TTL-driven invalidation:</strong> Cache entries are invalidated only when their TTL expires. This approach works best for frequently modified objects that can tolerate some staleness.</li></ul><h4>Work In Progress: Write-Through Caching</h4><p>We are also developing a write-through caching strategy designed to store most of the data required by the Abstraction during traversals. This caching mechanism can organize indexes by different sort orders (e.g., sorting data by last-write timestamp), at the cost of increased memory consumption. Stay tuned for more details on this approach.</p><p>Next, let’s examine the consistency guarantees in Graph Abstraction and how they are enforced for both reads and writes.</p><h3>Consistency Enforcement</h3><p>Enforcing data consistency in Graph Abstraction poses several challenges. The connected nature of the data, low-latency API requirements, and the need to handle intermittent failures have led to design choices that enforce strict eventual consistency across multiple regions.</p><h4>Entropy Repair</h4><p>Each write in the Abstraction persists data for both inward and outward indices in parallel to support high throughput. Further, each write happens on multiple KV namespaces. To prevent inconsistencies or lasting entropy from failures in any operation, the Abstraction uses a robust retry mechanism using Kafka:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*b-QZ0X3ODaHCbUmVGZ1jFw.png"></figure><h4>Node Deletions</h4><p>Deleting nodes in a highly connected graph is more complex than simply removing a KV record as each node may have thousands of connected edges that must be handled to maintain graph integrity. Further, synchronously deleting all such connections would introduce unacceptable latency for the Abstraction callers.</p><p>The Abstraction employs an asynchronous deletion strategy to manage this issue. The consequence of this approach, however, is that the observed mutated state is only eventually consistent. Further, to ensure correctness of asynchronous deletes during concurrent updates, the Last-Write-Wins (LWW) conflict resolution mechanism is essential.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*AUGcrovbyUWfdNnbk8gPdw.png"></figure><h4>Global Replication</h4><p>The consistency guarantees of Graph Abstraction are shaped by its multi-region availability. As illustrated in the diagram below, both the caching layer and durable storage replicate data asynchronously across regions, resulting in an eventually consistent system.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*xJq5kKpXQD-CQZxpp58Z-w.png"></figure><p>Now that we’ve covered storing the real-time graph index, let’s see how it enables graph traversals.</p><h3>Graph Traversals</h3><p>The Abstraction provides a custom gRPC traversal API, inspired by <a href="https://tinkerpop.apache.org/gremlin.html">Gremlin</a>, which enables exploration of the distributed graph by letting users chain traversals, apply filter criteria, sort results, limit results, and more.</p><p>Let’s explore a hypothetical scenario where the Abstraction is used to recommend shows to users on a shared device, by considering the duration of the most recent viewing session for each show across all profiles and accounts associated with that device:</p><pre>TraversalRequest.newBuilder()<br>  .setNamespace("&lt;graph-namespace&gt;")<br>  .setTraversalQuery(<br>     TraversalQuery.newBuilder()<br>       // Given id of the 'device' node type.<br>       .setStartNode(node("device", "my-device-id"))<br>       .setTraversal(<br>          Traversal.newBuilder()<br>            // fetch the first 5 connections<br>            .setEdgeLimit(5)<br>            .setDirectionTraversal(<br>               DirectionTraversal.newBuilder()<br>                  // traverse in the IN direction<br>                  .setDirection(IN)<br>                  // minimize data exchange: only interested in certain properties<br>                  .addNodePropertiesSelections(propSelection("account", "created_at"))<br>                  .addNodePropertiesSelections(propSelection("profile", "last_active"))<br>                  .setDirectionFilter(<br>                     DirectionFilter.newBuilder()<br>                       // only interested in certain connected types<br>                       .setTypeMatchingStrategy(EXCLUDE_NON_TARGETED)<br>                       .addAllNodeFilters(typeFilters("account", "profile"))))<br>            // chain traversals to the intermediate result<br>            .addNextTraversals(<br>               Traversal.newBuilder()<br>                 .setOrder(LATEST)<br>                 // limit to 200 connections for the 2nd hop<br>                 .setEdgeLimit(200)<br>                 .setDirectionTraversal(<br>                    DirectionTraversal.newBuilder()<br>                      // now traverse in the OUT direction<br>                      .setDirection(OUT)<br>                      .addEdgePropertiesSelections(propSelection("watched", "view_time"))<br>                      .addEdgePropertiesSelections(propSelection("has_plan", "active"))<br>                      .setDirectionFilter(<br>                         DirectionFilter.newBuilder()<br>                           .setTypeMatchingStrategy(EXCLUDE_NON_TARGETED)<br>                           .addAllNodeFilters(typeFilters("title", "plan")))))))<br>  .build();</pre><p>And let’s visualize the intended results set produced by the request above:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*_uEbkvl_zlhcJPHTQnubgg.png"></figure><p>We’ll explore the design and implementation of traversal planning and execution, along with different traversal types, in the <strong>Part II</strong> of this blog series.</p><p>Now let’s look at the performance metrics of Graph Abstraction based on current production use cases.</p><h3>Real World Performance</h3><p>Across all applications at Netflix, Graph Abstraction ensures high availability while processing up to 10 million operations per second across all writes, individual edge / node reads and traversals at peak hours:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*71VBcBOx20n6sEaS6rYRJg.png"></figure><p>Edge and node persistence achieve single-digit millisecond latencies (p99 shown in red, p90 shown in orange, and p50 shown in green):</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*SwBXgPnu_bMeR1nJkQpbsw.png"></figure><p>Traversal performance depends on the number of hops, the edge fanout at each stage, and associated filters and sort orders. We parallelize work as much as possible to reduce latencies. Typically 1-hop traversals are executed with single-digit millisecond latency:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*CW1f9ULedUK7u_W7hsz3Qw.png"><figcaption>1-hop traversal latencies</figcaption></figure><p>We also support a Count API that performs counting traversals at a very high rate with similar latencies, which we will cover in Part II of this series:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*wUCzYRbui9BRuL42uCdL8Q.png"></figure><p>Currently, the RDG is powered by 2-hop traversals with a higher degree of fan-out. While these operations can reach upwards of 100 ms in latency, the 90th percentile (p90) latency remains under 50ms.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*fzuOdqu0RmEza6cK44v48A.png"><figcaption>2-hop traversal latencies</figcaption></figure><p>We track the average and max edge fanout at different depths to give us insights into the traversal performance for different graph datasets.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*96pqab2LqCK3-RSAe3qcEQ.png"><figcaption>Median edge fan-out</figcaption></figure><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*OHOoHxkkC2EXvTRNgyRM7w.png"><figcaption>Max edge fan-out</figcaption></figure><p>Asynchronous operations such as node deletions can be slightly latent, but typically perform with sub-second latency:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*lWev5g0OStQycd3Ebo5arw.png"></figure><p>At the moment, we are storing close to 650 TB of data globally across all our graph datasets.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*y7lKTbAqi9Lid7dSAOY_Mw.png"></figure><h3>Conclusion</h3><p>As Netflix scales further into new verticals such as live content, games, and ads, Graph Abstraction will remain crucial for uncovering and leveraging rich connections — while continuing to support a high throughput and availability at low latencies.</p><p>Stay tuned for <strong>Part II</strong> of this blog series, where we’ll explore the implementation of graph traversals, counting and constraint mechanisms.</p><p>In <strong>Part III</strong>, we’ll take a closer look at the temporal index implementation and its integration with the Time Series Abstraction.</p><h3>Acknowledgments</h3><p>Special thanks to our stunning colleagues who contributed to Graph Abstraction’s success: <a href="https://www.linkedin.com/in/kaidanfullerton/">Kaidan Fullerton</a>, <a href="https://www.linkedin.com/in/joseph-lynch-9976a431/">Joey Lynch</a>, <a href="https://www.linkedin.com/in/sudheshsuresh/">Sudhesh Suresh</a>, <a href="https://www.linkedin.com/in/vinaychella/">Vinay Chella</a>, <a href="https://www.linkedin.com/in/sumanth-pasupuleti/">Sumanth Pasupuleti</a>, <a href="https://www.linkedin.com/in/vidhya-arvind-11908723/">Vidhya Arvind</a>, <a href="https://www.linkedin.com/in/rummadis/">Raj Ummadisetty</a>, <a href="https://www.linkedin.com/in/jordan-west-8aa1731a3/">Jordan West</a>, <a href="https://www.linkedin.com/in/clohfink/">Chris Lohfink</a>, <a href="https://www.linkedin.com/in/joe-lee-a70661a2/">Joe Lee</a>, <a href="https://www.linkedin.com/in/jingxi-huang/">Jingxi Huang</a>, <a href="https://www.linkedin.com/in/jessicaswalton/">Jessica Walton</a>, <a href="https://www.linkedin.com/in/prudhviraj9/">Prudhviraj Karumanchi</a>, <a href="https://www.linkedin.com/in/akashdeepgoel/">Akashdeep Goel</a>, <a href="https://www.linkedin.com/in/sriram-rangarajan-35169715/">Sriram Rangarajan</a>, <a href="https://www.linkedin.com/in/chrisvanvlack/">Chris Van Vlack</a>, <a href="https://www.linkedin.com/in/chrisleegray/">Christopher Gray</a>, <a href="https://www.linkedin.com/in/lu4nm3/">Luis Medina</a>, <a href="https://www.linkedin.com/in/ajitkoti/">Ajit Koti</a>, <a href="https://www.linkedin.com/in/mohidul-abedin">Mohidul Abedin</a>.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=e88063e6f6d5" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/high-throughput-graph-abstraction-at-netflix-part-i-e88063e6f6d5">High-Throughput Graph Abstraction at Netflix: Part I</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/high-throughput-graph-abstraction-at-netflix-part-i-e88063e6f6d5</link>
      <guid>https://netflixtechblog.com/high-throughput-graph-abstraction-at-netflix-part-i-e88063e6f6d5</guid>
      <pubDate>Fri, 29 May 2026 20:49:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[From Silos to Service Topology: Why Netflix Built a Real-Time Service Map]]></title>
      <description><![CDATA[<p><em>By </em><a href="https://www.linkedin.com/in/parth-jain-8a09abb6/"><em>Parth Jain</em></a>, <a href="https://www.linkedin.com/in/raskuma/"><em>Rakesh Sukumar</em></a><em>, </em><a href="https://www.linkedin.com/in/yingwu-zhao-62037418/"><em>Yingwu Zhao</em></a><em>, </em><a href="https://www.linkedin.com/in/renzosanchezsilva/"><em>Renzo Sanchez</em></a><em> &amp; </em><a href="https://www.linkedin.com/in/nathfisher/"><em>Nathan Fisher</em></a><em><br>How we built a living map of our distributed infrastructure to help engineers understand dependencies, troubleshoot faster, and keep Netflix running smoothly for our members around the world.</em></p><h3>The Puzzle with a Thousand Pieces</h3><p>Picture this: It’s 3am, and an engineer gets paged. One of our critical services is showing elevated error rates. Members trying to watch their favorite films and series are seeing degraded experiences. The clock is ticking.</p><figure><img alt="A central service node connected to multiple downstream services and data stores, illustrating the tangled dependency graph engineers must navigate without a service topology map." src="https://cdn-images-1.medium.com/max/799/1*t_oo96UujawlqwGtmBBm7Q.png"><figcaption>A single service at the center of a web of dependencies — services, data stores, and call chains branching in every direction. Without a unified map, engineers have to reason about this structure from memory and scattered signals.</figcaption></figure><p>In a system with thousands of microservices supporting our entertainment experience for members worldwide, answering these questions quickly can mean the difference between a minor blip and a major incident.</p><p>We kept hearing variations of this story from engineers across Netflix. The tooling gap was clear: we had plenty of signals, but no unified way to understand how everything connected.</p><h3>The Three Questions Every Engineer Asks</h3><p>When troubleshooting distributed systems, engineers fundamentally need to understand relationships:</p><p><strong>Which services depend on each other?</strong> Not just theoretical dependencies from configuration files or architecture diagrams, but actual runtime connections based on real traffic.</p><p><strong>What’s the blast radius?</strong> When something breaks or needs to go down for maintenance, what else will be affected? Which teams need to be notified?</p><p><strong>Where’s the source?</strong> Is my problem caused by an upstream issue, or am I the root cause that’s cascading to others?</p><p>Traditional observability tools show fragments of this picture. Metrics show symptoms and performance characteristics. Logs show individual service behavior. Traces show single request flows through the system. But none of them show the complete map of how everything connects — the steady-state topology of dependencies that forms the backbone of our distributed architecture.</p><p>For an engineer at 3am, having to mentally stitch together information from multiple tools is slow, error-prone, and stressful. We needed something better: a unified view of service dependencies — a map showing how everything connects — with easy navigation to the detailed signals when you need to dig deeper.</p><h3>Why This Matters More Than Ever</h3><p>Netflix runs on thousands of microservices working together to deliver entertainment to our members. When you press play on your favorite series, that single action triggers a cascade of service-to-service calls — authentication, recommendations tailored to your tastes, video encoding selection, playback optimization, and more.</p><p>This architecture gives us tremendous flexibility and allows hundreds of engineering teams to innovate independently. But it also creates fundamental observability challenges.</p><p>And these challenges were growing. New initiatives like our Live programming and Ads-supported plans require even more sophisticated monitoring and faster troubleshooting. Live events can’t wait for lengthy incident investigations. The scale and real-time nature of these systems demanded better tooling.</p><p>We analyzed thousands of support requests from our engineers over a four-year period. The patterns were consistent:</p><ul><li>“What are my upstream and downstream dependencies?”</li><li>“Is this failure in my service, or is something I depend on broken?”</li><li>“Which services will be impacted if I take this down for maintenance?”</li><li>“Why is this service showing as ‘Unknown’ in my metrics?”</li><li>“What changed in my call path recently that could explain this behavior?”</li></ul><p>Engineers were asking dependency questions constantly. We needed to provide answers — quickly, accurately, and in real-time.</p><h3>Building on What We Learned</h3><p>We didn’t start from scratch. Over the years, we explored various approaches to solving this problem — from evaluating external graph databases and vendor platforms to building internal prototypes with different storage technologies and data models.</p><h4>Each iteration taught us something valuable:</h4><p><strong>Real-time matters:</strong> Dependency maps that are hours old are useless in dynamic environments where services deploy multiple times per day. We needed near real-time updates.</p><p><strong>Scale changes everything:</strong> Solutions that work at modest scale hit fundamental walls at Netflix scale. Storage systems that handle thousands of nodes struggle with our service count and traffic volume.</p><p><strong>Integration is key:</strong> Any solution needs seamless integration with our existing observability ecosystem. Engineers shouldn’t have to learn entirely new tools or leave their existing workflows.</p><p><strong>Data quality is critical:</strong> Incomplete or incorrect dependency information is worse than no information — it leads to wrong conclusions during incidents.</p><p><strong>Multiple perspectives needed: </strong>We learned that no single source of dependency information tells the complete story. Network connectivity data lacks application context. Application metrics only cover instrumented services. We needed to combine multiple sources.</p><p>These lessons shaped every decision we made in building Service Topology.</p><h3>What We Needed: A Living Map</h3><p>We set out to build something specific: a living map of our infrastructure — one that updates in real-time as services deploy, as traffic patterns shift, as new dependencies form and old ones disappear.</p><p>The requirements were clear:</p><p><strong>Real-time updates, not stale snapshots:</strong> In an environment where services deploy continuously, yesterday’s topology map is archaeology, not observability.</p><p><strong>Fast queries at scale:</strong> When an engineer is troubleshooting at 3am, they can’t wait minutes for a query to return. We needed sub-second response times for traversing the call graph.</p><p><strong>Multiple layers:</strong> Network-level connectivity doesn’t tell the whole story. We needed to see both the network layer (what’s actually talking to what) and the application layer (which APIs and endpoints are being called).</p><p><strong>Rich context, not just connections:</strong> Knowing Service A talks to Service B isn’t enough. We needed to overlay health status, availability tiers, business domains, ownership information, and other metadata to make the information actionable.</p><p><strong>Visual and programmatic access:</strong> Engineers needed a UI for exploration and troubleshooting. But automated systems — resilience frameworks, blast radius calculators, incident response automation — needed programmatic API access.</p><h3>Our Approach: Three Sources of Truth</h3><figure><img alt="Three topology layers side by side: eBPF flow logs producing a network graph, IPC metrics producing an application graph, and distributed traces producing a request graph, all feeding into a unified view." src="https://cdn-images-1.medium.com/max/996/1*68tW9-7kC3oGL0T5J5di_A.png"><figcaption>Three data sources produce three independent topology graphs — network, application, and request — each stored separately and queryable on their own or merged into a single unified view.</figcaption></figure><p>Here’s the key insight we arrived at: no single source tells the complete story.</p><p>We built Service Topology by using three complementary sources to build separate dependency graphs — one from each perspective — that can be combined into a unified view or explored independently:</p><p>Each source creates its own graph that is physically separate — the network layer in one graph database partition, the IPC layer in another partition, and the tracing layer using columnar storage optimized for analytical queries. This physical separation allows each layer to evolve independently and be queried in parallel. When users request a unified view, we execute traversal queries across all layers simultaneously and merge results, achieving sub-second response times even when combining all three layers.</p><p>Each source creates its own graph of service relationships:</p><h4>1. eBPF Network Flows (Network Layer)</h4><p>We capture network flow records at the kernel level using eBPF technology — information about which services are connecting to which other services over the network. This gives us ground truth about actual network-level communication.</p><p>The value: Comprehensive coverage. Every service shows up here because we’re capturing actual network traffic, regardless of whether applications are instrumented. This layer provides topology at both cluster-level (which deployment clusters are communicating) and app-level (which applications are communicating).</p><p>The limitation: Network-level information lacks application context. We know Service A connected to Service B’s IP address using a specific protocol, but not which specific API endpoint or path was called (e.g., /api/v1/users vs /api/v1/orders).</p><h4>2. IPC Metrics (Application Layer)</h4><p>We collect Inter-Process Communication metrics from our instrumented services. These are the metrics applications emit when they make calls to other services via gRPC, GraphQL, REST, or other protocols.</p><p>The value: Rich application context. We can see which specific endpoints were called, error rates, latency distributions, protocol details, and request/response characteristics. This layer provides app-level topology — since IPC metrics are emitted by applications, the natural granularity is application-to-application connections with endpoint details.</p><p>The limitation: Only works for instrumented services. If a service doesn’t emit IPC metrics, we won’t see its application-level calls this way.</p><h4>3. End-to-End Tracing (Request Layer)</h4><p>We integrate distributed tracing information that follows individual requests as they flow through our system. We aggregate traces to build a unified topology graph, but also allow engineers to overlay individual traces on the topology to see specific request flows.</p><p>The value: Shows actual request paths. Not just “Service A <em>can</em> call Service B,” but “Service A <em>did</em> call Service B as part of serving this specific member request.” This captures runtime behavior, including conditional logic and feature flags. Engineers can both see the aggregated pattern and drill into individual traces. We aggregate traces to build topology at both cluster-level and app-level, allowing engineers to view request patterns at the granularity most useful for their investigation.</p><p>The limitation: Sampling. We can’t trace every request without impacting performance, so we sample. This is excellent for understanding common flows, but may miss rarely-used code paths in the aggregated view.</p><h4>Bringing It Together: Multi-Layer Architecture</h4><p>Here’s what makes this powerful: we build three separate graphs — one from each source — that create different perspectives on service relationships:</p><ul><li><strong>Network graph from eBPF flows:</strong> Every connection, regardless of instrumentation</li><li><strong>Application graph from IPC metrics:</strong> Rich endpoint and protocol details</li><li><strong>Request graph from tracing:</strong> Actual runtime behavior and call paths</li></ul><p>Engineers can:</p><ul><li>View each graph independently to focus on a specific perspective (pure network connectivity, application-level calls, or traced request flows)</li><li>Combine them into a unified graph by querying multiple partitions in parallel and merging results — our system returns the union of nodes and edges from all requested layers while preserving each layer’s distinct properties</li></ul><p>The unified view is especially powerful because:</p><ul><li>Network flows ensure completeness — we don’t miss anything</li><li>IPC metrics provide application details — we understand the “how” and “what”</li><li>Tracing shows actual behavior — we see real request patterns</li></ul><p>Each source compensates for the limitations of the others. The result is a comprehensive, accurate, and contextualized view of service dependencies that can be explored from multiple angles.</p><h3>From Flows to Graph: How We Built It</h3><p>Here’s the high-level architecture (we’ll dive deeper into engineering challenges in our next post):</p><figure><img alt="Pipeline diagram showing data flowing from a message stream through Stage 1 initial aggregation, Stage 2 intermediary resolution, and Stage 3 persistence and enrichment into a graph database, then exposed via an API." src="https://cdn-images-1.medium.com/max/1024/1*bvSG8r3B-fffrr-2ZKCtpA.png"><figcaption>Flow logs travel from multi-region Kafka through three aggregation stages — initial batching, intermediary resolution, and final enrichment — before being persisted to the graph database and served via API.</figcaption></figure><p><strong>Multi-Region Ingestion:</strong> We consume flow logs from Kafka across multiple AWS regions where Netflix operates. This runs continuously, processing millions of flow records as they arrive.</p><p><strong>Distributed Processing:</strong> We use Apache Pekko Streams (a fork of Akka) to process these flows in a distributed, fault-tolerant pipeline. The system automatically partitions work across our Auto Scaling Groups to handle the volume and provides natural backpressure handling.</p><p><strong>Three-Stage Distributed Aggregation</strong>: We aggregate network flows through a three-stage pipeline that solves a fundamental challenge: network flow logs only show individual network hops through intermediaries (App A → Load Balancer → App B, or App A → NAT Gateway → App B), not the true application-level connections we need (App A → App B).</p><figure><img alt="Before and after diagram showing intermediary resolution: raw flow logs recording two hops from App A through a load balancer to App B are collapsed into a single direct edge from App A to App B." src="https://cdn-images-1.medium.com/max/292/1*UcZvGrHzMq6geyat9MZv7g.png"><figcaption>Stage 2 resolves network intermediaries: raw flow logs show two separate hops (App A → Load Balancer → App B), but the resolved graph stores the direct application-to-application relationship (App A → App B).</figcaption></figure><p>Stage 1 performs initial aggregation from Kafka. Stage 2 applies resolution logic — identifying network intermediaries (load balancers, NAT gateways, API gateways, proxies) and combining their incoming and outgoing flows to reconstruct direct application-to-application paths. Stage 3 performs final aggregation with health status integration before graph persistence. This graduated approach also prevents hot spots by distributing load across multiple points even when specific applications or network intermediaries see 100x more traffic than others.</p><p>Graph Storage: We persist the topology in <a href="https://netflixtechblog.medium.com/high-throughput-graph-abstraction-at-netflix-part-i-e88063e6f6d5">Netflix’s graph database</a>, an abstraction layer built on top of our distributed key-value storage infrastructure. This graph database is specifically designed for high-throughput graph operations at our scale, with fast multi-hop traversal capabilities. Each of our three data sources (network flows, IPC metrics, tracing) creates a separate graph that can be queried independently or merged.</p><p>gRPC API: We expose the topology through a gRPC service that supports multi-hop traversal, filtering by availability tier and business domain, pagination for large result sets, and sub-second query response times.</p><p>The technical details of building this at Netflix scale — handling Kafka lag, managing memory and garbage collection, optimizing distributed processing, debugging reactive streams — deserve their own discussion. We learned a lot, and we’ll share those lessons in our next post.</p><h3>What Engineers Can Do Now</h3><p>Today, the service topology map is helping engineers across Netflix:</p><p><strong>Visualize Dependencies:</strong> See upstream and downstream dependencies for any service, with the ability to filter by availability tier (Tier 0, Tier 1, etc.) and business domain. Choose between the unified view (combining all sources) or individual graph views (network-only, IPC-only, or trace-only) depending on what you’re investigating.</p><p><strong>Jump to Detailed Signals: </strong>From any service in the topology, quickly navigate to logs, traces, and detailed metrics in their respective tools. No more hunting for the right service name or time window — the topology provides the context and the starting point.</p><p><strong>Understand Blast Radius:</strong> Before taking a service down for maintenance or making significant changes, see exactly what will be impacted. Identify which teams to notify and what to monitor.</p><p><strong>Overlay Health Status:</strong> See not just the topology, but which services in the call path are experiencing issues. This is integrated with health status tracking, so you can quickly identify if a problem you’re seeing is actually originating somewhere else.</p><p><strong>Query Programmatically:</strong> Use our gRPC API to integrate topology information into automated systems. For example, our Platform Modernization Engineering team uses this to verify that critical Live services have proper availability tier classifications throughout their dependency chains.</p><p><strong>Investigate Faster:</strong> During incidents, quickly identify if a failure is local or if it’s propagating from somewhere else in the call graph. Follow the failure pattern to find the root cause.</p><p><strong>Plan Changes Confidently:</strong> Understand the impact of proposed architectural changes or service migrations before implementing them.</p><p><strong>Time Travel Through Topology:</strong> Query what the topology looked like at specific points in the past. Understand what changed in dependencies around the time an issue started, or see how your service’s dependency footprint has evolved over time. This time-travel capability is powered by time-window aggregation — instead of storing every time slice separately, we use layer-specific aggregators that accumulate topology data across windows, allowing us to reconstruct historical views efficiently without exploding storage costs.</p><h3>The Living Map: Always Current</h3><p>What makes this truly useful is that it’s a living map. It’s not a static diagram drawn in a design document that goes out of date the moment it’s published. It’s continuously updated based on actual traffic:</p><ul><li>When a new service starts calling an API, it appears in the topology with near real-time freshness</li><li>When a service stops making calls to a dependency, that edge fades from the graph</li><li>When services deploy and their behavior changes, the topology reflects it</li><li>When incidents impact service health, the status overlay updates in real-time</li></ul><p>This means engineers can trust what they see. The map reflects reality, not someone’s idea of what the architecture should be.</p><h3>The Journey Continues</h3><p>We’re not done. We continue to evolve the system with new capabilities:</p><p>Change Event Overlay: We’re working to surface deployment events, configuration changes, and other mutations alongside the topology graph. Correlation becomes easier when you can see both the dependencies and what changed when.</p><p>Richer Context: As we expand coverage and integrate more signals, we continue to enrich the topology with additional endpoint-level details, protocol information, and network path context.</p><p>And looking further ahead, we’re excited about something bigger: Automated root cause analysis. Imagine an intelligent agent that continuously crawls the topology graph, correlates failures across dependencies, understands historical patterns, and surfaces likely root causes automatically. Service topology provides the knowledge graph foundation that makes this kind of intelligent automation possible.</p><h3>Why This Matters for Our Members</h3><p>This might seem like infrastructure — plumbing that our members never see directly. But it matters immensely to their experience.</p><p>When engineers can quickly understand dependencies and identify issues, incidents get resolved faster. When we can model blast radius before making changes, we avoid disruptions. When automated systems can query dependency information programmatically, we can build smarter, more resilient systems.</p><p>All of this translates to what matters most: our members getting to watch their favorite films and series, seamlessly, whenever they want. Whether it’s a weekend binge of a beloved show, a live sports event, or discovering something new through our recommendations tailored to their tastes — we want it to just work.</p><h3>What’s Next in This Series</h3><p>This is the first in a series of posts about building Service Topology at Netflix.</p><p>In our next post, we’ll pull back the curtain on the engineering challenges we faced at scale: How do you handle Kafka consumer lag when ingesting millions of flow logs per second? What happens when distributed processing meets garbage collection pauses? How do you debug reactive streams that stall under load? How do you manage hot nodes in a distributed system? We’ll share the real problems we hit in production and the solutions we developed.</p><p>In future posts, we’ll explore the lessons we learned that apply to any distributed system at scale, and where we’re heading next with time travel capabilities and Automated root cause analysis.</p><h3>Acknowledgements</h3><p><em>This post was written by </em><a href="https://www.linkedin.com/in/parth-jain-8a09abb6/"><em>Parth Jain</em></a><em>.</em></p><p><em>Service Topology was built by </em><a href="https://www.linkedin.com/in/parth-jain-8a09abb6/"><em>Parth Jain</em></a><em>, </em><a href="https://www.linkedin.com/in/raskuma/"><em>Rakesh Sukumar</em></a><em>, </em><a href="https://www.linkedin.com/in/yingwu-zhao-62037418/"><em>Yingwu Zhao</em></a><em>, </em><a href="https://www.linkedin.com/in/renzosanchezsilva/"><em>Renzo Sanchez-Silva</em></a><em>, and </em><a href="https://www.linkedin.com/in/nathfisher/"><em>Nathan Fisher</em></a><em>.</em></p><p><em>Special thanks to the many engineers across Netflix who made this possible — the Observability team who built the broader system, the graph database platform team who provided the storage foundation, and the Platform Modernization Engineering, Live, and Ads teams who provided invaluable feedback and use cases throughout development.</em></p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=0165ba13a7bc" width="1" height="1" alt=""><hr><p><a href="https://medium.com/netflix-techblog/from-silos-to-service-topology-why-netflix-built-a-real-time-service-map-0165ba13a7bc">From Silos to Service Topology: Why Netflix Built a Real-Time Service Map</a> was originally published in <a href="https://medium.com/netflix-techblog">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://medium.com/netflix-techblog/from-silos-to-service-topology-why-netflix-built-a-real-time-service-map-0165ba13a7bc</link>
      <guid>https://medium.com/netflix-techblog/from-silos-to-service-topology-why-netflix-built-a-real-time-service-map-0165ba13a7bc</guid>
      <pubDate>Fri, 29 May 2026 16:01:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[From Silos to Service Topology: Why Netflix Built a Real-Time Service Map]]></title>
      <description><![CDATA[<p><em>By </em><a href="https://www.linkedin.com/in/parth-jain-8a09abb6/"><em>Parth Jain</em></a>, <a href="https://www.linkedin.com/in/raskuma/"><em>Rakesh Sukumar</em></a><em>, </em><a href="https://www.linkedin.com/in/yingwu-zhao-62037418/"><em>Yingwu Zhao</em></a><em>, </em><a href="https://www.linkedin.com/in/renzosanchezsilva/"><em>Renzo Sanchez</em></a><em> &amp; </em><a href="https://www.linkedin.com/in/nathfisher/"><em>Nathan Fisher</em></a><em><br>How we built a living map of our distributed infrastructure to help engineers understand dependencies, troubleshoot faster, and keep Netflix running smoothly for our members around the world.</em></p><h3>The Puzzle with a Thousand Pieces</h3><p>Picture this: It’s 3am, and an engineer gets paged. One of our critical services is showing elevated error rates. Members trying to watch their favorite films and series are seeing degraded experiences. The clock is ticking.</p><figure><img alt="A central service node connected to multiple downstream services and data stores, illustrating the tangled dependency graph engineers must navigate without a service topology map." src="https://cdn-images-1.medium.com/max/799/1*t_oo96UujawlqwGtmBBm7Q.png"><figcaption>A single service at the center of a web of dependencies — services, data stores, and call chains branching in every direction. Without a unified map, engineers have to reason about this structure from memory and scattered signals.</figcaption></figure><p>In a system with thousands of microservices supporting our entertainment experience for members worldwide, answering these questions quickly can mean the difference between a minor blip and a major incident.</p><p>We kept hearing variations of this story from engineers across Netflix. The tooling gap was clear: we had plenty of signals, but no unified way to understand how everything connected.</p><h3>The Three Questions Every Engineer Asks</h3><p>When troubleshooting distributed systems, engineers fundamentally need to understand relationships:</p><p><strong>Which services depend on each other?</strong> Not just theoretical dependencies from configuration files or architecture diagrams, but actual runtime connections based on real traffic.</p><p><strong>What’s the blast radius?</strong> When something breaks or needs to go down for maintenance, what else will be affected? Which teams need to be notified?</p><p><strong>Where’s the source?</strong> Is my problem caused by an upstream issue, or am I the root cause that’s cascading to others?</p><p>Traditional observability tools show fragments of this picture. Metrics show symptoms and performance characteristics. Logs show individual service behavior. Traces show single request flows through the system. But none of them show the complete map of how everything connects — the steady-state topology of dependencies that forms the backbone of our distributed architecture.</p><p>For an engineer at 3am, having to mentally stitch together information from multiple tools is slow, error-prone, and stressful. We needed something better: a unified view of service dependencies — a map showing how everything connects — with easy navigation to the detailed signals when you need to dig deeper.</p><h3>Why This Matters More Than Ever</h3><p>Netflix runs on thousands of microservices working together to deliver entertainment to our members. When you press play on your favorite series, that single action triggers a cascade of service-to-service calls — authentication, recommendations tailored to your tastes, video encoding selection, playback optimization, and more.</p><p>This architecture gives us tremendous flexibility and allows hundreds of engineering teams to innovate independently. But it also creates fundamental observability challenges.</p><p>And these challenges were growing. New initiatives like our Live programming and Ads-supported plans require even more sophisticated monitoring and faster troubleshooting. Live events can’t wait for lengthy incident investigations. The scale and real-time nature of these systems demanded better tooling.</p><p>We analyzed thousands of support requests from our engineers over a four-year period. The patterns were consistent:</p><ul><li>“What are my upstream and downstream dependencies?”</li><li>“Is this failure in my service, or is something I depend on broken?”</li><li>“Which services will be impacted if I take this down for maintenance?”</li><li>“Why is this service showing as ‘Unknown’ in my metrics?”</li><li>“What changed in my call path recently that could explain this behavior?”</li></ul><p>Engineers were asking dependency questions constantly. We needed to provide answers — quickly, accurately, and in real-time.</p><h3>Building on What We Learned</h3><p>We didn’t start from scratch. Over the years, we explored various approaches to solving this problem — from evaluating external graph databases and vendor platforms to building internal prototypes with different storage technologies and data models.</p><h4>Each iteration taught us something valuable:</h4><p><strong>Real-time matters:</strong> Dependency maps that are hours old are useless in dynamic environments where services deploy multiple times per day. We needed near real-time updates.</p><p><strong>Scale changes everything:</strong> Solutions that work at modest scale hit fundamental walls at Netflix scale. Storage systems that handle thousands of nodes struggle with our service count and traffic volume.</p><p><strong>Integration is key:</strong> Any solution needs seamless integration with our existing observability ecosystem. Engineers shouldn’t have to learn entirely new tools or leave their existing workflows.</p><p><strong>Data quality is critical:</strong> Incomplete or incorrect dependency information is worse than no information — it leads to wrong conclusions during incidents.</p><p><strong>Multiple perspectives needed: </strong>We learned that no single source of dependency information tells the complete story. Network connectivity data lacks application context. Application metrics only cover instrumented services. We needed to combine multiple sources.</p><p>These lessons shaped every decision we made in building Service Topology.</p><h3>What We Needed: A Living Map</h3><p>We set out to build something specific: a living map of our infrastructure — one that updates in real-time as services deploy, as traffic patterns shift, as new dependencies form and old ones disappear.</p><p>The requirements were clear:</p><p><strong>Real-time updates, not stale snapshots:</strong> In an environment where services deploy continuously, yesterday’s topology map is archaeology, not observability.</p><p><strong>Fast queries at scale:</strong> When an engineer is troubleshooting at 3am, they can’t wait minutes for a query to return. We needed sub-second response times for traversing the call graph.</p><p><strong>Multiple layers:</strong> Network-level connectivity doesn’t tell the whole story. We needed to see both the network layer (what’s actually talking to what) and the application layer (which APIs and endpoints are being called).</p><p><strong>Rich context, not just connections:</strong> Knowing Service A talks to Service B isn’t enough. We needed to overlay health status, availability tiers, business domains, ownership information, and other metadata to make the information actionable.</p><p><strong>Visual and programmatic access:</strong> Engineers needed a UI for exploration and troubleshooting. But automated systems — resilience frameworks, blast radius calculators, incident response automation — needed programmatic API access.</p><h3>Our Approach: Three Sources of Truth</h3><figure><img alt="Three topology layers side by side: eBPF flow logs producing a network graph, IPC metrics producing an application graph, and distributed traces producing a request graph, all feeding into a unified view." src="https://cdn-images-1.medium.com/max/996/1*68tW9-7kC3oGL0T5J5di_A.png"><figcaption>Three data sources produce three independent topology graphs — network, application, and request — each stored separately and queryable on their own or merged into a single unified view.</figcaption></figure><p>Here’s the key insight we arrived at: no single source tells the complete story.</p><p>We built Service Topology by using three complementary sources to build separate dependency graphs — one from each perspective — that can be combined into a unified view or explored independently:</p><p>Each source creates its own graph that is physically separate — the network layer in one graph database partition, the IPC layer in another partition, and the tracing layer using columnar storage optimized for analytical queries. This physical separation allows each layer to evolve independently and be queried in parallel. When users request a unified view, we execute traversal queries across all layers simultaneously and merge results, achieving sub-second response times even when combining all three layers.</p><p>Each source creates its own graph of service relationships:</p><h4>1. eBPF Network Flows (Network Layer)</h4><p>We capture network flow records at the kernel level using eBPF technology — information about which services are connecting to which other services over the network. This gives us ground truth about actual network-level communication.</p><p>The value: Comprehensive coverage. Every service shows up here because we’re capturing actual network traffic, regardless of whether applications are instrumented. This layer provides topology at both cluster-level (which deployment clusters are communicating) and app-level (which applications are communicating).</p><p>The limitation: Network-level information lacks application context. We know Service A connected to Service B’s IP address using a specific protocol, but not which specific API endpoint or path was called (e.g., /api/v1/users vs /api/v1/orders).</p><h4>2. IPC Metrics (Application Layer)</h4><p>We collect Inter-Process Communication metrics from our instrumented services. These are the metrics applications emit when they make calls to other services via gRPC, GraphQL, REST, or other protocols.</p><p>The value: Rich application context. We can see which specific endpoints were called, error rates, latency distributions, protocol details, and request/response characteristics. This layer provides app-level topology — since IPC metrics are emitted by applications, the natural granularity is application-to-application connections with endpoint details.</p><p>The limitation: Only works for instrumented services. If a service doesn’t emit IPC metrics, we won’t see its application-level calls this way.</p><h4>3. End-to-End Tracing (Request Layer)</h4><p>We integrate distributed tracing information that follows individual requests as they flow through our system. We aggregate traces to build a unified topology graph, but also allow engineers to overlay individual traces on the topology to see specific request flows.</p><p>The value: Shows actual request paths. Not just “Service A <em>can</em> call Service B,” but “Service A <em>did</em> call Service B as part of serving this specific member request.” This captures runtime behavior, including conditional logic and feature flags. Engineers can both see the aggregated pattern and drill into individual traces. We aggregate traces to build topology at both cluster-level and app-level, allowing engineers to view request patterns at the granularity most useful for their investigation.</p><p>The limitation: Sampling. We can’t trace every request without impacting performance, so we sample. This is excellent for understanding common flows, but may miss rarely-used code paths in the aggregated view.</p><h4>Bringing It Together: Multi-Layer Architecture</h4><p>Here’s what makes this powerful: we build three separate graphs — one from each source — that create different perspectives on service relationships:</p><ul><li><strong>Network graph from eBPF flows:</strong> Every connection, regardless of instrumentation</li><li><strong>Application graph from IPC metrics:</strong> Rich endpoint and protocol details</li><li><strong>Request graph from tracing:</strong> Actual runtime behavior and call paths</li></ul><p>Engineers can:</p><ul><li>View each graph independently to focus on a specific perspective (pure network connectivity, application-level calls, or traced request flows)</li><li>Combine them into a unified graph by querying multiple partitions in parallel and merging results — our system returns the union of nodes and edges from all requested layers while preserving each layer’s distinct properties</li></ul><p>The unified view is especially powerful because:</p><ul><li>Network flows ensure completeness — we don’t miss anything</li><li>IPC metrics provide application details — we understand the “how” and “what”</li><li>Tracing shows actual behavior — we see real request patterns</li></ul><p>Each source compensates for the limitations of the others. The result is a comprehensive, accurate, and contextualized view of service dependencies that can be explored from multiple angles.</p><h3>From Flows to Graph: How We Built It</h3><p>Here’s the high-level architecture (we’ll dive deeper into engineering challenges in our next post):</p><figure><img alt="Pipeline diagram showing data flowing from a message stream through Stage 1 initial aggregation, Stage 2 intermediary resolution, and Stage 3 persistence and enrichment into a graph database, then exposed via an API." src="https://cdn-images-1.medium.com/max/1024/1*bvSG8r3B-fffrr-2ZKCtpA.png"><figcaption>Flow logs travel from multi-region Kafka through three aggregation stages — initial batching, intermediary resolution, and final enrichment — before being persisted to the graph database and served via API.</figcaption></figure><p><strong>Multi-Region Ingestion:</strong> We consume flow logs from Kafka across multiple AWS regions where Netflix operates. This runs continuously, processing millions of flow records as they arrive.</p><p><strong>Distributed Processing:</strong> We use Apache Pekko Streams (a fork of Akka) to process these flows in a distributed, fault-tolerant pipeline. The system automatically partitions work across our Auto Scaling Groups to handle the volume and provides natural backpressure handling.</p><p><strong>Three-Stage Distributed Aggregation</strong>: We aggregate network flows through a three-stage pipeline that solves a fundamental challenge: network flow logs only show individual network hops through intermediaries (App A → Load Balancer → App B, or App A → NAT Gateway → App B), not the true application-level connections we need (App A → App B).</p><figure><img alt="Before and after diagram showing intermediary resolution: raw flow logs recording two hops from App A through a load balancer to App B are collapsed into a single direct edge from App A to App B." src="https://cdn-images-1.medium.com/max/292/1*UcZvGrHzMq6geyat9MZv7g.png"><figcaption>Stage 2 resolves network intermediaries: raw flow logs show two separate hops (App A → Load Balancer → App B), but the resolved graph stores the direct application-to-application relationship (App A → App B).</figcaption></figure><p>Stage 1 performs initial aggregation from Kafka. Stage 2 applies resolution logic — identifying network intermediaries (load balancers, NAT gateways, API gateways, proxies) and combining their incoming and outgoing flows to reconstruct direct application-to-application paths. Stage 3 performs final aggregation with health status integration before graph persistence. This graduated approach also prevents hot spots by distributing load across multiple points even when specific applications or network intermediaries see 100x more traffic than others.</p><p>Graph Storage: We persist the topology in Netflix’s graph database, an abstraction layer built on top of our distributed key-value storage infrastructure. This graph database is specifically designed for high-throughput graph operations at our scale, with fast multi-hop traversal capabilities. Each of our three data sources (network flows, IPC metrics, tracing) creates a separate graph that can be queried independently or merged.</p><p>gRPC API: We expose the topology through a gRPC service that supports multi-hop traversal, filtering by availability tier and business domain, pagination for large result sets, and sub-second query response times.</p><p>The technical details of building this at Netflix scale — handling Kafka lag, managing memory and garbage collection, optimizing distributed processing, debugging reactive streams — deserve their own discussion. We learned a lot, and we’ll share those lessons in our next post.</p><h3>What Engineers Can Do Now</h3><p>Today, the service topology map is helping engineers across Netflix:</p><p><strong>Visualize Dependencies:</strong> See upstream and downstream dependencies for any service, with the ability to filter by availability tier (Tier 0, Tier 1, etc.) and business domain. Choose between the unified view (combining all sources) or individual graph views (network-only, IPC-only, or trace-only) depending on what you’re investigating.</p><p><strong>Jump to Detailed Signals: </strong>From any service in the topology, quickly navigate to logs, traces, and detailed metrics in their respective tools. No more hunting for the right service name or time window — the topology provides the context and the starting point.</p><p><strong>Understand Blast Radius:</strong> Before taking a service down for maintenance or making significant changes, see exactly what will be impacted. Identify which teams to notify and what to monitor.</p><p><strong>Overlay Health Status:</strong> See not just the topology, but which services in the call path are experiencing issues. This is integrated with health status tracking, so you can quickly identify if a problem you’re seeing is actually originating somewhere else.</p><p><strong>Query Programmatically:</strong> Use our gRPC API to integrate topology information into automated systems. For example, our Platform Modernization Engineering team uses this to verify that critical Live services have proper availability tier classifications throughout their dependency chains.</p><p><strong>Investigate Faster:</strong> During incidents, quickly identify if a failure is local or if it’s propagating from somewhere else in the call graph. Follow the failure pattern to find the root cause.</p><p><strong>Plan Changes Confidently:</strong> Understand the impact of proposed architectural changes or service migrations before implementing them.</p><p><strong>Time Travel Through Topology:</strong> Query what the topology looked like at specific points in the past. Understand what changed in dependencies around the time an issue started, or see how your service’s dependency footprint has evolved over time. This time-travel capability is powered by time-window aggregation — instead of storing every time slice separately, we use layer-specific aggregators that accumulate topology data across windows, allowing us to reconstruct historical views efficiently without exploding storage costs.</p><h3>The Living Map: Always Current</h3><p>What makes this truly useful is that it’s a living map. It’s not a static diagram drawn in a design document that goes out of date the moment it’s published. It’s continuously updated based on actual traffic:</p><ul><li>When a new service starts calling an API, it appears in the topology with near real-time freshness</li><li>When a service stops making calls to a dependency, that edge fades from the graph</li><li>When services deploy and their behavior changes, the topology reflects it</li><li>When incidents impact service health, the status overlay updates in real-time</li></ul><p>This means engineers can trust what they see. The map reflects reality, not someone’s idea of what the architecture should be.</p><h3>The Journey Continues</h3><p>We’re not done. We continue to evolve the system with new capabilities:</p><p>Change Event Overlay: We’re working to surface deployment events, configuration changes, and other mutations alongside the topology graph. Correlation becomes easier when you can see both the dependencies and what changed when.</p><p>Richer Context: As we expand coverage and integrate more signals, we continue to enrich the topology with additional endpoint-level details, protocol information, and network path context.</p><p>And looking further ahead, we’re excited about something bigger: Automated root cause analysis. Imagine an intelligent agent that continuously crawls the topology graph, correlates failures across dependencies, understands historical patterns, and surfaces likely root causes automatically. Service topology provides the knowledge graph foundation that makes this kind of intelligent automation possible.</p><h3>Why This Matters for Our Members</h3><p>This might seem like infrastructure — plumbing that our members never see directly. But it matters immensely to their experience.</p><p>When engineers can quickly understand dependencies and identify issues, incidents get resolved faster. When we can model blast radius before making changes, we avoid disruptions. When automated systems can query dependency information programmatically, we can build smarter, more resilient systems.</p><p>All of this translates to what matters most: our members getting to watch their favorite films and series, seamlessly, whenever they want. Whether it’s a weekend binge of a beloved show, a live sports event, or discovering something new through our recommendations tailored to their tastes — we want it to just work.</p><h3>What’s Next in This Series</h3><p>This is the first in a series of posts about building Service Topology at Netflix.</p><p>In our next post, we’ll pull back the curtain on the engineering challenges we faced at scale: How do you handle Kafka consumer lag when ingesting millions of flow logs per second? What happens when distributed processing meets garbage collection pauses? How do you debug reactive streams that stall under load? How do you manage hot nodes in a distributed system? We’ll share the real problems we hit in production and the solutions we developed.</p><p>In future posts, we’ll explore the lessons we learned that apply to any distributed system at scale, and where we’re heading next with time travel capabilities and Automated root cause analysis.</p><h3>Acknowledgements</h3><p><em>This post was written by </em><a href="https://www.linkedin.com/in/parth-jain-8a09abb6/"><em>Parth Jain</em></a><em>.</em></p><p><em>Service Topology was built by </em><a href="https://www.linkedin.com/in/parth-jain-8a09abb6/"><em>Parth Jain</em></a><em>, </em><a href="https://www.linkedin.com/in/raskuma/"><em>Rakesh Sukumar</em></a><em>, </em><a href="https://www.linkedin.com/in/yingwu-zhao-62037418/"><em>Yingwu Zhao</em></a><em>, </em><a href="https://www.linkedin.com/in/renzosanchezsilva/"><em>Renzo Sanchez-Silva</em></a><em>, and </em><a href="https://www.linkedin.com/in/nathfisher/"><em>Nathan Fisher</em></a><em>.</em></p><p><em>Special thanks to the many engineers across Netflix who made this possible — the Observability team who built the broader system, the graph database platform team who provided the storage foundation, and the Platform Modernization Engineering, Live, and Ads teams who provided invaluable feedback and use cases throughout development.</em></p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=0165ba13a7bc" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/from-silos-to-service-topology-why-netflix-built-a-real-time-service-map-0165ba13a7bc">From Silos to Service Topology: Why Netflix Built a Real-Time Service Map</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/from-silos-to-service-topology-why-netflix-built-a-real-time-service-map-0165ba13a7bc</link>
      <guid>https://netflixtechblog.com/from-silos-to-service-topology-why-netflix-built-a-real-time-service-map-0165ba13a7bc</guid>
      <pubDate>Fri, 29 May 2026 16:01:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Scaling ArchUnit with Nebula ArchRules]]></title>
      <description><![CDATA[<p>By <a href="https://github.com/wakingrufus">John Burns</a> and <a href="https://www.linkedin.com/in/emilyyuan03/">Emily Yuan</a></p><h3>Introduction</h3><p>At Netflix, we operate using a <a href="https://netflixtechblog.com/towards-true-continuous-integration-distributed-repositories-and-dependencies-2a2e3108c051">polyrepo</a> strategy with tens of thousands of Java repositories. This means that we need to have ways of sharing common build logic across these repositories. On the <a href="https://sites.google.com/netflix.com/javaplatformnetflix/jvm">JVM Ecosystem team</a> within Java Platform, we build tooling such as the <a href="https://github.com/nebula-plugins">Nebula suite of Gradle plugins</a> to provide standard ways to build projects, keep dependencies up-to-date, and publish artifacts reliably across the Java ecosystem. Our mission also entails providing build-time feedback to the developer when they deviate from the <a href="https://netflixtechblog.com/how-we-build-code-at-netflix-c5d9bd727f15">paved road</a>, or when their code base contains technical debt.</p><h3>Case Study</h3><p>After a Netflix incident relating to a library releasing a backwards-incompatible change, our team was asked to provide some tooling and practices to improve the Java library lifecycle management. This was not a simple case of a library making a reckless breaking change. The code removed had been deprecated for years. Library authors often struggle to know when it is safe to remove deprecated code, or refactor code that is not meant to be used by downstream applications. Fleet-wide migrations, such as upgrading major Spring Boot versions, also involve deprecated code removal. To help with this, we established a suite of API lifecycle annotations:</p><ul><li>@Deprecated from the Java standard library</li><li>@Public A custom annotation to use on APIs meant to be used downstream</li><li>@Experimental A custom annotation for new APIs which may not yet be stable</li><li>All other APIs are assumed to be “internal”</li></ul><p>Library authors can annotate their APIs with these annotations. However, how will they know which downstream projects are using their API incorrectly, based on these?</p><p>As we sought to improve the paved road for JVM-based libraries at Netflix, we needed a good way of identifying this kind of technical debt, not only for the benefit of the Java Platform-provided libraries, but any team delivering shared libraries to the organization. For this, we looked at ArchUnit.</p><p><a href="https://www.archunit.org/">ArchUnit</a> is a popular OSS library (3.5k stars, 84 contributors) used to enforce “architectural” code rules as part of a JUnit suite. It is used internally by Gradle, Spring, and is provided as part of the <a href="https://spring.io/projects/spring-modulith">Spring Modulith</a> platform. The rules engine, which is built directly on top of <a href="https://asm.ow2.io/">ASM</a>, can be used for a wide variety of use cases. It is powerful enough to be a general purpose static analysis tool with the following distinctive features:</p><p>1. Works cross-language (JVM), because it uses ASM/bytecode, not AST parsing.</p><p>2. Exposes a builder API pattern that makes it easy to write rules</p><p>3. Also has a lower level API ideal for writing more complex custom rules.</p><p>The limitation of ArchUnit is that it is designed to be used as part of a JUnit suite in a single repository. The Nebula ArchRules plugins give organizations the ability to share and apply rules across any number of repositories. Rules can be sourced from OSS libraries or private internal libraries. This makes the plugin generally useful for any JVM+Gradle engineering organization.</p><h3>Why ArchUnit?</h3><p>Before we go into how ArchRules works, it is good to understand why we would want to use ArchUnit in this way instead of other static analysis tools.</p><h4>AST vs Bytecode</h4><p>Some tools, such as PMD, process rules against an AST (abstract syntax tree). An AST is a structured representation of source code. This kind of tool will have rules that are syntax dependent. Rules that need to support multiple JVM languages, such as Kotlin or Scala, often need to be rewritten for each language. It also allows code which should be found to be hidden under syntactic sugar not anticipated by the rule author. ArchUnit uses <a href="https://asm.ow2.io/">ASM</a> to analyze actual compiled bytecode, which means it doesn’t matter how that code was produced. What is analyzed is the actual code that will be run.</p><h4>Rule Authorship</h4><p>Tools like PMD and Spotbugs are not optimized for custom rule authorships. Most usage of these tools run built-in provided rules, or add in pre-made third party plugins. Take a look at what a custom rule for PMD might look like:</p><pre>&lt;![CDATA[<br> //AllocationExpression/ClassOrInterfaceType[<br>   @Image='DateTime' and (<br>       (count(..//Name[@Image='DateTimeZone.UTC'])&lt;=0)<br>       and<br>       (count(..//Name[@Image='DateTimeZone.forID'])&lt;=0)<br>    ) or (<br>       (<br>           (count(..//Name[@Image='DateTimeZone.UTC'])&gt;0)<br>             or<br>           (count(..//Name[@Image='DateTimeZone.forID'])&gt;0)<br>       ) and (../Arguments/ArgumentList and count(../Arguments/ArgumentList/Expression) = 1)<br>   )<br> ]<br>]]&gt;</pre><p>This rule ensures that DateTimes are not instantiated without an explicit zone. This is a raw string meant to be used within PMD’s xpath parser. There is no IDE guidance on crafting it. To test it, a whole separate PMD process needs to be wired up to interpret the rule and evaluate it against a source file. Let’s see how a similar rule would look with ArchUnit:</p><pre>ArchRuleDefinition.priority(Priority.MEDIUM)<br>.noClasses()<br>.should()<br>.callConstructorWhere(<br>    // constructor does not have a zone arguement<br>    target(doesNot(have(rawParameterTypes(DateTimeZone.class))))<br>   // constructor is for DateTime<br>        .and(targetOwner(assignableTo(DateTime.class)))<br>)</pre><p>This is type-safe Java code with a fluent API. It is also simple to unit test, as ArchUnit has a method to pass a rule object and class references to evaluate the rule against those classes.</p><h4>Class Relations</h4><p>Because ArchUnit processes the entire classpath with ASM, it retains a graph of the class data, allowing rules to easily traverse class relationships and call sites. This allows rules to have much more context about the code it is evaluating.</p><h3>Rules Libraries</h3><p>The first step was to build the ability to write ArchUnit rules which can be shared and published. In order to do this, we have the <a href="https://github.com/nebula-plugins/nebula-archrules-plugin?tab=readme-ov-file#authoring-rules">ArchRules Library Plugin</a>. This plugin adds an additional source set to your Gradle project called archRules. In this source set, you can create a class which implements the ArchRulesService interface. This interface has a single abstract method which returns a Map&lt;String, ArchRule&gt;. The keys of this map are the names of your rules, and the ArchRule is the rule you would like to define using the standard ArchUnit API. Here is an example:</p><pre>public class GuavaRules implements ArchRulesService {<br>  static final ArchRule OPTIONAL = ArchRuleDefinition.priority(Priority.MEDIUM)<br>        .noClasses()<br>        .should()<br>        .dependOnClassesThat()<br>        .haveFullyQualifiedName("com.google.common.base.Optional")<br>        .because("Java Optional is preferred over Guava Optional");<br><br>    @Override<br>    public Map&lt;String, ArchRule&gt; getRules() {<br>        Map&lt;String, ArchRule&gt; rules = new HashMap&lt;&gt;();<br>        rules.put("guava optional", OPTIONAL);<br>        return rules;<br>    }<br>}</pre><p>This code and its dependencies will not be bundled with your main code. It is bundled into a separate Jar with the arch-rules classifier. When publishing, your library will publish this jar as a separate variant with the usage attribute set to arch-rules. This means that in order for downstream projects to use these rules, they must use <a href="https://docs.gradle.org/current/userguide/publishing_gradle_module_metadata.html">Gradle Module Metadata</a> for dependency resolution. There are 2 flavors of rules Libraries: Standalone rules libraries, bundled rule libraries.</p><h4>Standalone Rule Libraries</h4><p>A Standalone Rule library contains no main code: only archRules. These are useful for defining rules for code you don’t own, such as Core Java APIs or OSS libraries. They are also useful for generic rules that can apply to any code, such as “don’t use code marked as @Deprecated”. We maintain a <a href="https://github.com/nebula-plugins/nebula-archrules">collection</a> of OSS Standalone rule libraries which anyone is free to use, and serve as examples of the types of rules you may want to write yourself. However, the real power of ArchRules is in “bundled rule libraries”.</p><h4>Bundled Rule Libraries</h4><p>A bundled rule library is a library with both main and archRules sources. The main source set will contain useful library code, whatever it may be. The archRules will contain rules specific to the usage of that library. For example, rules scoped to that library’s package, or referencing that library’s specific API. Whenever possible, we recommend writing rules in this bundled way. That is because the <a href="https://github.com/nebula-plugins/nebula-archrules-plugin?tab=readme-ov-file#running-rules">ArchRules Runner Plugin</a> will be able to automatically detect these rules and run them in only the source sets that use this library as a dependency. An example of this can be seen in our <a href="https://github.com/nebula-plugins/nebula-test/blob/main/src/archRules/java/com/netflix/nebula/test/archrules/NebulaTestArchRules.java">Nebula Test</a> library.</p><p>In any case, the library plugin will automatically generate a service loader registration entry for your ArchRulesService so that the runner can discover your rules.</p><h3>Running Rules</h3><p>The <a href="https://github.com/nebula-plugins/nebula-archrules-plugin?tab=readme-ov-file#running-rules">ArchRules Runner Plugin</a> allows rules to be evaluated against your code. Standalone rule libraries can be evaluated against all source sets by adding them to the archRules configuration in your build. For example:</p><pre>dependencies {<br>    archRules("your:rules:1.0.0")<br>}</pre><p>As mentioned before, bundled rules will be evaluated automatically. To do this, the runner plugin creates a separate configuration for each of your source sets. In each of these configurations, the archRules classpath is combined with the runtimeClasspath with the arch-rules variant selected. This configuration is the classpath used when the ServiceLoader discovers implementations of ArchRulesService. In the following example, we have a Project which uses a test helper library as a testImplementation dependency, and also adds a standalone rules library to the archRules configuration. The test runtime classpath will only contain the implementation jar for the helper library, but the arch rules runtime will contain the archrules jar for the bundled rules and standalone rules. This all happens automatically.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*pWrmMaPCSm3XRIsTyMEpgQ.png"><figcaption>Gradle configurations used by ArchRules</figcaption></figure><p>Once the rules classpath is determined, the runner plugin will create a Gradle work action to evaluate rules against that specific source set. This action runs with classpath isolation using the *archRuleRuntime configuration. Within this action, a ServiceLoader is used to discover rule definitions. The action ends by writing a binary serialization of rule violations to a file for reporting.</p><p>In a project running rules, you also have the ability to customize rule configurations using the archRules extension. For example, you can override a rule’s priority level:</p><pre>archRules {<br>    ruleClass("com.netflix.nebula.archrules.deprecation") {<br>        priority("HIGH")<br>    }<br>}</pre><p>Other <a href="https://github.com/nebula-plugins/nebula-archrules-plugin?tab=readme-ov-file#running-rules">customizations</a> include disabling running rules on certain source sets and configuring the failure threshold (i.e., high priority failures will cause the build to fail).</p><h3>Reporting</h3><p>The ArchRules runner plugin has two built-in reports: JSON and console. The json report will collect the output from all source sets within a project and create a single json file with all of the data. The console report also collects the output from all source sets within a project, but it prints to the console an easy to read report, for example:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*BJ2ONrMNHEFsBEks3WMPaQ.png"><figcaption>Console Report output</figcaption></figure><p>Note that failure details feature a detailed plain English description, along with a pointer to the exact line of code in violation.</p><p>For custom reporting, you can either use the JSON file, or create your own task that reads the binary files. Take a look at the source code for the ArchRules runner plugin’s report tasks for an example of how to do this.</p><h3>Case Study Solution</h3><p>Going back to our original problem, using ArchRules, we were able to deliver a platform for library authors to track the usage of their APIs. They write ArchRules to detect usage of the annotations, scoped to their library’s package, such as:</p><pre>ArchRuleDefinition.priority(Priority.MEDIUM)<br>    .noClasses().that(resideOutsideOfPackage(packageName + ".."))<br>    .should()<br>    .dependOnClassesThat(resideInAPackage(packageName + "..").and(are(deprecated())))<br>    .orShould().accessTargetWhere(targetOwner(resideInAPackage(packageName + ".."))<br>        .and(target(is(deprecated())).or(targetOwner(is(deprecated())))))<br>    .allowEmptyShould(true)<br>    .because("Deprecated APIs are subject to removal");</pre><p>NB: the deprecated() predicate comes from <a href="https://github.com/nebula-plugins/nebula-archrules/blob/main/archrules-common/src/main/java/com/netflix/nebula/archrules/common/CanBeAnnotated.java">nebula-archrules</a>.</p><p>Our internal Nebula standard Gradle wrapper and plugin suite automatically enable the ArchRules runner on every project, and provides a custom reporter which sends the report data to our Internal Developer Portal on every main-branch CI build. This way, library authors can easily see a report of all downstream consumers using their experimental, deprecated, or non-public APIs, giving them confidence to make “breaking” changes, knowing that it will not actually break downstream consumers. If their changes are currently blocked by downstream usage, they can easily see exactly which projects are reporting those usages.</p><h3>OSS Rule Libraries</h3><p>While the most powerful way to use ArchRules is for you to write your own rules, we have built some <a href="https://github.com/nebula-plugins/nebula-archrules">OSS rule libraries</a> that anyone is free to use, or reference as examples.</p><h4>Nullability</h4><p>These rules enforce proper nullability annotation in Java, for example, that every public class is marked with <a href="https://jspecify.dev/">JSpecify</a>’s @NullMarked. It is smart enough to exclude Kotlin code, as Kotlin has built-in nullability.</p><h4>Gradle Plugin Best Practices</h4><p><a href="https://docs.gradle.org/current/userguide/writing_plugins.html">Writing Gradle plugins</a> can be hard, especially since there are many APIs and patterns that should not be used anymore. These rules help enforce current best practices when writing Gradle plugins.</p><h4>Joda / Guava Rules</h4><p>These rule libraries discourage the use of Joda Time and Guava classes (respectively) as these have been superseded by java.time and standard library enhancements.</p><h4>Security Rules</h4><p>These rules help mitigate CVEs by detecting usage of known vulnerable APIs. Ideally, we keep dependencies up to date to mitigate CVEs. But sometimes that is not immediately feasible, and in those cases, a compile time check to ensure the specific vulnerable API is not used is often good enough.</p><h3>Conclusion</h3><p>We are now running 358 (and counting) rules across over 5,000 repositories detecting over nearly 1 million issues. About 1,000 of these issues are for “High” priority rules. Being able to run these rules on this scale allows us to quickly gain insight into our large fleet of microservices, and identify the areas carrying the most critical technical debt. This makes it easier to focus and prioritize our efforts.</p><p>Going forward, we will be exploring how to tie auto-remediation solutions into the ArchRules findings. ArchUnit currently provides very specific and detailed information about failures in reports, which makes a very strong input signal to an auto remediation tool. We will explore deterministic solutions such as <a href="https://docs.openrewrite.org/">OpenRewrite</a> and non-deterministic solutions such as LLMs. Pairing the easy rule authorship and deterministic results of ArchUnit with an auto-remediation tool that can correctly interpret the results to solve the issue at hand will be a very powerful combination.</p><p>We also will investigate how to get ArchRule failure information surfaced in the IDE as inspections.</p><p>If you have questions or feedback about Nebula ArchRules, reach out to us by posting in the #nebula channel on the <a href="http://gradle-community.slack.com/">Gradle Community</a> Slack.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=b4642c464c5a" width="1" height="1" alt=""><hr><p><a href="https://medium.com/netflix-techblog/scaling-archunit-with-nebula-archrules-b4642c464c5a">Scaling ArchUnit with Nebula ArchRules</a> was originally published in <a href="https://medium.com/netflix-techblog">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://medium.com/netflix-techblog/scaling-archunit-with-nebula-archrules-b4642c464c5a</link>
      <guid>https://medium.com/netflix-techblog/scaling-archunit-with-nebula-archrules-b4642c464c5a</guid>
      <pubDate>Fri, 08 May 2026 17:55:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Scaling ArchUnit with Nebula ArchRules]]></title>
      <description><![CDATA[<p>By <a href="https://github.com/wakingrufus">John Burns</a> and <a href="https://www.linkedin.com/in/emilyyuan03/">Emily Yuan</a></p><h3>Introduction</h3><p>At Netflix, we operate using a <a href="https://netflixtechblog.com/towards-true-continuous-integration-distributed-repositories-and-dependencies-2a2e3108c051">polyrepo</a> strategy with tens of thousands of Java repositories. This means that we need to have ways of sharing common build logic across these repositories. On the <a href="https://sites.google.com/netflix.com/javaplatformnetflix/jvm">JVM Ecosystem team</a> within Java Platform, we build tooling such as the <a href="https://github.com/nebula-plugins">Nebula suite of Gradle plugins</a> to provide standard ways to build projects, keep dependencies up-to-date, and publish artifacts reliably across the Java ecosystem. Our mission also entails providing build-time feedback to the developer when they deviate from the <a href="https://netflixtechblog.com/how-we-build-code-at-netflix-c5d9bd727f15">paved road</a>, or when their code base contains technical debt.</p><h3>Case Study</h3><p>After a Netflix incident relating to a library releasing a backwards-incompatible change, our team was asked to provide some tooling and practices to improve the Java library lifecycle management. This was not a simple case of a library making a reckless breaking change. The code removed had been deprecated for years. Library authors often struggle to know when it is safe to remove deprecated code, or refactor code that is not meant to be used by downstream applications. Fleet-wide migrations, such as upgrading major Spring Boot versions, also involve deprecated code removal. To help with this, we established a suite of API lifecycle annotations:</p><ul><li>@Deprecated from the Java standard library</li><li>@Public A custom annotation to use on APIs meant to be used downstream</li><li>@Experimental A custom annotation for new APIs which may not yet be stable</li><li>All other APIs are assumed to be “internal”</li></ul><p>Library authors can annotate their APIs with these annotations. However, how will they know which downstream projects are using their API incorrectly, based on these?</p><p>As we sought to improve the paved road for JVM-based libraries at Netflix, we needed a good way of identifying this kind of technical debt, not only for the benefit of the Java Platform-provided libraries, but any team delivering shared libraries to the organization. For this, we looked at ArchUnit.</p><p><a href="https://www.archunit.org/">ArchUnit</a> is a popular OSS library (3.5k stars, 84 contributors) used to enforce “architectural” code rules as part of a JUnit suite. It is used internally by Gradle, Spring, and is provided as part of the <a href="https://spring.io/projects/spring-modulith">Spring Modulith</a> platform. The rules engine, which is built directly on top of <a href="https://asm.ow2.io/">ASM</a>, can be used for a wide variety of use cases. It is powerful enough to be a general purpose static analysis tool with the following distinctive features:</p><p>1. Works cross-language (JVM), because it uses ASM/bytecode, not AST parsing.</p><p>2. Exposes a builder API pattern that makes it easy to write rules</p><p>3. Also has a lower level API ideal for writing more complex custom rules.</p><p>The limitation of ArchUnit is that it is designed to be used as part of a JUnit suite in a single repository. The Nebula ArchRules plugins give organizations the ability to share and apply rules across any number of repositories. Rules can be sourced from OSS libraries or private internal libraries. This makes the plugin generally useful for any JVM+Gradle engineering organization.</p><h3>Why ArchUnit?</h3><p>Before we go into how ArchRules works, it is good to understand why we would want to use ArchUnit in this way instead of other static analysis tools.</p><h4>AST vs Bytecode</h4><p>Some tools, such as PMD, process rules against an AST (abstract syntax tree). An AST is a structured representation of source code. This kind of tool will have rules that are syntax dependent. Rules that need to support multiple JVM languages, such as Kotlin or Scala, often need to be rewritten for each language. It also allows code which should be found to be hidden under syntactic sugar not anticipated by the rule author. ArchUnit uses <a href="https://asm.ow2.io/">ASM</a> to analyze actual compiled bytecode, which means it doesn’t matter how that code was produced. What is analyzed is the actual code that will be run.</p><h4>Rule Authorship</h4><p>Tools like PMD and Spotbugs are not optimized for custom rule authorships. Most usage of these tools run built-in provided rules, or add in pre-made third party plugins. Take a look at what a custom rule for PMD might look like:</p><pre>&lt;![CDATA[<br> //AllocationExpression/ClassOrInterfaceType[<br>   @Image='DateTime' and (<br>       (count(..//Name[@Image='DateTimeZone.UTC'])&lt;=0)<br>       and<br>       (count(..//Name[@Image='DateTimeZone.forID'])&lt;=0)<br>    ) or (<br>       (<br>           (count(..//Name[@Image='DateTimeZone.UTC'])&gt;0)<br>             or<br>           (count(..//Name[@Image='DateTimeZone.forID'])&gt;0)<br>       ) and (../Arguments/ArgumentList and count(../Arguments/ArgumentList/Expression) = 1)<br>   )<br> ]<br>]]&gt;</pre><p>This rule ensures that DateTimes are not instantiated without an explicit zone. This is a raw string meant to be used within PMD’s xpath parser. There is no IDE guidance on crafting it. To test it, a whole separate PMD process needs to be wired up to interpret the rule and evaluate it against a source file. Let’s see how a similar rule would look with ArchUnit:</p><pre>ArchRuleDefinition.priority(Priority.MEDIUM)<br>.noClasses()<br>.should()<br>.callConstructorWhere(<br>    // constructor does not have a zone arguement<br>    target(doesNot(have(rawParameterTypes(DateTimeZone.class))))<br>   // constructor is for DateTime<br>        .and(targetOwner(assignableTo(DateTime.class)))<br>)</pre><p>This is type-safe Java code with a fluent API. It is also simple to unit test, as ArchUnit has a method to pass a rule object and class references to evaluate the rule against those classes.</p><h4>Class Relations</h4><p>Because ArchUnit processes the entire classpath with ASM, it retains a graph of the class data, allowing rules to easily traverse class relationships and call sites. This allows rules to have much more context about the code it is evaluating.</p><h3>Rules Libraries</h3><p>The first step was to build the ability to write ArchUnit rules which can be shared and published. In order to do this, we have the <a href="https://github.com/nebula-plugins/nebula-archrules-plugin?tab=readme-ov-file#authoring-rules">ArchRules Library Plugin</a>. This plugin adds an additional source set to your Gradle project called archRules. In this source set, you can create a class which implements the ArchRulesService interface. This interface has a single abstract method which returns a Map&lt;String, ArchRule&gt;. The keys of this map are the names of your rules, and the ArchRule is the rule you would like to define using the standard ArchUnit API. Here is an example:</p><pre>public class GuavaRules implements ArchRulesService {<br>  static final ArchRule OPTIONAL = ArchRuleDefinition.priority(Priority.MEDIUM)<br>        .noClasses()<br>        .should()<br>        .dependOnClassesThat()<br>        .haveFullyQualifiedName("com.google.common.base.Optional")<br>        .because("Java Optional is preferred over Guava Optional");<br><br>    @Override<br>    public Map&lt;String, ArchRule&gt; getRules() {<br>        Map&lt;String, ArchRule&gt; rules = new HashMap&lt;&gt;();<br>        rules.put("guava optional", OPTIONAL);<br>        return rules;<br>    }<br>}</pre><p>This code and its dependencies will not be bundled with your main code. It is bundled into a separate Jar with the arch-rules classifier. When publishing, your library will publish this jar as a separate variant with the usage attribute set to arch-rules. This means that in order for downstream projects to use these rules, they must use <a href="https://docs.gradle.org/current/userguide/publishing_gradle_module_metadata.html">Gradle Module Metadata</a> for dependency resolution. There are 2 flavors of rules Libraries: Standalone rules libraries, bundled rule libraries.</p><h4>Standalone Rule Libraries</h4><p>A Standalone Rule library contains no main code: only archRules. These are useful for defining rules for code you don’t own, such as Core Java APIs or OSS libraries. They are also useful for generic rules that can apply to any code, such as “don’t use code marked as @Deprecated”. We maintain a <a href="https://github.com/nebula-plugins/nebula-archrules">collection</a> of OSS Standalone rule libraries which anyone is free to use, and serve as examples of the types of rules you may want to write yourself. However, the real power of ArchRules is in “bundled rule libraries”.</p><h4>Bundled Rule Libraries</h4><p>A bundled rule library is a library with both main and archRules sources. The main source set will contain useful library code, whatever it may be. The archRules will contain rules specific to the usage of that library. For example, rules scoped to that library’s package, or referencing that library’s specific API. Whenever possible, we recommend writing rules in this bundled way. That is because the <a href="https://github.com/nebula-plugins/nebula-archrules-plugin?tab=readme-ov-file#running-rules">ArchRules Runner Plugin</a> will be able to automatically detect these rules and run them in only the source sets that use this library as a dependency. An example of this can be seen in our <a href="https://github.com/nebula-plugins/nebula-test/blob/main/src/archRules/java/com/netflix/nebula/test/archrules/NebulaTestArchRules.java">Nebula Test</a> library.</p><p>In any case, the library plugin will automatically generate a service loader registration entry for your ArchRulesService so that the runner can discover your rules.</p><h3>Running Rules</h3><p>The <a href="https://github.com/nebula-plugins/nebula-archrules-plugin?tab=readme-ov-file#running-rules">ArchRules Runner Plugin</a> allows rules to be evaluated against your code. Standalone rule libraries can be evaluated against all source sets by adding them to the archRules configuration in your build. For example:</p><pre>dependencies {<br>    archRules("your:rules:1.0.0")<br>}</pre><p>As mentioned before, bundled rules will be evaluated automatically. To do this, the runner plugin creates a separate configuration for each of your source sets. In each of these configurations, the archRules classpath is combined with the runtimeClasspath with the arch-rules variant selected. This configuration is the classpath used when the ServiceLoader discovers implementations of ArchRulesService. In the following example, we have a Project which uses a test helper library as a testImplementation dependency, and also adds a standalone rules library to the archRules configuration. The test runtime classpath will only contain the implementation jar for the helper library, but the arch rules runtime will contain the archrules jar for the bundled rules and standalone rules. This all happens automatically.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*pWrmMaPCSm3XRIsTyMEpgQ.png"><figcaption>Gradle configurations used by ArchRules</figcaption></figure><p>Once the rules classpath is determined, the runner plugin will create a Gradle work action to evaluate rules against that specific source set. This action runs with classpath isolation using the *archRuleRuntime configuration. Within this action, a ServiceLoader is used to discover rule definitions. The action ends by writing a binary serialization of rule violations to a file for reporting.</p><p>In a project running rules, you also have the ability to customize rule configurations using the archRules extension. For example, you can override a rule’s priority level:</p><pre>archRules {<br>    ruleClass("com.netflix.nebula.archrules.deprecation") {<br>        priority("HIGH")<br>    }<br>}</pre><p>Other <a href="https://github.com/nebula-plugins/nebula-archrules-plugin?tab=readme-ov-file#running-rules">customizations</a> include disabling running rules on certain source sets and configuring the failure threshold (i.e., high priority failures will cause the build to fail).</p><h3>Reporting</h3><p>The ArchRules runner plugin has two built-in reports: JSON and console. The json report will collect the output from all source sets within a project and create a single json file with all of the data. The console report also collects the output from all source sets within a project, but it prints to the console an easy to read report, for example:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*BJ2ONrMNHEFsBEks3WMPaQ.png"><figcaption>Console Report output</figcaption></figure><p>Note that failure details feature a detailed plain English description, along with a pointer to the exact line of code in violation.</p><p>For custom reporting, you can either use the JSON file, or create your own task that reads the binary files. Take a look at the source code for the ArchRules runner plugin’s report tasks for an example of how to do this.</p><h3>Case Study Solution</h3><p>Going back to our original problem, using ArchRules, we were able to deliver a platform for library authors to track the usage of their APIs. They write ArchRules to detect usage of the annotations, scoped to their library’s package, such as:</p><pre>ArchRuleDefinition.priority(Priority.MEDIUM)<br>    .noClasses().that(resideOutsideOfPackage(packageName + ".."))<br>    .should()<br>    .dependOnClassesThat(resideInAPackage(packageName + "..").and(are(deprecated())))<br>    .orShould().accessTargetWhere(targetOwner(resideInAPackage(packageName + ".."))<br>        .and(target(is(deprecated())).or(targetOwner(is(deprecated())))))<br>    .allowEmptyShould(true)<br>    .because("Deprecated APIs are subject to removal");</pre><p>NB: the deprecated() predicate comes from <a href="https://github.com/nebula-plugins/nebula-archrules/blob/main/archrules-common/src/main/java/com/netflix/nebula/archrules/common/CanBeAnnotated.java">nebula-archrules</a>.</p><p>Our internal Nebula standard Gradle wrapper and plugin suite automatically enable the ArchRules runner on every project, and provides a custom reporter which sends the report data to our Internal Developer Portal on every main-branch CI build. This way, library authors can easily see a report of all downstream consumers using their experimental, deprecated, or non-public APIs, giving them confidence to make “breaking” changes, knowing that it will not actually break downstream consumers. If their changes are currently blocked by downstream usage, they can easily see exactly which projects are reporting those usages.</p><h3>OSS Rule Libraries</h3><p>While the most powerful way to use ArchRules is for you to write your own rules, we have built some <a href="https://github.com/nebula-plugins/nebula-archrules">OSS rule libraries</a> that anyone is free to use, or reference as examples.</p><h4>Nullability</h4><p>These rules enforce proper nullability annotation in Java, for example, that every public class is marked with <a href="https://jspecify.dev/">JSpecify</a>’s @NullMarked. It is smart enough to exclude Kotlin code, as Kotlin has built-in nullability.</p><h4>Gradle Plugin Best Practices</h4><p><a href="https://docs.gradle.org/current/userguide/writing_plugins.html">Writing Gradle plugins</a> can be hard, especially since there are many APIs and patterns that should not be used anymore. These rules help enforce current best practices when writing Gradle plugins.</p><h4>Joda / Guava Rules</h4><p>These rule libraries discourage the use of Joda Time and Guava classes (respectively) as these have been superseded by java.time and standard library enhancements.</p><h4>Security Rules</h4><p>These rules help mitigate CVEs by detecting usage of known vulnerable APIs. Ideally, we keep dependencies up to date to mitigate CVEs. But sometimes that is not immediately feasible, and in those cases, a compile time check to ensure the specific vulnerable API is not used is often good enough.</p><h3>Conclusion</h3><p>We are now running 358 (and counting) rules across over 5,000 repositories detecting over nearly 1 million issues. About 1,000 of these issues are for “High” priority rules. Being able to run these rules on this scale allows us to quickly gain insight into our large fleet of microservices, and identify the areas carrying the most critical technical debt. This makes it easier to focus and prioritize our efforts.</p><p>Going forward, we will be exploring how to tie auto-remediation solutions into the ArchRules findings. ArchUnit currently provides very specific and detailed information about failures in reports, which makes a very strong input signal to an auto remediation tool. We will explore deterministic solutions such as <a href="https://docs.openrewrite.org/">OpenRewrite</a> and non-deterministic solutions such as LLMs. Pairing the easy rule authorship and deterministic results of ArchUnit with an auto-remediation tool that can correctly interpret the results to solve the issue at hand will be a very powerful combination.</p><p>We also will investigate how to get ArchRule failure information surfaced in the IDE as inspections.</p><p>If you have questions or feedback about Nebula ArchRules, reach out to us by posting in the #nebula channel on the <a href="http://gradle-community.slack.com/">Gradle Community</a> Slack.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=b4642c464c5a" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/scaling-archunit-with-nebula-archrules-b4642c464c5a">Scaling ArchUnit with Nebula ArchRules</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/scaling-archunit-with-nebula-archrules-b4642c464c5a</link>
      <guid>https://netflixtechblog.com/scaling-archunit-with-nebula-archrules-b4642c464c5a</guid>
      <pubDate>Fri, 08 May 2026 17:55:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Democratizing Machine Learning at Netflix: Building the Model Lifecycle Graph]]></title>
      <description><![CDATA[<p><a href="https://www.linkedin.com/in/saishsali/">Saish Sali</a>, <a href="https://www.linkedin.com/in/nipunk/">Nipun Kumar</a>, <a href="https://www.linkedin.com/in/suraelamurugu/">Sura Elamurugu</a></p><h3>Introduction</h3><p>As Netflix has grown, machine learning continues to support our ability to deliver value to members and drive excellence across multiple areas of our business. When Netflix began investing in machine learning over a decade ago, it was primarily focused on a single domain: personalization. Scala was the industry standard, our ML teams were relatively small, and optimizing member engagement was our primary use case. Fast forward to today, and machine learning has become the backbone of Netflix’s business transformation. We now apply ML across various business domains, including:</p><ul><li><strong>Personalization</strong>: Optimizing engagement and helping members discover content they’ll love</li><li><strong>Studio</strong>: Pre and post-production workflows</li><li><strong>Payments</strong><em>: </em>Fraud detection, payment routing, and recurring billing optimization</li><li><strong>Ads</strong>: Our newest domain, requiring real-time decisioning and targeting</li></ul><p>… and a growing number of additional use cases across the company</p><p>Each domain operates with a different tech stack, different business metrics, and a distinct organizational structure. While this diversity is a testament to how machine learning has evolved to drive value across many verticals at Netflix, this growth introduces a new challenge: <strong>enabling cross-pollination of models and data across domains.</strong></p><h3>The Challenge: A Fragmented ML Landscape</h3><p>As our ML investments scaled across these domains, a critical problem emerged: the models produced largely became black boxes. Without any discovery infrastructure, ML practitioners couldn’t easily collaborate or share work across business verticals.</p><p>Consider a concrete example: <a href="https://netflixtechblog.com/mediafm-the-multimodal-ai-foundation-for-media-understanding-at-netflix-e8c28df82e2d">content embeddings</a>. Our Studio teams create sophisticated embeddings that identify scene boundaries, detect visual transitions, and understand content structure. These embeddings were originally built for production workflows.</p><p>But those same embeddings could be incredibly valuable elsewhere. Ads could hypothetically use content embeddings for context matching (ensuring advertisements align with the tone and content of what’s currently playing). Personalization could leverage them for episodic merchandising and recommendations (matching the topic or mood of an episode with a user’s preferred viewing preferences). Yet making this cross-pollination happen is extraordinarily difficult.</p><p>Why? Our ML tools exist in silos, each with its own backend services and user interface. The model registry is unaware of which A/B tests were using its models, and the pipeline orchestrator is unaware of downstream model dependencies. ML practitioners have to traverse multiple systems to answer basic questions about their work. Finding a model requires opening the model registry, understanding its lineage means switching to the pipeline orchestrator, and tracking which A/B tests use that model requires navigating to the experimentation platform. This fragmentation prevents practitioners from answering critical questions:</p><ul><li><strong>Discovery: </strong>What features exist? What data sources are available for generating features for a model?</li><li><strong>Lineage:</strong> Which pipeline is generating data for a specific model? What data sources feed those features?</li><li><strong>Impact:</strong> Which A/B tests are running this model? Which models will break if I change this feature? Who owns each piece of this chain?</li></ul><h3>The Hard Problem: Connecting everything</h3><p>The real challenge wasn’t just building a consolidated UI. We needed to connect the different pieces of infrastructure our ML practitioners were using to perform different parts of the ML lifecycle.</p><p>Our ML ecosystem generates metadata from dozens of sources:</p><ul><li>Pipeline orchestration systems emit execution details, stage dependencies, and data transformations</li><li>Deployed model registry tracks model versions, artifacts, staleness, and deployment history</li><li>Experimentation platform manages A/B tests and their configurations</li><li>Feature store catalog feature definitions and usage</li><li>AI Dataset platform tracks the creation, management, discovery, and loading of datasets.</li><li>Identity platform maintains user, team, and organization metadata</li></ul><p>Each system employs different formats, identifiers, and mental models. The hard technical problem we had to solve was: <strong>How do we collect this heterogeneous metadata, transform it into a unified entity model, and build a connected graph that enables true exploration and collaboration across business domains?</strong></p><h4>The Solution: Metadata Service and the Model Lifecycle Graph</h4><p>Our answer was the Metadata Service (MDS), which builds a Model Lifecycle Graph that indexes and connects ML-related entities across Netflix. MDS is optimized for real-time ingestion of ML metadata (e.g., models, features, pipelines, experiments, datasets) and to answer cross-domain questions such as “Which experiments are running this model?” or “Which models share these features?” It is the foundation that enables discovery, ingesting events from diverse sources, enriching them with context, and materializing relationships across entities.</p><p>Our vision: to make every ML asset at Netflix discoverable, understandable, and reusable by every ML practitioner, regardless of their team or domain.</p><h3>Core Abstractions: The Vocabulary of the System</h3><p>Before diving into the technical implementation, it’s helpful to understand the conceptual model that underpins MDS. This vocabulary enables consistent communication across teams and systems:</p><p><strong>Component:</strong> Any object that is uniquely addressable using an AI Platform’s (AIP) Uniform Resource Identifier (URI). An AIP URI follows the formataip://&lt;componentType&gt;/&lt;platformId&gt;/&lt;resourceId&gt;, ensuring global uniqueness. For example:</p><ul><li>Models: aip://model/registry/ranking-v5</li><li>Users: aip://user/identity/alice</li><li>Pipelines: aip://pipeline/orchestrator/weekly-training</li></ul><p><strong>Entity:</strong> A component within the ML ecosystem, characterized by additional properties such as name, description, creation date, and owners. Entities represent ML-specific assets, such as models, features, and pipelines.</p><p><strong>Entity Type:</strong> A group of entities that share the same data shape. A data shape is a set of property constraints that specify the attributes and relationships an entity must have.</p><p><strong>Domain:</strong> A functional grouping of related entity types that defines the abstract interface for a category of ML assets. For example, the Models domain defines what a Model and Model Instance look like, while the Pipelines domain defines Schedules, Requests, and Executions.</p><p><strong>Provider:</strong> A concrete implementation of a domain, backed by a specific source system. For example, the Models domain is currently backed by our internal model registry. This separation allows MDS to support multiple providers for the same domain. If a new model registry were introduced, it could be added as an additional provider without changing the domain interface.</p><p>We can summarize these concepts with a concrete example:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*RQuXyzwooTZcUZ5rCejOug.png"></figure><p>This URI-based addressing scheme is crucial as it allows any service to reference any ML asset with a single string, and MDS can resolve that reference back to rich, connected metadata.</p><h3><strong>From Events to Entities to Graph</strong></h3><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*dXe2XJMjTpZ6o5wwvVxGUA.png"></figure><p>The journey from raw system events to a queryable graph happens in stages. Let’s walk through each with a concrete example: connecting a model to its A/B tests through relationship inference.</p><h4>1 Event Ingestion</h4><p>MDS integrates with various source systems via Kafka and AWS SNS/SQS, consuming events in real-time. Source systems emit thin events that include an identifier and an event type.</p><p>Example event:</p><pre>{<br>  "event_type": "model_instance_created",<br>  "instance_id": "ranking-model-v5-20XX0101",<br>  ...<br>}</pre><p>This design keeps producers simple. Source systems only need to announce that a change occurred, without building complete payloads or understanding downstream requirements.</p><p>Each source system has dedicated event handlers in MDS:</p><ul><li><strong>Pipeline Orchestration</strong>: Ingests pipeline execution events, including node definitions, schedules, requests, and job attempts</li><li><strong>Model Registry</strong>: Captures model deployments, configurations, and version updates</li><li><strong>Feature Store</strong>: Tracks feature definitions and their versions</li><li><strong>Experimentation Platform</strong>: Monitors A/B test configurations and allocations</li><li><strong>Datasets:</strong> Tracks ML datasets and their versions</li><li><strong>Identity Platform</strong>: Maintains ownership and team membership information</li></ul><h4>2 Entity Enrichment</h4><p>MDS implements a hydration contract for each event type. When an event arrives, MDS:</p><ol><li>Validates the event schema</li><li>Calls the source system’s API to fetch the complete, current state</li><li>Transforms the response into a normalized entity</li></ol><p>This design has a crucial property: the order of events doesn’t matter. MDS always fetches the latest facts from the source of truth. This pattern decouples the event stream from state consistency. If the event bus drops a message or delivers it out of order, the next event corrects the state. The event stream becomes a notification of change rather than a log of changes.</p><p>This notification of change pattern has a few important tradeoffs. On the plus side, it keeps producers simple, makes us robust to out-of-order or dropped events, and ensures that MDS can always reconcile to the latest state by reading from the source of truth. The tradeoff is that we place additional read load on source systems during hydration and need to be deliberate about rate limiting, caching, and backoff in our enrichment workers so that we don’t overload them.</p><p>For our ranking model example, when the model_instance_created event arrives, MDS calls the Model Registry API: GET /api/v1/instances/ranking-model-v5-20XX0101</p><p>The registry responds with a full descriptor. Example response (key fields only):</p><pre>{<br>  "id": "ranking-model-v5-20XX0101",<br>  "pipeline_run_id": "train-weekly-ranking-20XX0101",<br>  "owner_emails": ["alice@netflix.com"],<br>  "labels": [{"key": "team", "value": "personalization"}],<br>  ...<br>}</pre><h4>3 Data Transformation and Normalization</h4><p>Raw events are heterogeneous and each source system has its own schema and semantics. MDS workers transform these events into a unified entity model with standardized fields.</p><p>Without normalization, downstream consumers would need to understand every source system’s schema. Normalization creates a consistent interface, allowing queries and relationships to work across all entity types. Here is an example.</p><p>Normalized MDS entity:</p><pre>{<br>  "id": "aip://model/registry/ranking-model-v5-20XX0101",<br>  "pipeline_run": "aip://pipeline-run/orchestrator/train-weekly-ranking-20XX0101",<br>  "entity_type": "ModelInstance",<br>  "owners": ["aip://user/identity/alice"],<br>  "tags": [{"tag": "team", "value": "personalization"}],<br>  ...<br>}</pre><p>The normalization process standardizes field names and formats. For example, platform-specific IDs become global AIP URIs, owner_emails becomes owners with resolved user URIs, and labels become tags. Foreign keys like pipeline_run_id are transformed into entity references. However, there’s still no reference to which A/B tests are using this model. The Model Registry doesn’t track experiments, and the Experimentation Platform doesn’t track which pipeline produced a given model. This is where knowledge enrichment becomes critical.</p><h4>4 Storage and Indexing</h4><p>Once normalized, entities are persisted to Datomic and immediately indexed in Elasticsearch. This happens synchronously within the event processing flow.</p><p><strong>Datomic for Caching and Relationships</strong><br>Normalized entities are first written to Datomic, which serves as both a local cache and a graph database.</p><p>Why Datomic? Datomic serves as both the system of record for MDS and the working dataset for enrichment processes. Its immutable fact model means we can continuously add relationships without losing the original entity state.</p><p><strong>What we store:</strong></p><ul><li>All entity attributes as facts</li><li>Entity references (foreign keys that may point to entities not yet fully resolved)</li><li>All relationships as reified edges (added by enrichment processes)</li><li>Entity lifecycle state (tracking which entities are fully enriched vs awaiting hydration)</li></ul><p><strong>This enables:</strong></p><ul><li><strong>Complex graph traversals:</strong> Navigate from a model to its features to their data sources in a single query</li><li><strong>Entity relationships:</strong> Join across multiple domains without N+1 query problems</li><li><strong>Flexible schema evolution:</strong> Easy to add new entity types and attributes as the catalog grows</li><li><strong>Progressive enrichment</strong>: Background jobs efficiently identify and process entities requiring additional hydration, enabling gradual graph completion without reprocessing fully enriched entities</li></ul><p>In practice, we use Datomic for relationship-heavy, navigational queries such as:</p><ul><li>Starting from this model instance, show me all upstream datasets and downstream experiments.</li><li>Given this feature, list all consuming models and their owning teams.</li></ul><p>These queries often span multiple hops in the graph and benefit from Datomic’s immutable fact model and efficient joins across entity relationships.</p><p><strong>Elasticsearch for Discovery</strong><br>Immediately after writing to Datomic, entities are indexed in Elasticsearch to power fast, full-text search across the catalog.</p><p><strong>What we index:</strong></p><ul><li>Primary fields: Entity name, description, entity type, owner names</li><li>Relationship metadata: Names of related entities (e.g., a model’s features, pipelines, A/B tests) stored in the related field</li><li>Tags: Domain-specific metadata stored as key-value pairs (e.g., <em>team::personalization, env::production, model.state::released</em>)</li></ul><p><strong>Index structure:</strong></p><ul><li>Single entities index: All entity types (models, features, pipelines, etc.) are indexed in one unified index, differentiated by the entityType field</li><li>Separate owners index: Dedicated index for users and groups to enable cross-entity owner searches</li><li>Relevance boosting: Exact name matches score higher than other relevant matches</li></ul><p><strong>This enables:</strong></p><ul><li>Multi-field text search across entity names, descriptions, tags, and related metadata</li><li>Relevance ranking with boosting (exact name matches score significantly higher)</li><li>Complex filtering by entity type, ownership, tags, and domain-specific attributes (stored as tags)</li><li>Fuzzy matching to handle typos and partial queries</li></ul><p>Elasticsearch powers the entry point into the system: users typically start with a free-text search in the AIP Portal (for a model name, a team, or a domain term), and then switch to graph navigation once they land on an entity page. Indexing happens in near real-time as part of the ingestion and enrichment workflows, so changes are usually visible in the Portal with a short delay that is acceptable for interactive use.</p><h4>5 Knowledge Enrichment and Graph Formation</h4><p>Once entity metadata is persisted in Datomic, scheduled background processes take over to discover and materialize relationships. These enrichment jobs run periodically, scanning for uncached or partially resolved entities (entities that exist only as references without full metadata).</p><p>The enrichment workflow:</p><ul><li><strong>Identify candidates:</strong> Find entities marked as uncached or with unresolved references</li><li><strong>Hydrate relationships:</strong> Query source-of-truth systems to fetch related entity details</li><li><strong>Materialize edges:</strong> Write discovered relationships back to Datomic</li><li><strong>Re-index:</strong> Trigger Elasticsearch indexing for updated entities</li><li><strong>Mark as enriched:</strong> Update entity status to prevent redundant processing</li></ul><p>This asynchronous approach allows MDS to handle the computational cost of graph formation without blocking real-time event ingestion. It also enables retry logic and gradual enrichment as new entities become available.</p><p>Because enrichment is asynchronous, newly discovered relationships may appear with a short delay after the underlying entities are created (typically minutes rather than seconds). We track when each entity was last enriched and surface this timestamp in the AIP Portal, so practitioners can reason about staleness and know when it’s safe to rely on a particular relationship for debugging or impact analysis.</p><p><strong>Why enrich?</strong> Source systems are purpose-built and don’t know about entities in other domains. Enrichment discovers and materializes cross-system relationships that enable powerful lineage and impact queries.</p><h4>Example: Connecting Models to A/B Tests</h4><p>When MDS processes a new model instance, background enrichment jobs discover relationships through multi-hop inference:</p><p><strong>Step 1: Direct link to pipeline</strong></p><p>The model references a pipeline_run_id. An enrichment job hydrates the pipeline and discovers its A/B test associations: GET /api/v1/pipeline-runs/train-weekly-ranking-20XX0101</p><p>Response:</p><pre>{<br>"run_id": "train-weekly-ranking-20XX0101", "pipeline":  "weekly-ranking-trainer",<br>"ab_test_cells": [<br>   {"test_id": "12345","cell_number": 2,"cell_name": "treatment_ranking_v5"}<br> ]<br> ...<br>}</pre><p><strong>Step 2: Discover A/B test context</strong><br>The enrichment job discovers the pipeline ran for A/B test cell #2 and queries the Experimentation Platform for test details: GET /api/v1/tests/12345</p><pre>{<br> "test_id": "12345",<br> "name": "Ranking Model v5 vs v4",<br> "status": "ACTIVE",<br> "cells": [{"cell_number": 1, "name": "control_ranking_v4"}],<br> ...<br>}</pre><p><strong>Step 3: Infer transitive relationships</strong><br>The enrichment job now has the complete chain:</p><ul><li>Model Instance was produced by Pipeline Run</li><li>Pipeline Run was executed for A/B Test Cell #2</li><li>The A/B Test Cell #2 belongs to A/B Test “Ranking Model v5 vs v4”</li><li>Model Instance now gets associated with this A/B Test</li></ul><p>The job writes the inferred relationship back to Datomic and triggers re-indexing, and materializes these edges in the graph. MDS doesn’t just store what it’s told; it derives new knowledge by <em>walking</em> the graph in the background.</p><p><strong>Why this matters:</strong> Without MDS, answering “Which A/B tests are using this model?” requires:</p><ol><li>Looking up the model in the Model Registry</li><li>Finding which pipeline produced it</li><li>Checking the Pipeline Orchestrator for A/B test tags</li><li>Querying the Experimentation Platform for test details</li></ol><p>With the model lifecycle graph, it’s a single query:</p><pre>query {<br>  model(id: "aip://model/registry/ranking-model-v5-20XX0101") {<br>    name<br>    owners { name }<br>    currentInstance {<br>      version<br>      pipeline {<br>        name<br>        owners { name }<br>      }<br>      features {<br>        edges {<br>          node {<br>            name<br>            data { edges { node { name } } }<br>          }<br>        }<br>      }<br>      associatedAbTests {<br>        name<br>        cells { number name }<br>      }<br>    }<br>  }<br>}</pre><p>The reverse query also works: “What models are being tested in experiment 12345?”</p><h3>Enabling Exploration, Not Just Search</h3><figure><img alt="" src="https://cdn-images-1.medium.com/max/682/1*j8moOl19CHOfIDRvfnPk5A.png"></figure><p>With the Model Lifecycle Graph in place, we shift from entity search to entity exploration. Discovery isn’t just about finding a model; It’s about traversing relationships:</p><ul><li>Start with a model, explore its features</li><li>From features, navigate to the core data driving them</li><li>From the data, trace back to the pipelines generating it</li><li>From pipelines, see which teams own and depend on them</li><li>From experiments, understand which models are being tested</li></ul><p>For example, imagine an engineer investigating a degraded engagement metric for a personalization model. They might:</p><ol><li>Start with the model instance powering the affected recommendations in the AIP Portal.</li><li>Inspect the model’s features and follow a suspicious feature to its upstream dataset.</li><li>From the dataset page, see that its pipeline recently had failed runs and identify the owning team.</li><li>Confirm which A/B tests are currently running this model instance to understand which members and surfaces are impacted.</li></ol><p>Before MDS and the Model Lifecycle Graph, this required manual checks across multiple tools (model registry, pipeline orchestrator, experiment platform). Now it’s a contiguous journey in a single interface.</p><p>This graph-based exploration answers questions that were previously impossible:</p><ul><li>Lineage queries: What is the complete lineage of this model, from training data to production experiments?</li><li>Impact analysis: Which models will be affected if I change this feature?</li><li>Usage discovery: Which A/B tests are using this model?</li><li>Dependency mapping: What data sources does my pipeline transitively depend on?</li><li>Deprecation planning: Which entities are no longer being used and can be retired?</li></ul><p>Every entity has deep context: its creation time, ownership, update history, and most importantly, its relationships to other entities.</p><p>The Model Lifecycle Graph is surfaced to practitioners through the AIP Portal, a unified interface that provides full-text search across all entity types, detailed entity pages with navigable relationships, and personalized views for teams and individuals.</p><p>A typical interaction in the AIP Portal looks like:</p><ul><li><strong>Search:</strong> Type a model, feature, dataset, or team name into the single search box backed by Elasticsearch.</li><li><strong>Inspect:</strong> Land on an entity page that shows key metadata (description, owners, domains, tags) alongside a relationships panel.</li><li><strong>Explore:</strong> Click through to related entities (upstream datasets, downstream experiments, and sibling model versions) to navigate the Model Lifecycle Graph without leaving the portal.</li></ul><p>When new entity types are introduced into MDS, the portal automatically provides baseline search, entity pages, and relationship navigation, and we can then layer on domain-specific visualizations (such as model deployment history or dataset version timelines) over time.</p><h3>The Road Ahead: Open Challenges</h3><p>Building the ML lifecycle graph is an ongoing journey. Significant challenges remain, and these represent the future opportunities for us:</p><ul><li><strong>Tool Proliferation:</strong> As new ML tools emerge, we need robust integration patterns that scale. How do we design plugin architectures that make adding new sources seamless? If we don’t keep up with new tools, practitioners will be forced back into fragmented views, and the Model Lifecycle Graph will lose coverage and trust.</li><li><strong>Domain-Specific Visualizations:</strong> Different entity types require distinct visualization experiences. Model pages should display deployment history, A/B test associations, and performance metrics. Feature pages should highlight data lineage and consuming models. Pipeline pages must show execution history, dependencies, and schedules. Dataset pages require versioning timelines and downstream consumers. How do we design a flexible UI framework that allows each entity type to have its own tailored experience while maintaining consistent navigation and interaction patterns across the portal? Without rich, domain-specific experiences, the portal risks becoming a generic catalog rather than a tool that ML practitioners rely on in their daily workflows.</li><li><strong>Metadata Quality:</strong> Today, MDS ensures data consistency through source-of-truth hydration and schema validation at ingestion. Background enrichment jobs continuously infer relationships and materialize entities from source systems. However, challenges remain in ensuring completeness and timeliness at scale. When source systems fail to emit events, when ownership information becomes stale, or when entities lack descriptions and contextual metadata, the graph’s utility degrades. How do we build automated validation and enrichment systems to detect metadata anomalies, suggest missing relationships, and maintain quality benchmarks across millions of entities? Poor or stale metadata erodes practitioner trust: if the graph is incomplete or incorrect, teams will revert to ad hoc knowledge and one-off integrations rather than using MDS as their source of truth.</li><li><strong>Advanced Relationship Inference:</strong> Beyond explicit relationships declared in source systems, how do we infer implicit connections? Can we detect that two models serve similar purposes based on shared features? Can we recommend features based on usage patterns from similar pipelines? We are in the early stages of exploring these ideas. Done well, they would turn MDS from a passive catalog into an active recommendation engine for ML assets, accelerating reuse and reducing duplicate work across domains.</li></ul><h3>Acknowledgments</h3><p>This work represents the collective effort of stunning colleagues across the AI Platform organization: <a href="https://www.linkedin.com/in/emma-carney-6a700b17a/">Emma Carney</a>, <a href="https://www.linkedin.com/in/megan-ren-7b78a81a8/">Megan Ren</a>, <a href="https://www.linkedin.com/in/nadeem-ahmad-80000983/">Nadeem Ahmad</a>, <a href="https://www.linkedin.com/in/poleniuk/">Pat Olenik</a>, <a href="https://www.linkedin.com/in/prateekagarwal17/">Prateek Agarwal</a>, <a href="https://www.linkedin.com/in/tikhakobyan/">Tigran Hakobyan</a>, <a href="https://www.linkedin.com/in/yinglao-liu-6b48b6126/">Yinglao Liu</a></p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=5cc6d5828bb1" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/democratizing-machine-learning-at-netflix-building-the-model-lifecycle-graph-5cc6d5828bb1">Democratizing Machine Learning at Netflix: Building the Model Lifecycle Graph</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/democratizing-machine-learning-at-netflix-building-the-model-lifecycle-graph-5cc6d5828bb1</link>
      <guid>https://netflixtechblog.com/democratizing-machine-learning-at-netflix-building-the-model-lifecycle-graph-5cc6d5828bb1</guid>
      <pubDate>Mon, 04 May 2026 18:01:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Democratizing Machine Learning at Netflix: Building the Model Lifecycle Graph]]></title>
      <description><![CDATA[<p><a href="https://www.linkedin.com/in/saishsali/">Saish Sali</a>, <a href="https://www.linkedin.com/in/nipunk/">Nipun Kumar</a>, <a href="https://www.linkedin.com/in/suraelamurugu/">Sura Elamurugu</a></p><h3>Introduction</h3><p>As Netflix has grown, machine learning continues to support our ability to deliver value to members and drive excellence across multiple areas of our business. When Netflix began investing in machine learning over a decade ago, it was primarily focused on a single domain: personalization. Scala was the industry standard, our ML teams were relatively small, and optimizing member engagement was our primary use case. Fast forward to today, and machine learning has become the backbone of Netflix’s business transformation. We now apply ML across various business domains, including:</p><ul><li><strong>Personalization</strong>: Optimizing engagement and helping members discover content they’ll love</li><li><strong>Studio</strong>: Pre and post-production workflows</li><li><strong>Payments</strong><em>: </em>Fraud detection, payment routing, and recurring billing optimization</li><li><strong>Ads</strong>: Our newest domain, requiring real-time decisioning and targeting</li></ul><p>… and a growing number of additional use cases across the company</p><p>Each domain operates with a different tech stack, different business metrics, and a distinct organizational structure. While this diversity is a testament to how machine learning has evolved to drive value across many verticals at Netflix, this growth introduces a new challenge: <strong>enabling cross-pollination of models and data across domains.</strong></p><h3>The Challenge: A Fragmented ML Landscape</h3><p>As our ML investments scaled across these domains, a critical problem emerged: the models produced largely became black boxes. Without any discovery infrastructure, ML practitioners couldn’t easily collaborate or share work across business verticals.</p><p>Consider a concrete example: <a href="https://netflixtechblog.com/mediafm-the-multimodal-ai-foundation-for-media-understanding-at-netflix-e8c28df82e2d">content embeddings</a>. Our Studio teams create sophisticated embeddings that identify scene boundaries, detect visual transitions, and understand content structure. These embeddings were originally built for production workflows.</p><p>But those same embeddings could be incredibly valuable elsewhere. Ads could hypothetically use content embeddings for context matching (ensuring advertisements align with the tone and content of what’s currently playing). Personalization could leverage them for episodic merchandising and recommendations (matching the topic or mood of an episode with a user’s preferred viewing preferences). Yet making this cross-pollination happen is extraordinarily difficult.</p><p>Why? Our ML tools exist in silos, each with its own backend services and user interface. The model registry is unaware of which A/B tests were using its models, and the pipeline orchestrator is unaware of downstream model dependencies. ML practitioners have to traverse multiple systems to answer basic questions about their work. Finding a model requires opening the model registry, understanding its lineage means switching to the pipeline orchestrator, and tracking which A/B tests use that model requires navigating to the experimentation platform. This fragmentation prevents practitioners from answering critical questions:</p><ul><li><strong>Discovery: </strong>What features exist? What data sources are available for generating features for a model?</li><li><strong>Lineage:</strong> Which pipeline is generating data for a specific model? What data sources feed those features?</li><li><strong>Impact:</strong> Which A/B tests are running this model? Which models will break if I change this feature? Who owns each piece of this chain?</li></ul><h3>The Hard Problem: Connecting everything</h3><p>The real challenge wasn’t just building a consolidated UI. We needed to connect the different pieces of infrastructure our ML practitioners were using to perform different parts of the ML lifecycle.</p><p>Our ML ecosystem generates metadata from dozens of sources:</p><ul><li>Pipeline orchestration systems emit execution details, stage dependencies, and data transformations</li><li>Deployed model registry tracks model versions, artifacts, staleness, and deployment history</li><li>Experimentation platform manages A/B tests and their configurations</li><li>Feature store catalog feature definitions and usage</li><li>AI Dataset platform tracks the creation, management, discovery, and loading of datasets.</li><li>Identity platform maintains user, team, and organization metadata</li></ul><p>Each system employs different formats, identifiers, and mental models. The hard technical problem we had to solve was: <strong>How do we collect this heterogeneous metadata, transform it into a unified entity model, and build a connected graph that enables true exploration and collaboration across business domains?</strong></p><h4>The Solution: Metadata Service and the Model Lifecycle Graph</h4><p>Our answer was the Metadata Service (MDS), which builds a Model Lifecycle Graph that indexes and connects ML-related entities across Netflix. MDS is optimized for real-time ingestion of ML metadata (e.g., models, features, pipelines, experiments, datasets) and to answer cross-domain questions such as “Which experiments are running this model?” or “Which models share these features?” It is the foundation that enables discovery, ingesting events from diverse sources, enriching them with context, and materializing relationships across entities.</p><p>Our vision: to make every ML asset at Netflix discoverable, understandable, and reusable by every ML practitioner, regardless of their team or domain.</p><h3>Core Abstractions: The Vocabulary of the System</h3><p>Before diving into the technical implementation, it’s helpful to understand the conceptual model that underpins MDS. This vocabulary enables consistent communication across teams and systems:</p><p><strong>Component:</strong> Any object that is uniquely addressable using an AI Platform’s (AIP) Uniform Resource Identifier (URI). An AIP URI follows the formataip://&lt;componentType&gt;/&lt;platformId&gt;/&lt;resourceId&gt;, ensuring global uniqueness. For example:</p><ul><li>Models: aip://model/registry/ranking-v5</li><li>Users: aip://user/identity/alice</li><li>Pipelines: aip://pipeline/orchestrator/weekly-training</li></ul><p><strong>Entity:</strong> A component within the ML ecosystem, characterized by additional properties such as name, description, creation date, and owners. Entities represent ML-specific assets, such as models, features, and pipelines.</p><p><strong>Entity Type:</strong> A group of entities that share the same data shape. A data shape is a set of property constraints that specify the attributes and relationships an entity must have.</p><p><strong>Domain:</strong> A functional grouping of related entity types that defines the abstract interface for a category of ML assets. For example, the Models domain defines what a Model and Model Instance look like, while the Pipelines domain defines Schedules, Requests, and Executions.</p><p><strong>Provider:</strong> A concrete implementation of a domain, backed by a specific source system. For example, the Models domain is currently backed by our internal model registry. This separation allows MDS to support multiple providers for the same domain. If a new model registry were introduced, it could be added as an additional provider without changing the domain interface.</p><p>We can summarize these concepts with a concrete example:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*RQuXyzwooTZcUZ5rCejOug.png"></figure><p>This URI-based addressing scheme is crucial as it allows any service to reference any ML asset with a single string, and MDS can resolve that reference back to rich, connected metadata.</p><h3><strong>From Events to Entities to Graph</strong></h3><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*dXe2XJMjTpZ6o5wwvVxGUA.png"></figure><p>The journey from raw system events to a queryable graph happens in stages. Let’s walk through each with a concrete example: connecting a model to its A/B tests through relationship inference.</p><h4>1 Event Ingestion</h4><p>MDS integrates with various source systems via Kafka and AWS SNS/SQS, consuming events in real-time. Source systems emit thin events that include an identifier and an event type.</p><p>Example event:</p><pre>{<br>  "event_type": "model_instance_created",<br>  "instance_id": "ranking-model-v5-20XX0101",<br>  ...<br>}</pre><p>This design keeps producers simple. Source systems only need to announce that a change occurred, without building complete payloads or understanding downstream requirements.</p><p>Each source system has dedicated event handlers in MDS:</p><ul><li><strong>Pipeline Orchestration</strong>: Ingests pipeline execution events, including node definitions, schedules, requests, and job attempts</li><li><strong>Model Registry</strong>: Captures model deployments, configurations, and version updates</li><li><strong>Feature Store</strong>: Tracks feature definitions and their versions</li><li><strong>Experimentation Platform</strong>: Monitors A/B test configurations and allocations</li><li><strong>Datasets:</strong> Tracks ML datasets and their versions</li><li><strong>Identity Platform</strong>: Maintains ownership and team membership information</li></ul><h4>2 Entity Enrichment</h4><p>MDS implements a hydration contract for each event type. When an event arrives, MDS:</p><ol><li>Validates the event schema</li><li>Calls the source system’s API to fetch the complete, current state</li><li>Transforms the response into a normalized entity</li></ol><p>This design has a crucial property: the order of events doesn’t matter. MDS always fetches the latest facts from the source of truth. This pattern decouples the event stream from state consistency. If the event bus drops a message or delivers it out of order, the next event corrects the state. The event stream becomes a notification of change rather than a log of changes.</p><p>This notification of change pattern has a few important tradeoffs. On the plus side, it keeps producers simple, makes us robust to out-of-order or dropped events, and ensures that MDS can always reconcile to the latest state by reading from the source of truth. The tradeoff is that we place additional read load on source systems during hydration and need to be deliberate about rate limiting, caching, and backoff in our enrichment workers so that we don’t overload them.</p><p>For our ranking model example, when the model_instance_created event arrives, MDS calls the Model Registry API: GET /api/v1/instances/ranking-model-v5-20XX0101</p><p>The registry responds with a full descriptor. Example response (key fields only):</p><pre>{<br>  "id": "ranking-model-v5-20XX0101",<br>  "pipeline_run_id": "train-weekly-ranking-20XX0101",<br>  "owner_emails": ["alice@netflix.com"],<br>  "labels": [{"key": "team", "value": "personalization"}],<br>  ...<br>}</pre><h4>3 Data Transformation and Normalization</h4><p>Raw events are heterogeneous and each source system has its own schema and semantics. MDS workers transform these events into a unified entity model with standardized fields.</p><p>Without normalization, downstream consumers would need to understand every source system’s schema. Normalization creates a consistent interface, allowing queries and relationships to work across all entity types. Here is an example.</p><p>Normalized MDS entity:</p><pre>{<br>  "id": "aip://model/registry/ranking-model-v5-20XX0101",<br>  "pipeline_run": "aip://pipeline-run/orchestrator/train-weekly-ranking-20XX0101",<br>  "entity_type": "ModelInstance",<br>  "owners": ["aip://user/identity/alice"],<br>  "tags": [{"tag": "team", "value": "personalization"}],<br>  ...<br>}</pre><p>The normalization process standardizes field names and formats. For example, platform-specific IDs become global AIP URIs, owner_emails becomes owners with resolved user URIs, and labels become tags. Foreign keys like pipeline_run_id are transformed into entity references. However, there’s still no reference to which A/B tests are using this model. The Model Registry doesn’t track experiments, and the Experimentation Platform doesn’t track which pipeline produced a given model. This is where knowledge enrichment becomes critical.</p><h4>4 Storage and Indexing</h4><p>Once normalized, entities are persisted to Datomic and immediately indexed in Elasticsearch. This happens synchronously within the event processing flow.</p><p><strong>Datomic for Caching and Relationships</strong><br>Normalized entities are first written to Datomic, which serves as both a local cache and a graph database.</p><p>Why Datomic? Datomic serves as both the system of record for MDS and the working dataset for enrichment processes. Its immutable fact model means we can continuously add relationships without losing the original entity state.</p><p><strong>What we store:</strong></p><ul><li>All entity attributes as facts</li><li>Entity references (foreign keys that may point to entities not yet fully resolved)</li><li>All relationships as reified edges (added by enrichment processes)</li><li>Entity lifecycle state (tracking which entities are fully enriched vs awaiting hydration)</li></ul><p><strong>This enables:</strong></p><ul><li><strong>Complex graph traversals:</strong> Navigate from a model to its features to their data sources in a single query</li><li><strong>Entity relationships:</strong> Join across multiple domains without N+1 query problems</li><li><strong>Flexible schema evolution:</strong> Easy to add new entity types and attributes as the catalog grows</li><li><strong>Progressive enrichment</strong>: Background jobs efficiently identify and process entities requiring additional hydration, enabling gradual graph completion without reprocessing fully enriched entities</li></ul><p>In practice, we use Datomic for relationship-heavy, navigational queries such as:</p><ul><li>Starting from this model instance, show me all upstream datasets and downstream experiments.</li><li>Given this feature, list all consuming models and their owning teams.</li></ul><p>These queries often span multiple hops in the graph and benefit from Datomic’s immutable fact model and efficient joins across entity relationships.</p><p><strong>Elasticsearch for Discovery</strong><br>Immediately after writing to Datomic, entities are indexed in Elasticsearch to power fast, full-text search across the catalog.</p><p><strong>What we index:</strong></p><ul><li>Primary fields: Entity name, description, entity type, owner names</li><li>Relationship metadata: Names of related entities (e.g., a model’s features, pipelines, A/B tests) stored in the related field</li><li>Tags: Domain-specific metadata stored as key-value pairs (e.g., <em>team::personalization, env::production, model.state::released</em>)</li></ul><p><strong>Index structure:</strong></p><ul><li>Single entities index: All entity types (models, features, pipelines, etc.) are indexed in one unified index, differentiated by the entityType field</li><li>Separate owners index: Dedicated index for users and groups to enable cross-entity owner searches</li><li>Relevance boosting: Exact name matches score higher than other relevant matches</li></ul><p><strong>This enables:</strong></p><ul><li>Multi-field text search across entity names, descriptions, tags, and related metadata</li><li>Relevance ranking with boosting (exact name matches score significantly higher)</li><li>Complex filtering by entity type, ownership, tags, and domain-specific attributes (stored as tags)</li><li>Fuzzy matching to handle typos and partial queries</li></ul><p>Elasticsearch powers the entry point into the system: users typically start with a free-text search in the AIP Portal (for a model name, a team, or a domain term), and then switch to graph navigation once they land on an entity page. Indexing happens in near real-time as part of the ingestion and enrichment workflows, so changes are usually visible in the Portal with a short delay that is acceptable for interactive use.</p><h4>5 Knowledge Enrichment and Graph Formation</h4><p>Once entity metadata is persisted in Datomic, scheduled background processes take over to discover and materialize relationships. These enrichment jobs run periodically, scanning for uncached or partially resolved entities (entities that exist only as references without full metadata).</p><p>The enrichment workflow:</p><ul><li><strong>Identify candidates:</strong> Find entities marked as uncached or with unresolved references</li><li><strong>Hydrate relationships:</strong> Query source-of-truth systems to fetch related entity details</li><li><strong>Materialize edges:</strong> Write discovered relationships back to Datomic</li><li><strong>Re-index:</strong> Trigger Elasticsearch indexing for updated entities</li><li><strong>Mark as enriched:</strong> Update entity status to prevent redundant processing</li></ul><p>This asynchronous approach allows MDS to handle the computational cost of graph formation without blocking real-time event ingestion. It also enables retry logic and gradual enrichment as new entities become available.</p><p>Because enrichment is asynchronous, newly discovered relationships may appear with a short delay after the underlying entities are created (typically minutes rather than seconds). We track when each entity was last enriched and surface this timestamp in the AIP Portal, so practitioners can reason about staleness and know when it’s safe to rely on a particular relationship for debugging or impact analysis.</p><p><strong>Why enrich?</strong> Source systems are purpose-built and don’t know about entities in other domains. Enrichment discovers and materializes cross-system relationships that enable powerful lineage and impact queries.</p><h4>Example: Connecting Models to A/B Tests</h4><p>When MDS processes a new model instance, background enrichment jobs discover relationships through multi-hop inference:</p><p><strong>Step 1: Direct link to pipeline</strong></p><p>The model references a pipeline_run_id. An enrichment job hydrates the pipeline and discovers its A/B test associations: GET /api/v1/pipeline-runs/train-weekly-ranking-20XX0101</p><p>Response:</p><pre>{<br>"run_id": "train-weekly-ranking-20XX0101", "pipeline":  "weekly-ranking-trainer",<br>"ab_test_cells": [<br>   {"test_id": "12345","cell_number": 2,"cell_name": "treatment_ranking_v5"}<br> ]<br> ...<br>}</pre><p><strong>Step 2: Discover A/B test context</strong><br>The enrichment job discovers the pipeline ran for A/B test cell #2 and queries the Experimentation Platform for test details: GET /api/v1/tests/12345</p><pre>{<br> "test_id": "12345",<br> "name": "Ranking Model v5 vs v4",<br> "status": "ACTIVE",<br> "cells": [{"cell_number": 1, "name": "control_ranking_v4"}],<br> ...<br>}</pre><p><strong>Step 3: Infer transitive relationships</strong><br>The enrichment job now has the complete chain:</p><ul><li>Model Instance was produced by Pipeline Run</li><li>Pipeline Run was executed for A/B Test Cell #2</li><li>The A/B Test Cell #2 belongs to A/B Test “Ranking Model v5 vs v4”</li><li>Model Instance now gets associated with this A/B Test</li></ul><p>The job writes the inferred relationship back to Datomic and triggers re-indexing, and materializes these edges in the graph. MDS doesn’t just store what it’s told; it derives new knowledge by <em>walking</em> the graph in the background.</p><p><strong>Why this matters:</strong> Without MDS, answering “Which A/B tests are using this model?” requires:</p><ol><li>Looking up the model in the Model Registry</li><li>Finding which pipeline produced it</li><li>Checking the Pipeline Orchestrator for A/B test tags</li><li>Querying the Experimentation Platform for test details</li></ol><p>With the model lifecycle graph, it’s a single query:</p><pre>query {<br>  model(id: "aip://model/registry/ranking-model-v5-20XX0101") {<br>    name<br>    owners { name }<br>    currentInstance {<br>      version<br>      pipeline {<br>        name<br>        owners { name }<br>      }<br>      features {<br>        edges {<br>          node {<br>            name<br>            data { edges { node { name } } }<br>          }<br>        }<br>      }<br>      associatedAbTests {<br>        name<br>        cells { number name }<br>      }<br>    }<br>  }<br>}</pre><p>The reverse query also works: “What models are being tested in experiment 12345?”</p><h3>Enabling Exploration, Not Just Search</h3><figure><img alt="" src="https://cdn-images-1.medium.com/max/682/1*j8moOl19CHOfIDRvfnPk5A.png"></figure><p>With the Model Lifecycle Graph in place, we shift from entity search to entity exploration. Discovery isn’t just about finding a model; It’s about traversing relationships:</p><ul><li>Start with a model, explore its features</li><li>From features, navigate to the core data driving them</li><li>From the data, trace back to the pipelines generating it</li><li>From pipelines, see which teams own and depend on them</li><li>From experiments, understand which models are being tested</li></ul><p>For example, imagine an engineer investigating a degraded engagement metric for a personalization model. They might:</p><ol><li>Start with the model instance powering the affected recommendations in the AIP Portal.</li><li>Inspect the model’s features and follow a suspicious feature to its upstream dataset.</li><li>From the dataset page, see that its pipeline recently had failed runs and identify the owning team.</li><li>Confirm which A/B tests are currently running this model instance to understand which members and surfaces are impacted.</li></ol><p>Before MDS and the Model Lifecycle Graph, this required manual checks across multiple tools (model registry, pipeline orchestrator, experiment platform). Now it’s a contiguous journey in a single interface.</p><p>This graph-based exploration answers questions that were previously impossible:</p><ul><li>Lineage queries: What is the complete lineage of this model, from training data to production experiments?</li><li>Impact analysis: Which models will be affected if I change this feature?</li><li>Usage discovery: Which A/B tests are using this model?</li><li>Dependency mapping: What data sources does my pipeline transitively depend on?</li><li>Deprecation planning: Which entities are no longer being used and can be retired?</li></ul><p>Every entity has deep context: its creation time, ownership, update history, and most importantly, its relationships to other entities.</p><p>The Model Lifecycle Graph is surfaced to practitioners through the AIP Portal, a unified interface that provides full-text search across all entity types, detailed entity pages with navigable relationships, and personalized views for teams and individuals.</p><p>A typical interaction in the AIP Portal looks like:</p><ul><li><strong>Search:</strong> Type a model, feature, dataset, or team name into the single search box backed by Elasticsearch.</li><li><strong>Inspect:</strong> Land on an entity page that shows key metadata (description, owners, domains, tags) alongside a relationships panel.</li><li><strong>Explore:</strong> Click through to related entities (upstream datasets, downstream experiments, and sibling model versions) to navigate the Model Lifecycle Graph without leaving the portal.</li></ul><p>When new entity types are introduced into MDS, the portal automatically provides baseline search, entity pages, and relationship navigation, and we can then layer on domain-specific visualizations (such as model deployment history or dataset version timelines) over time.</p><h3>The Road Ahead: Open Challenges</h3><p>Building the ML lifecycle graph is an ongoing journey. Significant challenges remain, and these represent the future opportunities for us:</p><ul><li><strong>Tool Proliferation:</strong> As new ML tools emerge, we need robust integration patterns that scale. How do we design plugin architectures that make adding new sources seamless? If we don’t keep up with new tools, practitioners will be forced back into fragmented views, and the Model Lifecycle Graph will lose coverage and trust.</li><li><strong>Domain-Specific Visualizations:</strong> Different entity types require distinct visualization experiences. Model pages should display deployment history, A/B test associations, and performance metrics. Feature pages should highlight data lineage and consuming models. Pipeline pages must show execution history, dependencies, and schedules. Dataset pages require versioning timelines and downstream consumers. How do we design a flexible UI framework that allows each entity type to have its own tailored experience while maintaining consistent navigation and interaction patterns across the portal? Without rich, domain-specific experiences, the portal risks becoming a generic catalog rather than a tool that ML practitioners rely on in their daily workflows.</li><li><strong>Metadata Quality:</strong> Today, MDS ensures data consistency through source-of-truth hydration and schema validation at ingestion. Background enrichment jobs continuously infer relationships and materialize entities from source systems. However, challenges remain in ensuring completeness and timeliness at scale. When source systems fail to emit events, when ownership information becomes stale, or when entities lack descriptions and contextual metadata, the graph’s utility degrades. How do we build automated validation and enrichment systems to detect metadata anomalies, suggest missing relationships, and maintain quality benchmarks across millions of entities? Poor or stale metadata erodes practitioner trust: if the graph is incomplete or incorrect, teams will revert to ad hoc knowledge and one-off integrations rather than using MDS as their source of truth.</li><li><strong>Advanced Relationship Inference:</strong> Beyond explicit relationships declared in source systems, how do we infer implicit connections? Can we detect that two models serve similar purposes based on shared features? Can we recommend features based on usage patterns from similar pipelines? We are in the early stages of exploring these ideas. Done well, they would turn MDS from a passive catalog into an active recommendation engine for ML assets, accelerating reuse and reducing duplicate work across domains.</li></ul><h3>Acknowledgments</h3><p>This work represents the collective effort of stunning colleagues across the AI Platform organization: <a href="https://www.linkedin.com/in/emma-carney-6a700b17a/">Emma Carney</a>, <a href="https://www.linkedin.com/in/megan-ren-7b78a81a8/">Megan Ren</a>, <a href="https://www.linkedin.com/in/nadeem-ahmad-80000983/">Nadeem Ahmad</a>, <a href="https://www.linkedin.com/in/poleniuk/">Pat Oleniuk</a>, <a href="https://www.linkedin.com/in/prateekagarwal17/">Prateek Agarwal</a>, <a href="https://www.linkedin.com/in/tikhakobyan/">Tigran Hakobyan</a>, <a href="https://www.linkedin.com/in/yinglao-liu-6b48b6126/">Yinglao Liu</a></p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=5cc6d5828bb1" width="1" height="1" alt=""><hr><p><a href="https://medium.com/netflix-techblog/democratizing-machine-learning-at-netflix-building-the-model-lifecycle-graph-5cc6d5828bb1">Democratizing Machine Learning at Netflix: Building the Model Lifecycle Graph</a> was originally published in <a href="https://medium.com/netflix-techblog">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://medium.com/netflix-techblog/democratizing-machine-learning-at-netflix-building-the-model-lifecycle-graph-5cc6d5828bb1</link>
      <guid>https://medium.com/netflix-techblog/democratizing-machine-learning-at-netflix-building-the-model-lifecycle-graph-5cc6d5828bb1</guid>
      <pubDate>Mon, 04 May 2026 18:01:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[State of Routing in Model Serving]]></title>
      <description><![CDATA[<p>By <a href="https://www.linkedin.com/in/nipunk/">Nipun Kumar</a>, <a href="https://www.linkedin.com/in/rajatsshah/">Rajat Shah</a>, <a href="https://www.linkedin.com/in/peterchng/">Peter Chng</a></p><h3>Introduction</h3><p><em>This is the first blog post in a multi-part series that shares technical insights into how our ML model serving infrastructure powers several personalized experiences at scale across various domains (e.g., title recommendations, commerce). In this introductory blog post, we will dive into our domain-independent API abstraction and its traffic routing capabilities that the central ML model serving platform exposes to several domain-specific microservices for model inference. This singular API, or entry point, into the ML model serving platform has significantly increased the speed of innovation for iterating on newer versions of existing ML experiences, as well as enabling completely new product experiences with ML.</em></p><p>Machine Learning use cases powering member experiences on Netflix require rapid iteration and evolution in response to new learnings. The success of our ML model serving infrastructure largely depends on enabling researchers to rapidly experiment with new hypotheses and safely, at scale, release their models into production. Equally important is enabling multiple microservices at Netflix to seamlessly get model inference without exposing the complexities of ML model inference. To achieve this in a uniform and scalable manner, we created a centralized ML serving platform. As of 2025, the platform serves hundreds of model types and versions, netting 1 million requests per second. In this post, we’ll zoom in on a core challenge of any large-scale ML serving system: How to route traffic to the right model instance, on the right cluster shard, for the right user and use case, while preserving a simple abstraction for both client services and model researchers.</p><h3>Background</h3><h3>Models at Netflix</h3><p>To properly frame our discussion, let’s first clarify the distinction between model <em>serving</em> and model <em>inference</em>. At Netflix, the definition of an ML model has historically been somewhat unique. While model <em>inference</em> typically focuses only on an infer(features) -&gt; score capability, models at Netflix act as self-contained workflows that transform inputs to outputs. A “model” encapsulates pre- and post-processing, feature computation logic, and an optional ML-trained component, all packaged in a standard format suitable for use across multiple contexts. We refer to the end-to-end execution of this workflow as model <em>serving</em>. This distinction matters because our routing and API abstractions operate at the level of workflows, not just individual scoring functions.</p><p>A few <em>simplified</em> examples of model serving use cases:</p><p><strong>Use case</strong>: Personalized Continue Watching row on Netflix Homepage</p><ul><li>Input: UserId, Country, Device ID</li><li>Output: Ranked List of movies and shows (aka title): [titleId1, titleId2, titleId3,…]</li></ul><p><strong>Use case</strong>: Payment Fraud Detection</p><ul><li>Input: UserId, Country, Payment Transaction details</li><li>Output: Probability of the transaction being fraudulent</li></ul><p>A typical flow of this serving workflow is depicted below:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/804/1*V0IaBLQSyADbjdnlhRDJdA.png"></figure><p>To achieve this higher level of abstraction, the model definition contains a list of facts (raw, unprocessed data or observations built as states in different business workflows) that it needs to compute features, and it relies on the model serving platform to supply these facts at serving time by calling several other microservices. Likewise, during offline training, <a href="https://netflixtechblog.com/evolution-of-ml-fact-store-5941d3231762">Netflix’s ML fact store</a> provides snapshots for bulk access to facilitate feature computation.</p><p>The important takeaway from this model definition is that the calling services only need to provide standard request context (such as userId, country, device), and the relevant domain context (such as titles to rank, or payment transaction for fraud detection), and the model can itself compute features and perform inference as part of the execution flow. This common set of request contexts across domains enables them to share a standard API abstraction and standardizes how various client microservices can uniformly integrate with the serving app. Furthermore, clients are shielded from the model selection and execution, allowing the model architecture and data inputs to evolve with minimal client coordination.</p><p>This post focuses on showcasing the technical details to support this design paradigm. We’ll first describe how we implemented this abstraction with Switchboard, a centralized routing service, and then discuss the operational challenges we encountered at scale and how they led us to the Lightbulb architecture.</p><h3>ML Model Serving Platform Principles</h3><p>We envisioned a central model serving platform for all of Netflix’s member-facing ML Model serving needs. This ambitious effort required principled thinking to provide the right level of abstraction for both the researchers and client applications. The following ideas, which are relevant to the topic of this blog post, ensured that the platform acts as an enabler of rapid ML innovation and limits the exposure of ML model iterations to the client apps:</p><ul><li><strong>Model innovation independent of client apps: </strong>There should be only a one-time integration effort by the calling app with the ML serving platform for a new use case. After that, almost all model iterations, including intermediate model A/B experiments, should be mostly opaque to the calling apps. This implies that the platform should handle tasks such as model selection based on a user’s A/B allocation, fetching additional data needed by experimental models, logging for further training or observability, and more. This also benefits the ML researcher, as they only need to coordinate with one platform for model innovation.</li><li><strong>Decouple clients from model sharding: </strong>Models are distributed across multiple serving compute cluster shards, each with its own Virtual IP (VIP) Address. Various factors, such as traffic patterns, SLAs, model architecture, and CPU/Memory availability, affect model-to-cluster mapping, and changes to this mapping result in changes to the VIP address at which a model is reachable. The serving platform should make clients agnostic to such frequent VIP address changes while ensuring high availability.</li><li><strong>Flexible traffic routing rules: </strong>Support flexible mechanisms to introduce new traffic routing rules. This includes supporting traffic routing based on A/B experiments, providing a knob to slowly shift traffic to new models and VIP addresses, and allowing client overrides.</li></ul><h3>Introducing Switchboard</h3><p>Standard out-of-the-box API Gateway solutions (such as AWS API Gateway, a standalone Service Mesh proxy) did not meet all our requirements. In particular, we needed first-class integration with Netflix’s experimentation platform, the ability to expose gRPC endpoints to clients, and the ability to use rich domain-specific context for routing customizations, which generic proxies were not designed to handle. Furthermore, the platform required customizations to model-specific lifecycle stages (shadow mode, canaries, rollbacks) to enable safe rollouts and migrations.</p><p>Hence, we embarked on building a custom service that serves as a flexible proxy layer for all traffic, handling over 1 million requests per second while maintaining high availability and reliability. We named it Switchboard.</p><p>Switchboard serves as the central entry point for the system, <strong>acting as a mandatory interface </strong>for all clients to access the appropriate model based on their context. Its role is to perform context-aware routing and to apply any configured context enrichment to the model inputs.</p><p>Here is a visual representation of the request flow from different clients to different serving clusters:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/935/1*HEV6B_6F5ci3dyoKXyq55A.png"></figure><h3>Objective Abstraction</h3><p>To support this system design, we introduce the concept of an “Objective”. It’s an Enumeration defined by the serving platform that every request into the system must provide. It has three key purposes:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/703/1*V6bhyHhYQ4W5baVzvygW9g.png"></figure><p>In short, an <strong>Objective</strong> is the serving platform’s name for a specific business use case (e.g., ContinueWatchingRanking), which decouples clients from concrete models and guides the platform’s routing and model selection decisions.</p><h3>Key Capabilities of Switchboard</h3><p>To summarize, these are the key capabilities of Switchboard:</p><ol><li><strong>Common Client Abstraction: </strong>Switchboard provides a single point of contact for all our clients’ model needs. When clients wish to consume additional models for new ML applications addressing the same business need, there is no new service dependency to introduce or new clients to manage to make requests to the models. From an ML Ops perspective, this also gives us knobs to control client rate limits across model versions and manage central concurrency limits to deal with bad clients.</li><li><strong>Context-Aware Routing:</strong> Switchboard can route a request based on a rich set of contextual features, such as the user’s current device, locale, ranking surface type (e.g., home page vs. search results), or the current A/B test a user is in.</li><li><strong>Dynamic Traffic Splitting:</strong> It enables real-time traffic splitting for canary deployments and experimentation. This allows engineers to safely roll out a new model version to a small, controlled percentage of users before a full launch.</li><li><strong>Model Versioning and Lifecycle Management:</strong> Switchboard inherently manages concurrent request traffic to multiple versions of the same model. This is crucial for:</li></ol><ul><li><strong>Shadow Mode Testing:</strong> Routing production traffic to a new model version without affecting the user experience, enabling performance comparisons.</li><li><strong>Instant Rollback:</strong> Immediate switching of traffic away from a problematic new model version back to a stable one.</li></ul><p>But is this the whole story? Not quite. Introducing this routing layer adds complexity to our model deployment cycles. In addition, we need a mechanism to collect the context-based routing information from the researchers when they choose to deploy model variants.</p><h3>The Glue — Switchboard Rules</h3><p>Given that Objectives serve as the contract between clients and the serving platform, we needed a way for researchers to attach model variants, experiments, and traffic splits to those Objectives without changing client code. This is where Switchboard Rules comes in.</p><p>The primary UX for model researchers to define models associated with an objective in a flexible manner is a JavaScript configuration, which we call <em>Switchboard Rules</em>. It’s used to produce a set of rules (typically a JSON file) that primarily dictate the following things to the serving platform:</p><ol><li>The default model to use for a given Objective</li><li>A/B experiments to configure for a set of Objectives and the corresponding models to load for those experiments</li><li>Customizations to gradually shift traffic to a new model</li></ol><p>Here is an example of an A/B test rule in the context of the Continue Watching row:</p><pre>/**<br>Configuration rule written by a Model Researcher to add an A/B experiment in the Model Serving system.<br>Cell 1: Uses the default, currently productized model<br>Cell 2 and Cell 3: Use different experimental (candidate) models<br>**/<br><br>function defineAB12345Rule() {<br>    const abTestId = 12345;<br><br>    const objectives = Objectives.ContinueWatchingRanking;<br>    const abTestCellToModel = {<br>        1: {name: "netflix-continue-watching-model-default"},<br>        2: {name: "netflix-continue-watching-model-cell-2"},<br>        3: {name: "netflix-continue-watching-model-cell-3"}<br>    };<br><br>    return {<br>        cellToModel: abTestCellToModel,<br>        abTestId: abTestId,<br>        targetObjectives: [objectives],<br>        modelInputType: constants.TITLE_INPUT_TYPE,<br>        modelType: 'SCORER'<br>    };<br>}</pre><p>These rules are consumed by both the Switchboard and the Model Serving clusters. Given these rules, the serving platform components can take various actions, some detailed below:</p><p><strong>Control Plane Flow</strong>:</p><ol><li><strong>Assignment:</strong> Produce model-to-cluster shard assignment.</li><li><strong>Validation:</strong> Load all specified models into the Serving Cluster Shard and validate model dependencies to ensure successful execution.</li><li><strong>Mapping:</strong> Provide the model-to-shard VIP address mapping to Switchboard.</li></ol><p><strong>Data Plane Flow</strong>:</p><ol><li><strong>Allocation:</strong> If the request is for Objective=ContinueWatchingRanking, query the <a href="https://netflixtechblog.com/its-all-a-bout-testing-the-netflix-experimentation-platform-4e1ca458c15">Experimentation Platform</a> for the userId’s cell allocation.</li><li><strong>Model Selection:</strong> Use the allocation and A/B test rule to select the appropriate model.</li><li><strong>Request Routing:</strong> Route the request to the serving cluster shard with the selected model and context.</li><li><strong>Model Execution (on the serving host):</strong> Run the model workflow steps and return the response.</li></ol><p>A key highlight of this setup is the decoupling of the experimentation config from the serving platform code. This includes having an independent release cycle for the rules, separate from the code deployments. <a href="https://netflixtechblog.com/how-netflix-microservices-tackle-dataset-pub-sub-4a068adcc9a">Netflix’s Gutenberg</a> system provides an excellent ecosystem that enables a flexible pub-sub architecture, facilitating proper versioning, dynamic loading, easy rollbacks, and more. Both Switchboard and the Serving Cluster Host subscribe to the same Switchboard Rules configuration.</p><p>To prevent race conditions and ensure proper sync of the dynamic Switchboard Rules configuration, the following flow is considered:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/944/1*iSNuD8ZuSC9E2zN-wZdH-A.png"></figure><h3>Evolving Challenges</h3><p>Switchboard solved the primary problem of improving model iteration and innovation velocity, and provided an excellent ML serving abstraction to over 30 service clients. However, as the system scale increased, a few challenges and problems with this design became apparent:</p><ul><li><strong>Single point of failure: </strong>The presence of Switchboard in the critical request path clearly highlights the risks of shutting down access to all serving hosts in extreme cases, such as unintentional bugs or noisy neighbors sending excessive traffic.</li><li><em>Why this matters: Switchboard became a shared dependency whose failure would degrade or disable multiple ML-powered experiences at Netflix.</em></li><li><strong>Added latency due to additional network hop:</strong> Switchboard in the request path adds between 10–20ms of latency due to serialization-deserialization operations, depending on payload size. Additionally, it further exposes a request to tail latency amplification.</li><li><em>Why this matters: The added latency is unacceptable for some latency-sensitive clients, resulting in end-user impact due to service timeouts.</em></li><li><strong>Reduced Client flexibility</strong>: Switchboard obscures visibility into client request origins from the serving clusters. Consequently, distinguishing data logged for real vs artificial traffic, which is essential for model training, is difficult and requires ongoing customization and increased MLOps overhead.</li><li><em>Why this matters: It makes it harder to do tenant separation and test traffic isolation.</em></li></ul><h3>What Next? — Lightbulb</h3><p>The aforementioned challenges of operating Switchboard at scale forced us to rethink the core implementation while retaining its key features. Our goal was not to throw away Switchboard’s design, but to refactor where and how its responsibilities were executed, keeping the benefits while reducing risk and latency. Particularly:</p><ul><li><em>Common Client Abstraction</em></li><li><em>Decouple clients from model sharding</em></li><li><em>Flexible traffic routing rules</em></li><li><em>Lightweight system client</em></li><li><em>Single place to define model and experimentation config</em></li></ul><p>However, we did want to address some of the previous design choices to move forward with:</p><ul><li><strong>Remove the routing service from the direct request path: </strong>Having a single service in the active request path introduces another failure mode and limits fallback flexibility. While routing rules change infrequently, maintaining consistency comes at the cost of increased availability risks.</li><li><strong>Separate model inputs from the request metadata</strong>: In certain cases, the request payload could be quite large. Needing to deserialize and then re-serialize the payload as it flowed through Switchboard to make a routing decision was a significant contributor to latency and increased serving costs.</li><li><strong>Provide better isolation for the routing layer: </strong>Consolidating multiple use cases (tenants) into a single routing cluster poses two main challenges. First, error propagation posed a risk, as a surge of problematic requests from one tenant could cascade errors back to Switchboard, potentially impacting other users. Second, the cluster had to accommodate diverse latency requirements because the requests from different use cases varied significantly in complexity.</li></ul><p>This required some changes in our setup flow: While it largely remained unchanged, however, we created separate components for Routing and Model Selection (Lightbulb):</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/897/1*wtVpe5xJMNEENitvkd1FPQ.png"></figure><p>We now take the rules for an Objective and break them into distinct sets of configuration:</p><ul><li><strong>Model Serving Configuration</strong>: This allows us to determine which model should be used at request time, along with the required metadata</li><li><strong>Routing Rules</strong>: Given a model we want to serve at request time, this tells us which VIP the request should be routed to.</li></ul><p>The Data Plane changes also reflect this separation, as we now rely on <a href="https://github.com/envoyproxy/envoy">Envoy</a> to take care of the routing details:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/912/1*jbW4NlcnKucGKBi_vNjXbw.png"></figure><p>Envoy is <a href="https://netflixtechblog.com/zero-configuration-service-mesh-with-on-demand-cluster-discovery-ac6483b52a51">already used</a> for all egress communication between apps at Netflix, and it can route requests to different clusters (VIPs) based on the configurable Routing Rules published from our control plane. However, it lacks the information needed to make routing decisions and the ability to enrich the request body with additional serving parameters required for A/B testing model variants. We introduced Lightbulb to cover this gap:</p><ul><li>Lightbulb consumes the minimal request context, which contains use-case information, and provides the metadata mapping required for routing at the Envoy layer.</li><li>Lightbulb resolves the request context to determine a routingKey configuration along with the <strong>ObjectiveConfig</strong> — this is where we place the model id along with other request-specific configurations required for model execution. This is done to separate the config resolution associated with the request from the placement and routing information needed to reach it on the inference cluster.</li><li>While the routingKey is added to the headers for Envoy proxy to consume, the client adds the ObjectiveConfig parameters to the request itself. This is done to avoid bloating the request headers while passing additional parameters for the model to process the request appropriately.</li><li>The routing of the actual request is performed by the Envoy proxy, which has the metadata to map the routingKey to the actual cluster VIP running the model. Because the routingKey is in a header, this determination can be made with minimal overhead.</li></ul><p>These changes retain the advantages of Switchboard, such as a single integration point, abstraction of model id from use case, context-aware routing, while addressing the challenges we observed over time.</p><h3>Conclusion</h3><p>The evolution from Switchboard to Lightbulb marks a significant architectural refinement in our ML model serving infrastructure. While Switchboard provided the initial abstraction layer critical for rapid innovation, its latency and single-point-of-failure risk posed scaling hurdles. The subsequent adoption of Lightbulb, a decoupled service focused solely on routing metadata, and its integration with Envoy successfully resolved these challenges. This sophisticated new architecture preserves the key benefits — seamless client integration and flexible experimentation — while ensuring reliable, efficient, and scalable delivery of personalized member experiences, positioning us well for future ML growth.</p><p>In future posts in this series, we’ll dive deeper into other aspects of our ML serving platform, including inference and feature fetching, and how they interact with the routing architecture described here.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=16e22fe18741" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/state-of-routing-in-model-serving-16e22fe18741">State of Routing in Model Serving</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/state-of-routing-in-model-serving-16e22fe18741</link>
      <guid>https://netflixtechblog.com/state-of-routing-in-model-serving-16e22fe18741</guid>
      <pubDate>Fri, 01 May 2026 23:03:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[State of Routing in Model Serving]]></title>
      <description><![CDATA[<p>By <a href="https://www.linkedin.com/in/nipunk/">Nipun Kumar</a>, <a href="https://www.linkedin.com/in/rajatsshah/">Rajat Shah</a>, <a href="https://www.linkedin.com/in/peterchng/">Peter Chng</a></p><h3>Introduction</h3><p><em>This is the first blog post in a multi-part series that shares technical insights into how our ML model serving infrastructure powers several personalized experiences at scale across various domains (e.g., title recommendations, commerce). In this introductory blog post, we will dive into our domain-independent API abstraction and its traffic routing capabilities that the central ML model serving platform exposes to several domain-specific microservices for model inference. This singular API, or entry point, into the ML model serving platform has significantly increased the speed of innovation for iterating on newer versions of existing ML experiences, as well as enabling completely new product experiences with ML.</em></p><p>Machine Learning use cases powering member experiences on Netflix require rapid iteration and evolution in response to new learnings. The success of our ML model serving infrastructure largely depends on enabling researchers to rapidly experiment with new hypotheses and safely, at scale, release their models into production. Equally important is enabling multiple microservices at Netflix to seamlessly get model inference without exposing the complexities of ML model inference. To achieve this in a uniform and scalable manner, we created a centralized ML serving platform. As of 2025, the platform serves hundreds of model types and versions, netting 1 million requests per second. In this post, we’ll zoom in on a core challenge of any large-scale ML serving system: How to route traffic to the right model instance, on the right cluster shard, for the right user and use case, while preserving a simple abstraction for both client services and model researchers.</p><h3>Background</h3><h3>Models at Netflix</h3><p>To properly frame our discussion, let’s first clarify the distinction between model <em>serving</em> and model <em>inference</em>. At Netflix, the definition of an ML model has historically been somewhat unique. While model <em>inference</em> typically focuses only on an infer(features) -&gt; score capability, models at Netflix act as self-contained workflows that transform inputs to outputs. A “model” encapsulates pre- and post-processing, feature computation logic, and an optional ML-trained component, all packaged in a standard format suitable for use across multiple contexts. We refer to the end-to-end execution of this workflow as model <em>serving</em>. This distinction matters because our routing and API abstractions operate at the level of workflows, not just individual scoring functions.</p><p>A few <em>simplified</em> examples of model serving use cases:</p><p><strong>Use case</strong>: Personalized Continue Watching row on Netflix Homepage</p><ul><li>Input: UserId, Country, Device ID</li><li>Output: Ranked List of movies and shows (aka title): [titleId1, titleId2, titleId3,…]</li></ul><p><strong>Use case</strong>: Payment Fraud Detection</p><ul><li>Input: UserId, Country, Payment Transaction details</li><li>Output: Probability of the transaction being fraudulent</li></ul><p>A typical flow of this serving workflow is depicted below:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/804/1*V0IaBLQSyADbjdnlhRDJdA.png"></figure><p>To achieve this higher level of abstraction, the model definition contains a list of facts (raw, unprocessed data or observations built as states in different business workflows) that it needs to compute features, and it relies on the model serving platform to supply these facts at serving time by calling several other microservices. Likewise, during offline training, <a href="https://netflixtechblog.com/evolution-of-ml-fact-store-5941d3231762">Netflix’s ML fact store</a> provides snapshots for bulk access to facilitate feature computation.</p><p>The important takeaway from this model definition is that the calling services only need to provide standard request context (such as userId, country, device), and the relevant domain context (such as titles to rank, or payment transaction for fraud detection), and the model can itself compute features and perform inference as part of the execution flow. This common set of request contexts across domains enables them to share a standard API abstraction and standardizes how various client microservices can uniformly integrate with the serving app. Furthermore, clients are shielded from the model selection and execution, allowing the model architecture and data inputs to evolve with minimal client coordination.</p><p>This post focuses on showcasing the technical details to support this design paradigm. We’ll first describe how we implemented this abstraction with Switchboard, a centralized routing service, and then discuss the operational challenges we encountered at scale and how they led us to the Lightbulb architecture.</p><h3>ML Model Serving Platform Principles</h3><p>We envisioned a central model serving platform for all of Netflix’s member-facing ML Model serving needs. This ambitious effort required principled thinking to provide the right level of abstraction for both the researchers and client applications. The following ideas, which are relevant to the topic of this blog post, ensured that the platform acts as an enabler of rapid ML innovation and limits the exposure of ML model iterations to the client apps:</p><ul><li><strong>Model innovation independent of client apps: </strong>There should be only a one-time integration effort by the calling app with the ML serving platform for a new use case. After that, almost all model iterations, including intermediate model A/B experiments, should be mostly opaque to the calling apps. This implies that the platform should handle tasks such as model selection based on a user’s A/B allocation, fetching additional data needed by experimental models, logging for further training or observability, and more. This also benefits the ML researcher, as they only need to coordinate with one platform for model innovation.</li><li><strong>Decouple clients from model sharding: </strong>Models are distributed across multiple serving compute cluster shards, each with its own Virtual IP (VIP) Address. Various factors, such as traffic patterns, SLAs, model architecture, and CPU/Memory availability, affect model-to-cluster mapping, and changes to this mapping result in changes to the VIP address at which a model is reachable. The serving platform should make clients agnostic to such frequent VIP address changes while ensuring high availability.</li><li><strong>Flexible traffic routing rules: </strong>Support flexible mechanisms to introduce new traffic routing rules. This includes supporting traffic routing based on A/B experiments, providing a knob to slowly shift traffic to new models and VIP addresses, and allowing client overrides.</li></ul><h3>Introducing Switchboard</h3><p>Standard out-of-the-box API Gateway solutions (such as AWS API Gateway, a standalone Service Mesh proxy) did not meet all our requirements. In particular, we needed first-class integration with Netflix’s experimentation platform, the ability to expose gRPC endpoints to clients, and the ability to use rich domain-specific context for routing customizations, which generic proxies were not designed to handle. Furthermore, the platform required customizations to model-specific lifecycle stages (shadow mode, canaries, rollbacks) to enable safe rollouts and migrations.</p><p>Hence, we embarked on building a custom service that serves as a flexible proxy layer for all traffic, handling over 1 million requests per second while maintaining high availability and reliability. We named it Switchboard.</p><p>Switchboard serves as the central entry point for the system, <strong>acting as a mandatory interface </strong>for all clients to access the appropriate model based on their context. Its role is to perform context-aware routing and to apply any configured context enrichment to the model inputs.</p><p>Here is a visual representation of the request flow from different clients to different serving clusters:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/935/1*HEV6B_6F5ci3dyoKXyq55A.png"></figure><h3>Objective Abstraction</h3><p>To support this system design, we introduce the concept of an “Objective”. It’s an Enumeration defined by the serving platform that every request into the system must provide. It has three key purposes:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/703/1*V6bhyHhYQ4W5baVzvygW9g.png"></figure><p>In short, an <strong>Objective</strong> is the serving platform’s name for a specific business use case (e.g., ContinueWatchingRanking), which decouples clients from concrete models and guides the platform’s routing and model selection decisions.</p><h3>Key Capabilities of Switchboard</h3><p>To summarize, these are the key capabilities of Switchboard:</p><ol><li><strong>Common Client Abstraction: </strong>Switchboard provides a single point of contact for all our clients’ model needs. When clients wish to consume additional models for new ML applications addressing the same business need, there is no new service dependency to introduce or new clients to manage to make requests to the models. From an ML Ops perspective, this also gives us knobs to control client rate limits across model versions and manage central concurrency limits to deal with bad clients.</li><li><strong>Context-Aware Routing:</strong> Switchboard can route a request based on a rich set of contextual features, such as the user’s current device, locale, ranking surface type (e.g., home page vs. search results), or the current A/B test a user is in.</li><li><strong>Dynamic Traffic Splitting:</strong> It enables real-time traffic splitting for canary deployments and experimentation. This allows engineers to safely roll out a new model version to a small, controlled percentage of users before a full launch.</li><li><strong>Model Versioning and Lifecycle Management:</strong> Switchboard inherently manages concurrent request traffic to multiple versions of the same model. This is crucial for:</li></ol><ul><li><strong>Shadow Mode Testing:</strong> Routing production traffic to a new model version without affecting the user experience, enabling performance comparisons.</li><li><strong>Instant Rollback:</strong> Immediate switching of traffic away from a problematic new model version back to a stable one.</li></ul><p>But is this the whole story? Not quite. Introducing this routing layer adds complexity to our model deployment cycles. In addition, we need a mechanism to collect the context-based routing information from the researchers when they choose to deploy model variants.</p><h3>The Glue — Switchboard Rules</h3><p>Given that Objectives serve as the contract between clients and the serving platform, we needed a way for researchers to attach model variants, experiments, and traffic splits to those Objectives without changing client code. This is where Switchboard Rules comes in.</p><p>The primary UX for model researchers to define models associated with an objective in a flexible manner is a JavaScript configuration, which we call <em>Switchboard Rules</em>. It’s used to produce a set of rules (typically a JSON file) that primarily dictate the following things to the serving platform:</p><ol><li>The default model to use for a given Objective</li><li>A/B experiments to configure for a set of Objectives and the corresponding models to load for those experiments</li><li>Customizations to gradually shift traffic to a new model</li></ol><p>Here is an example of an A/B test rule in the context of the Continue Watching row:</p><pre>/**<br>Configuration rule written by a Model Researcher to add an A/B experiment in the Model Serving system.<br>Cell 1: Uses the default, currently productized model<br>Cell 2 and Cell 3: Use different experimental (candidate) models<br>**/<br><br>function defineAB12345Rule() {<br>    const abTestId = 12345;<br><br>    const objectives = Objectives.ContinueWatchingRanking;<br>    const abTestCellToModel = {<br>        1: {name: "netflix-continue-watching-model-default"},<br>        2: {name: "netflix-continue-watching-model-cell-2"},<br>        3: {name: "netflix-continue-watching-model-cell-3"}<br>    };<br><br>    return {<br>        cellToModel: abTestCellToModel,<br>        abTestId: abTestId,<br>        targetObjectives: [objectives],<br>        modelInputType: constants.TITLE_INPUT_TYPE,<br>        modelType: 'SCORER'<br>    };<br>}</pre><p>These rules are consumed by both the Switchboard and the Model Serving clusters. Given these rules, the serving platform components can take various actions, some detailed below:</p><p><strong>Control Plane Flow</strong>:</p><ol><li><strong>Assignment:</strong> Produce model-to-cluster shard assignment.</li><li><strong>Validation:</strong> Load all specified models into the Serving Cluster Shard and validate model dependencies to ensure successful execution.</li><li><strong>Mapping:</strong> Provide the model-to-shard VIP address mapping to Switchboard.</li></ol><p><strong>Data Plane Flow</strong>:</p><ol><li><strong>Allocation:</strong> If the request is for Objective=ContinueWatchingRanking, query the <a href="https://netflixtechblog.com/its-all-a-bout-testing-the-netflix-experimentation-platform-4e1ca458c15">Experimentation Platform</a> for the userId’s cell allocation.</li><li><strong>Model Selection:</strong> Use the allocation and A/B test rule to select the appropriate model.</li><li><strong>Request Routing:</strong> Route the request to the serving cluster shard with the selected model and context.</li><li><strong>Model Execution (on the serving host):</strong> Run the model workflow steps and return the response.</li></ol><p>A key highlight of this setup is the decoupling of the experimentation config from the serving platform code. This includes having an independent release cycle for the rules, separate from the code deployments. <a href="https://netflixtechblog.com/how-netflix-microservices-tackle-dataset-pub-sub-4a068adcc9a">Netflix’s Gutenberg</a> system provides an excellent ecosystem that enables a flexible pub-sub architecture, facilitating proper versioning, dynamic loading, easy rollbacks, and more. Both Switchboard and the Serving Cluster Host subscribe to the same Switchboard Rules configuration.</p><p>To prevent race conditions and ensure proper sync of the dynamic Switchboard Rules configuration, the following flow is considered:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/944/1*iSNuD8ZuSC9E2zN-wZdH-A.png"></figure><h3>Evolving Challenges</h3><p>Switchboard solved the primary problem of improving model iteration and innovation velocity, and provided an excellent ML serving abstraction to over 30 service clients. However, as the system scale increased, a few challenges and problems with this design became apparent:</p><ul><li><strong>Single point of failure: </strong>The presence of Switchboard in the critical request path clearly highlights the risks of shutting down access to all serving hosts in extreme cases, such as unintentional bugs or noisy neighbors sending excessive traffic.</li><li><em>Why this matters: Switchboard became a shared dependency whose failure would degrade or disable multiple ML-powered experiences at Netflix.</em></li><li><strong>Added latency due to additional network hop:</strong> Switchboard in the request path adds between 10–20ms of latency due to serialization-deserialization operations, depending on payload size. Additionally, it further exposes a request to tail latency amplification.</li><li><em>Why this matters: The added latency is unacceptable for some latency-sensitive clients, resulting in end-user impact due to service timeouts.</em></li><li><strong>Reduced Client flexibility</strong>: Switchboard obscures visibility into client request origins from the serving clusters. Consequently, distinguishing data logged for real vs artificial traffic, which is essential for model training, is difficult and requires ongoing customization and increased MLOps overhead.</li><li><em>Why this matters: It makes it harder to do tenant separation and test traffic isolation.</em></li></ul><h3>What Next? — Lightbulb</h3><p>The aforementioned challenges of operating Switchboard at scale forced us to rethink the core implementation while retaining its key features. Our goal was not to throw away Switchboard’s design, but to refactor where and how its responsibilities were executed, keeping the benefits while reducing risk and latency. Particularly:</p><ul><li><em>Common Client Abstraction</em></li><li><em>Decouple clients from model sharding</em></li><li><em>Flexible traffic routing rules</em></li><li><em>Lightweight system client</em></li><li><em>Single place to define model and experimentation config</em></li><li><em>Fast experimentation config propagation</em></li><li><em>Fallback and client-side caching in case of failures</em></li></ul><p>However, we did want to address some of the previous design choices to move forward with:</p><ul><li><strong>Remove the routing service from the direct request path: </strong>Having a single service in the active request path introduces another failure mode and limits fallback flexibility. While routing rules change infrequently, maintaining consistency comes at the cost of increased availability risks.</li><li><strong>Separate model inputs from the request metadata</strong>: In certain cases, the request payload could be quite large. Needing to deserialize and then re-serialize the payload as it flowed through Switchboard to make a routing decision was a significant contributor to latency and increased serving costs.</li><li><strong>Provide better isolation for the routing layer: </strong>Consolidating multiple use cases (tenants) into a single routing cluster poses two main challenges. First, error propagation posed a risk, as a surge of problematic requests from one tenant could cascade errors back to Switchboard, potentially impacting other users. Second, the cluster had to accommodate diverse latency requirements because the requests from different use cases varied significantly in complexity.</li></ul><p>This required some changes in our setup flow: While it largely remained unchanged, however, we created separate components for Routing and Model Selection (Lightbulb):</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/897/1*wtVpe5xJMNEENitvkd1FPQ.png"></figure><p>We now take the rules for an Objective and break them into distinct sets of configuration:</p><ul><li><strong>Model Serving Configuration</strong>: This allows us to determine which model should be used at request time, along with the required metadata</li><li><strong>Routing Rules</strong>: Given a model we want to serve at request time, this tells us which VIP the request should be routed to.</li></ul><p>The Data Plane changes also reflect this separation, as we now rely on <a href="https://github.com/envoyproxy/envoy">Envoy</a> to take care of the routing details:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/912/1*jbW4NlcnKucGKBi_vNjXbw.png"></figure><p>Envoy is <a href="https://netflixtechblog.com/zero-configuration-service-mesh-with-on-demand-cluster-discovery-ac6483b52a51">already used</a> for all egress communication between apps at Netflix, and it can route requests to different clusters (VIPs) based on the configurable Routing Rules published from our control plane. However, it lacks the information needed to make routing decisions and the ability to enrich the request body with additional serving parameters required for A/B testing model variants. We introduced Lightbulb to cover this gap:</p><ul><li>Lightbulb consumes the minimal request context, which contains use-case information, and provides the metadata mapping required for routing at the Envoy layer.</li><li>Lightbulb resolves the request context to determine a routingKey configuration along with the <strong>ObjectiveConfig</strong> — this is where we place the model id along with other request-specific configurations required for model execution. This is done to separate the config resolution associated with the request from the placement and routing information needed to reach it on the inference cluster.</li><li>While the routingKey is added to the headers for Envoy proxy to consume, the client adds the ObjectiveConfig parameters to the request itself. This is done to avoid bloating the request headers while passing additional parameters for the model to process the request appropriately.</li><li>The routing of the actual request is performed by the Envoy proxy, which has the metadata to map the routingKey to the actual cluster VIP running the model. Because the routingKey is in a header, this determination can be made with minimal overhead.</li></ul><p>These changes retain the advantages of Switchboard, such as a single integration point, abstraction of model id from use case, context-aware routing, while addressing the challenges we observed over time.</p><h3>Conclusion</h3><p>The evolution from Switchboard to Lightbulb marks a significant architectural refinement in our ML model serving infrastructure. While Switchboard provided the initial abstraction layer critical for rapid innovation, its latency and single-point-of-failure risk posed scaling hurdles. The subsequent adoption of Lightbulb, a decoupled service focused solely on routing metadata, and its integration with Envoy successfully resolved these challenges. This sophisticated new architecture preserves the key benefits — seamless client integration and flexible experimentation — while ensuring reliable, efficient, and scalable delivery of personalized member experiences, positioning us well for future ML growth.</p><p>In future posts in this series, we’ll dive deeper into other aspects of our ML serving platform, including inference and feature fetching, and how they interact with the routing architecture described here.</p><p>Special thanks to <strong>Sura Elamurugu</strong>, <strong>Sri Krishna Vempati</strong>, <strong>Ed Maddox</strong>, and <strong>Sreepathi Prasanna</strong> for their invaluable feedback and partnership in iterating on this idea and bringing this blog post to life.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=16e22fe18741" width="1" height="1" alt=""><hr><p><a href="https://medium.com/netflix-techblog/state-of-routing-in-model-serving-16e22fe18741">State of Routing in Model Serving</a> was originally published in <a href="https://medium.com/netflix-techblog">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://medium.com/netflix-techblog/state-of-routing-in-model-serving-16e22fe18741</link>
      <guid>https://medium.com/netflix-techblog/state-of-routing-in-model-serving-16e22fe18741</guid>
      <pubDate>Fri, 01 May 2026 23:03:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Scaling Camera File Processing at Netflix]]></title>
      <description><![CDATA[<p><em>Orchestrating Media Workflows Through Strategic Collaboration</em></p><p>Authors: <a href="https://www.linkedin.com/in/ericreinecke/">Eric Reinecke</a>, <a href="https://www.linkedin.com/in/bhanusrikanth/">Bhanu Srikanth</a></p><h3>Introduction to Content Hub’s Media Production Suite</h3><p>At Netflix, we want to provide filmmakers with the tools they need to produce content at a global scale, with quick turnaround and choice from an extraordinary variety of cameras, formats, workflows, and collaborators. Every series or film arrives with its own creative ambitions and technical requirements. To reduce friction and keep productions moving smoothly, we built <a href="https://netflixtechblog.com/globalizing-productions-with-netflixs-media-production-suite-fc3c108c0a22">Netflix’s Media Production Suite (MPS)</a> with the goal of automating repeatable tasks, standardizing key workflows, and giving productions more time to focus on creative collaboration and craftsmanship.</p><p>A critical part of this effort is how we handle image processing and camera metadata across the hundreds of hours and terabytes of camera footage that Netflix productions ingest on a daily basis. Rather than build every component from scratch, we chose to partner where it made sense–especially in areas where the industry already had trusted, battle-tested solutions.</p><p>This article explores how Netflix’s Media Production Suite integrates with FilmLight’s API (FLAPI) as the core studio media processing engine in Netflix’s cloud compute infrastructure, and how that collaboration helps us deliver smarter, more reliable workflows at scale.</p><h3>Why We Built MPS</h3><p>As Netflix’s production slate grew, so did the complexity of file-based workflows. We saw recurring challenges across productions:</p><ul><li>File wrangling sapping time from creative decision-making</li><li>Inconsistent media handling across shows, regions, or vendors</li><li>Difficult to audit manual processes that are prone to human error</li><li>Duplication of effort as teams reinvented similar workflows for each production</li></ul><p>Content Hub Media Production Suite was created to address these pain points. MPS is designed to:</p><ul><li>Bring efficiency, consistency, and quality control to global productions</li><li>Streamline media management and movement from production through post-production</li><li>Reduce time spent on non-creative file management</li><li>Minimize human error while maximizing creative time</li></ul><p>To achieve this, MPS needed a robust, flexible, and trusted way to handle camera-original media and metadata at scale.</p><h3>The Right Tool for the Job</h3><p>From the start, we knew that building a world-class image processing engine in-house is a significant, long-term commitment: one that would require deep, continuous collaboration with camera manufacturers and the wider industry.</p><p>When designing the system, we set out some core requirements:</p><ul><li><strong>Inspect, trim, and transcode original camera files and metadata</strong> for any Netflix production with trusted color science</li><li><strong>Support a wide variety of cameras and recording formats</strong> used worldwide while staying current as new ones are released</li><li><strong>Run well in our paved-path encoding infrastructure,</strong> enabling us to take advantage of proven compute and storage scalability with robust observability</li></ul><p>FilmLight develops Baselight and Daylight, which are commonly used in the industry for color grading, dailies, and transcoding. Their FilmLight API (FLAPI) allows us to use that same media processing engine as a backend API.</p><p>Rather than duplicating that work, we chose to integrate. FilmLight became a trusted technology partner, and FLAPI is now a foundational part of how MPS processes media.</p><h3>The Media Processing Engine</h3><p>MPS is not a single application; it’s an ecosystem of tools and services that support Netflix productions globally. Within that ecosystem, the FilmLight API plays the following key roles.</p><ol><li>Parsing camera metadata on ingest</li></ol><p>Productions upload media to Netflix’s <strong>Content Hub</strong> with <a href="https://theasc.com/society/ascmitc/asc-media-hash-list">ASC MHL</a> (Media Hash List) files to ensure completeness and integrity of initial ingest, but soon after, it’s important to understand the technical characteristics of each piece of media. We call this workflow phase “inspection.”</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*OYBDXSUJ6D0tVXVGO3w5TQ.jpeg"><figcaption>Footage ingested with MPS is inspected using FLAPI and all metadata is indexed and stored</figcaption></figure><p>At this stage, we:</p><ul><li>Use FLAPI to gather <strong>camera metadata</strong> from the original camera files</li><li>Conform the workflow critical fields to <strong>Netflix’s normalized schema</strong></li><li>Make it <strong>searchable and reusable</strong> for downstream processes</li></ul><p>This metadata is integral to:</p><ul><li>Matching footage based on timing and reel name for automated retrieval</li><li>Debugging (e.g., why a shot looks a certain way after processing)</li><li>Validations and checks across the pipeline</li></ul><p>FLAPI provides consistent, camera-aware insight into footage that may have originated anywhere in the world. Additionally, since we’re able to package FLAPI in a Docker image, we can deploy almost identical code to both cloud and our production compute and storage centers around the world, ensuring a consistent assessment of footage wherever it may exist.</p><p>2. Generating VFX plates and other deliverables</p><p>Visual effects workflows constantly push image processing pipelines to their absolute limits. For MPS to succeed, it must generate images with <strong>accurate</strong> framing, <strong>consistent</strong> color management, and <strong>correct</strong> debayering/decoding parameters — all while maintaining rapid turnaround times.</p><p>To achieve this, we leverage Netflix’s <a href="https://netflixtechblog.com/the-netflix-cosmos-platform-35c14d9351ad">Cosmos</a> compute and storage platform and use open standards to provide predictable and consistent creative control.</p><p>At this phase, we use the FilmLight API to:</p><ul><li><strong>Debayer</strong> original camera files with the correct format-specific decoding parameters</li><li>Crop and de-squeeze images using <strong>Framing Decision Lists (ASC FDL)</strong> to ensure spatial creative decisions are preserved</li><li><strong>Apply ACES Metadata Files (AMF), </strong>providing repeatable color pipelines from dailies through finishing</li><li>Generate <strong>an array of media deliverables</strong> in varied formats</li></ul><p>These processes are automated, repeatable, and auditable. We deliver AMFs alongside the OpenEXRs to ensure recipients know exactly what color transforms are already applied, and which need to be applied to match dailies.</p><p>Because we use FilmLight’s tools on the backend, our workflow specialists can use Baselight on their workstations to manually validate pipeline decisions for productions before the first day of principal photography.</p><h3>The Media Processing Factory in the Cloud</h3><p>Finding an engine that competently processes media in line with open standards is an important part of the equation. To maximize impact, we want to make these tools available to all of the filmmakers we work with. Luckily, we’re no strangers to scaled processing at Netflix, and our <a href="https://netflixtechblog.com/the-netflix-cosmos-platform-35c14d9351ad">Cosmos compute platform</a> was ready for the job!</p><h4>Cloud-first integration</h4><p>The traditional model for this kind of processing in filmmaking has been to invest in beefy computers with large GPUs and high-performance storage arrays to rip through debayering and encoding at breakneck speed. However, constraints in the cloud environment are different.</p><p>Factors that are essential for tools in our runtime environment include that they:</p><ul><li>Are <strong>packageable as Serverless Functions in Linux Docker images</strong> that can be quickly invoked to run a single unit of work and shut down on completion</li><li>Can <strong>run on CPU-only instances</strong> to allow us to take advantage of a wide array of available compute</li><li>Support <strong>headless invocation </strong>via Java, Python, or CLI</li><li><strong>Operate statelessly,</strong> so when things do go wrong, we can simply terminate and re-launch the worker</li></ul><p>Operating within these constraints lets us focus on increasing throughput via parallel encoding rather than focusing on single-instance processing power. We can then target the sweet spot of the cost/performance efficiency curve while still hitting our target turnaround times.</p><p>When tools are API-driven, easily packaged in Linux containers, and don’t require a lot of external state management, Netflix can quickly integrate and deploy them with operational reliability. FilmLight API fit the bill for us. At Netflix, we leverage:</p><ul><li><strong>Java</strong> and <strong>Python</strong> as the primary integration languages</li><li><strong>Ubuntu-based Docker images</strong> with Java and Python code to expose functionality to our workflows</li><li><strong>CPU instances in the cloud and local compute centers</strong> for running inspection, rendering, and trimming jobs</li></ul><p>While FLAPI also supports GPU rendering, CPU instances give us access to a much wider segment of Netflix’s vast encoding compute pool and free up GPU instances for other workloads.</p><p>To use FilmLight API, we bundle it in a package that can be easily installed via a Dockerfile. Then, we built Cosmos Stratum Functions that accept an input clip, output location, and varying parameters such as frame ranges and AMF or FDL files when debayering footage. These functions can be quickly invoked to process a single clip or sub-segment of a clip and shut down again to free up resources.</p><h4>Elastic scaling for production workloads</h4><p>Production workloads are inherently spiky:</p><ul><li>A quiet day on set may mean minimal new footage to inspect.</li><li>A full VFX turnover or pulling trimmed OCF for finishing might require <strong>thousands of parallel renders</strong> in a short time window.</li></ul><p>By deploying FLAPI in the cloud as functions, MPS can:</p><ul><li>Allocate compute on demand and release it when our work queue dies down</li><li>Avoid tying capacity to a fixed pool of local hardware</li><li>Smooth demand across many types of encoding workload in a shared resource pool</li></ul><p>This elasticity lets us swarm pull requests to get them through quickly, then immediately yield resources back to lower priority workloads. Even in peak production periods, we avoid the pain of manually managing render queues and prioritization by avoiding fixed resource allocation. All this means <strong>lightning-fast</strong> turnaround times and <strong>less anxiety</strong> around deadlines for our filmmakers.</p><h3>Designed for Seasoned Pros and Emerging Filmmakers</h3><p>Netflix productions range from highly experienced teams with very specific workflows to newer teams who may be less familiar with potential pitfalls in complex file-based pipelines.</p><p>MPS is designed to support both:</p><ul><li>Industry veterans who need to configure precise, bespoke workflows and trust that underlying image processing will respect those decisions.</li><li>Productions without a color scientist on staff — those who benefit from guardrails and sane defaults that help them avoid common workflow issues (e.g., mismatched color transforms, inconsistent debayering, or incomplete metadata handling).</li></ul><p>The partnership with FilmLight lets Netflix focus on workflow design, orchestration, and production support, while FilmLight focuses on providing competent handling of a wide variety of camera formats with world-class image science!</p><h3>Collaboration and Co-Evolution</h3><p>Netflix aimed to integrate MPS into a wider tool ecosystem by developing a comprehensive solution based on emerging open standards, rather than making MPS a self-contained system. Integrating FLAPI into our system requires more than an API reference–it requires ongoing partnership. FilmLight worked closely with Netflix teams to:</p><ul><li>Align on <strong>feature roadmaps</strong>, particularly around new camera formats and open standards</li><li>Validate the <strong>accuracy and performance</strong> of key operations</li><li>Debug <strong>edge cases</strong> discovered in large-scale, real-world workloads</li><li><strong>Evolve the API</strong> in ways that serve both Netflix and the wider industry</li><li>Create <strong>a positive feedback cycle with open standards</strong> like ACES and ASC FDL to solve for gaps when the rubber hits the road</li></ul><p>One example of this has been with the implementation of <a href="https://draftdocs.acescentral.com/background/about-aces-2/">ACES 2</a>. FilmLight’s developers quickly provided a roadmap for support. As our engineering teams collaborated on integration, we also provided feedback to the ACES technical leadership to quickly address integration challenges and test drive updates in our pipeline.</p><p>This collaborative relationship–built on open communication, joint validation, and feedback to the greater industry–is how we routinely work with FilmLight to ensure we’re not just building something that works for our shows, but also driving a healthy tooling and standards ecosystem.</p><h3>Impact</h3><p>While much of this work takes place behind the scenes, its impact is felt directly by our productions. Our goal with building MPS is for producers, post supervisors, and vendors to experience:</p><ul><li>Fewer delays caused by missing, incomplete, or incorrect media</li><li>Faster turnaround on VFX plates and other technical deliverables</li><li>More predictable, consistent handoffs between editorial, color, and VFX</li><li>Less time spent troubleshooting technical issues, and more time focused on creative review</li></ul><p>In practice, this often shows up as the absence of crisis: the time a VFX vendor doesn’t have to request a re-delivery, or the time editorial doesn’t have to wait for corrected plates, or the time the color facility doesn’t have to reinvent a tone-mapping path because the AMF and ACES pipeline are already in place.</p><h3>Looking Ahead</h3><p>As camera technology, codecs, open standards, and production workflows continue to evolve, so will MPS. The guiding principles remain:</p><ul><li>Automate what’s repeatable</li><li>Centralize what benefits from standardization</li><li>Partner where deep domain expertise already exists</li></ul><p>The integration with FilmLight API is one example of this philosophy in action. By treating image processing as a specialized discipline and collaborating with a trusted industry partner, Netflix is delivering smarter, more reliable workflows to productions worldwide.</p><p>At its core, this partnership supports a simple goal: reduce manual workflow and tool management, giving filmmakers more time to tell stories.</p><h3>Acknowledgements</h3><p>This project is the result of collaboration and iteration over many years. In addition to the authors, the following people have contributed to this work:</p><ul><li>Matthew Donato</li><li>Prabh Nallani</li><li>Andy Schuler</li><li>Jesse Korosi</li></ul><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=6dab2b1e80be" width="1" height="1" alt=""><hr><p><a href="https://medium.com/netflix-techblog/scaling-camera-file-processing-at-netflix-6dab2b1e80be">Scaling Camera File Processing at Netflix</a> was originally published in <a href="https://medium.com/netflix-techblog">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://medium.com/netflix-techblog/scaling-camera-file-processing-at-netflix-6dab2b1e80be</link>
      <guid>https://medium.com/netflix-techblog/scaling-camera-file-processing-at-netflix-6dab2b1e80be</guid>
      <pubDate>Fri, 24 Apr 2026 17:06:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Scaling Camera File Processing at Netflix]]></title>
      <description><![CDATA[<p><em>Orchestrating Media Workflows Through Strategic Collaboration</em></p><p>Authors: <a href="https://www.linkedin.com/in/ericreinecke/">Eric Reinecke</a>, <a href="https://www.linkedin.com/in/bhanusrikanth/">Bhanu Srikanth</a></p><h3>Introduction to Content Hub’s Media Production Suite</h3><p>At Netflix, we want to provide filmmakers with the tools they need to produce content at a global scale, with quick turnaround and choice from an extraordinary variety of cameras, formats, workflows, and collaborators. Every series or film arrives with its own creative ambitions and technical requirements. To reduce friction and keep productions moving smoothly, we built <a href="https://netflixtechblog.com/globalizing-productions-with-netflixs-media-production-suite-fc3c108c0a22">Netflix’s Media Production Suite (MPS)</a> with the goal of automating repeatable tasks, standardizing key workflows, and giving productions more time to focus on creative collaboration and craftsmanship.</p><p>A critical part of this effort is how we handle image processing and camera metadata across the hundreds of hours and terabytes of camera footage that Netflix productions ingest on a daily basis. Rather than build every component from scratch, we chose to partner where it made sense–especially in areas where the industry already had trusted, battle-tested solutions.</p><p>This article explores how Netflix’s Media Production Suite integrates with FilmLight’s API (FLAPI) as the core studio media processing engine in Netflix’s cloud compute infrastructure, and how that collaboration helps us deliver smarter, more reliable workflows at scale.</p><h3>Why We Built MPS</h3><p>As Netflix’s production slate grew, so did the complexity of file-based workflows. We saw recurring challenges across productions:</p><ul><li>File wrangling sapping time from creative decision-making</li><li>Inconsistent media handling across shows, regions, or vendors</li><li>Difficult to audit manual processes that are prone to human error</li><li>Duplication of effort as teams reinvented similar workflows for each production</li></ul><p>Content Hub Media Production Suite was created to address these pain points. MPS is designed to:</p><ul><li>Bring efficiency, consistency, and quality control to global productions</li><li>Streamline media management and movement from production through post-production</li><li>Reduce time spent on non-creative file management</li><li>Minimize human error while maximizing creative time</li></ul><p>To achieve this, MPS needed a robust, flexible, and trusted way to handle camera-original media and metadata at scale.</p><h3>The Right Tool for the Job</h3><p>From the start, we knew that building a world-class image processing engine in-house is a significant, long-term commitment: one that would require deep, continuous collaboration with camera manufacturers and the wider industry.</p><p>When designing the system, we set out some core requirements:</p><ul><li><strong>Inspect, trim, and transcode original camera files and metadata</strong> for any Netflix production with trusted color science</li><li><strong>Support a wide variety of cameras and recording formats</strong> used worldwide while staying current as new ones are released</li><li><strong>Run well in our paved-path encoding infrastructure,</strong> enabling us to take advantage of proven compute and storage scalability with robust observability</li></ul><p>FilmLight develops Baselight and Daylight, which are commonly used in the industry for color grading, dailies, and transcoding. Their FilmLight API (FLAPI) allows us to use that same media processing engine as a backend API.</p><p>Rather than duplicating that work, we chose to integrate. FilmLight became a trusted technology partner, and FLAPI is now a foundational part of how MPS processes media.</p><h3>The Media Processing Engine</h3><p>MPS is not a single application; it’s an ecosystem of tools and services that support Netflix productions globally. Within that ecosystem, the FilmLight API plays the following key roles.</p><ol><li>Parsing camera metadata on ingest</li></ol><p>Productions upload media to Netflix’s <strong>Content Hub</strong> with <a href="https://theasc.com/society/ascmitc/asc-media-hash-list">ASC MHL</a> (Media Hash List) files to ensure completeness and integrity of initial ingest, but soon after, it’s important to understand the technical characteristics of each piece of media. We call this workflow phase “inspection.”</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*OYBDXSUJ6D0tVXVGO3w5TQ.jpeg"><figcaption>Footage ingested with MPS is inspected using FLAPI and all metadata is indexed and stored</figcaption></figure><p>At this stage, we:</p><ul><li>Use FLAPI to gather <strong>camera metadata</strong> from the original camera files</li><li>Conform the workflow critical fields to <strong>Netflix’s normalized schema</strong></li><li>Make it <strong>searchable and reusable</strong> for downstream processes</li></ul><p>This metadata is integral to:</p><ul><li>Matching footage based on timing and reel name for automated retrieval</li><li>Debugging (e.g., why a shot looks a certain way after processing)</li><li>Validations and checks across the pipeline</li></ul><p>FLAPI provides consistent, camera-aware insight into footage that may have originated anywhere in the world. Additionally, since we’re able to package FLAPI in a Docker image, we can deploy almost identical code to both cloud and our production compute and storage centers around the world, ensuring a consistent assessment of footage wherever it may exist.</p><p>2. Generating VFX plates and other deliverables</p><p>Visual effects workflows constantly push image processing pipelines to their absolute limits. For MPS to succeed, it must generate images with <strong>accurate</strong> framing, <strong>consistent</strong> color management, and <strong>correct</strong> debayering/decoding parameters — all while maintaining rapid turnaround times.</p><p>To achieve this, we leverage Netflix’s <a href="https://netflixtechblog.com/the-netflix-cosmos-platform-35c14d9351ad">Cosmos</a> compute and storage platform and use open standards to provide predictable and consistent creative control.</p><p>At this phase, we use the FilmLight API to:</p><ul><li><strong>Debayer</strong> original camera files with the correct format-specific decoding parameters</li><li>Crop and de-squeeze images using <strong>Framing Decision Lists (ASC FDL)</strong> to ensure spatial creative decisions are preserved</li><li><strong>Apply ACES Metadata Files (AMF), </strong>providing repeatable color pipelines from dailies through finishing</li><li>Generate <strong>an array of media deliverables</strong> in varied formats</li></ul><p>These processes are automated, repeatable, and auditable. We deliver AMFs alongside the OpenEXRs to ensure recipients know exactly what color transforms are already applied, and which need to be applied to match dailies.</p><p>Because we use FilmLight’s tools on the backend, our workflow specialists can use Baselight on their workstations to manually validate pipeline decisions for productions before the first day of principal photography.</p><h3>The Media Processing Factory in the Cloud</h3><p>Finding an engine that competently processes media in line with open standards is an important part of the equation. To maximize impact, we want to make these tools available to all of the filmmakers we work with. Luckily, we’re no strangers to scaled processing at Netflix, and our <a href="https://netflixtechblog.com/the-netflix-cosmos-platform-35c14d9351ad">Cosmos compute platform</a> was ready for the job!</p><h4>Cloud-first integration</h4><p>The traditional model for this kind of processing in filmmaking has been to invest in beefy computers with large GPUs and high-performance storage arrays to rip through debayering and encoding at breakneck speed. However, constraints in the cloud environment are different.</p><p>Factors that are essential for tools in our runtime environment include that they:</p><ul><li>Are <strong>packageable as Serverless Functions in Linux Docker images</strong> that can be quickly invoked to run a single unit of work and shut down on completion</li><li>Can <strong>run on CPU-only instances</strong> to allow us to take advantage of a wide array of available compute</li><li>Support <strong>headless invocation </strong>via Java, Python, or CLI</li><li><strong>Operate statelessly,</strong> so when things do go wrong, we can simply terminate and re-launch the worker</li></ul><p>Operating within these constraints lets us focus on increasing throughput via parallel encoding rather than focusing on single-instance processing power. We can then target the sweet spot of the cost/performance efficiency curve while still hitting our target turnaround times.</p><p>When tools are API-driven, easily packaged in Linux containers, and don’t require a lot of external state management, Netflix can quickly integrate and deploy them with operational reliability. FilmLight API fit the bill for us. At Netflix, we leverage:</p><ul><li><strong>Java</strong> and <strong>Python</strong> as the primary integration languages</li><li><strong>Ubuntu-based Docker images</strong> with Java and Python code to expose functionality to our workflows</li><li><strong>CPU instances in the cloud and local compute centers</strong> for running inspection, rendering, and trimming jobs</li></ul><p>While FLAPI also supports GPU rendering, CPU instances give us access to a much wider segment of Netflix’s vast encoding compute pool and free up GPU instances for other workloads.</p><p>To use FilmLight API, we bundle it in a package that can be easily installed via a Dockerfile. Then, we built Cosmos Stratum Functions that accept an input clip, output location, and varying parameters such as frame ranges and AMF or FDL files when debayering footage. These functions can be quickly invoked to process a single clip or sub-segment of a clip and shut down again to free up resources.</p><h4>Elastic scaling for production workloads</h4><p>Production workloads are inherently spiky:</p><ul><li>A quiet day on set may mean minimal new footage to inspect.</li><li>A full VFX turnover or pulling trimmed OCF for finishing might require <strong>thousands of parallel renders</strong> in a short time window.</li></ul><p>By deploying FLAPI in the cloud as functions, MPS can:</p><ul><li>Allocate compute on demand and release it when our work queue dies down</li><li>Avoid tying capacity to a fixed pool of local hardware</li><li>Smooth demand across many types of encoding workload in a shared resource pool</li></ul><p>This elasticity lets us swarm pull requests to get them through quickly, then immediately yield resources back to lower priority workloads. Even in peak production periods, we avoid the pain of manually managing render queues and prioritization by avoiding fixed resource allocation. All this means <strong>lightning-fast</strong> turnaround times and <strong>less anxiety</strong> around deadlines for our filmmakers.</p><h3>Designed for Seasoned Pros and Emerging Filmmakers</h3><p>Netflix productions range from highly experienced teams with very specific workflows to newer teams who may be less familiar with potential pitfalls in complex file-based pipelines.</p><p>MPS is designed to support both:</p><ul><li>Industry veterans who need to configure precise, bespoke workflows and trust that underlying image processing will respect those decisions.</li><li>Productions without a color scientist on staff — those who benefit from guardrails and sane defaults that help them avoid common workflow issues (e.g., mismatched color transforms, inconsistent debayering, or incomplete metadata handling).</li></ul><p>The partnership with FilmLight lets Netflix focus on workflow design, orchestration, and production support, while FilmLight focuses on providing competent handling of a wide variety of camera formats with world-class image science!</p><h3>Collaboration and Co-Evolution</h3><p>Netflix aimed to integrate MPS into a wider tool ecosystem by developing a comprehensive solution based on emerging open standards, rather than making MPS a self-contained system. Integrating FLAPI into our system requires more than an API reference–it requires ongoing partnership. FilmLight worked closely with Netflix teams to:</p><ul><li>Align on <strong>feature roadmaps</strong>, particularly around new camera formats and open standards</li><li>Validate the <strong>accuracy and performance</strong> of key operations</li><li>Debug <strong>edge cases</strong> discovered in large-scale, real-world workloads</li><li><strong>Evolve the API</strong> in ways that serve both Netflix and the wider industry</li><li>Create <strong>a positive feedback cycle with open standards</strong> like ACES and ASC FDL to solve for gaps when the rubber hits the road</li></ul><p>One example of this has been with the implementation of <a href="https://draftdocs.acescentral.com/background/about-aces-2/">ACES 2</a>. FilmLight’s developers quickly provided a roadmap for support. As our engineering teams collaborated on integration, we also provided feedback to the ACES technical leadership to quickly address integration challenges and test drive updates in our pipeline.</p><p>This collaborative relationship–built on open communication, joint validation, and feedback to the greater industry–is how we routinely work with FilmLight to ensure we’re not just building something that works for our shows, but also driving a healthy tooling and standards ecosystem.</p><h3>Impact</h3><p>While much of this work takes place behind the scenes, its impact is felt directly by our productions. Our goal with building MPS is for producers, post supervisors, and vendors to experience:</p><ul><li>Fewer delays caused by missing, incomplete, or incorrect media</li><li>Faster turnaround on VFX plates and other technical deliverables</li><li>More predictable, consistent handoffs between editorial, color, and VFX</li><li>Less time spent troubleshooting technical issues, and more time focused on creative review</li></ul><p>In practice, this often shows up as the absence of crisis: the time a VFX vendor doesn’t have to request a re-delivery, or the time editorial doesn’t have to wait for corrected plates, or the time the color facility doesn’t have to reinvent a tone-mapping path because the AMF and ACES pipeline are already in place.</p><h3>Looking Ahead</h3><p>As camera technology, codecs, open standards, and production workflows continue to evolve, so will MPS. The guiding principles remain:</p><ul><li>Automate what’s repeatable</li><li>Centralize what benefits from standardization</li><li>Partner where deep domain expertise already exists</li></ul><p>The integration with FilmLight API is one example of this philosophy in action. By treating image processing as a specialized discipline and collaborating with a trusted industry partner, Netflix is delivering smarter, more reliable workflows to productions worldwide.</p><p>At its core, this partnership supports a simple goal: reduce manual workflow and tool management, giving filmmakers more time to tell stories.</p><h3>Acknowledgements</h3><p>This project is the result of collaboration and iteration over many years. In addition to the authors, the following people have contributed to this work:</p><ul><li>Matthew Donato</li><li>Prabh Nallani</li><li>Andy Schuler</li><li>Jesse Korosi</li></ul><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=6dab2b1e80be" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/scaling-camera-file-processing-at-netflix-6dab2b1e80be">Scaling Camera File Processing at Netflix</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/scaling-camera-file-processing-at-netflix-6dab2b1e80be</link>
      <guid>https://netflixtechblog.com/scaling-camera-file-processing-at-netflix-6dab2b1e80be</guid>
      <pubDate>Fri, 24 Apr 2026 17:06:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[The Human Infrastructure: How Netflix Built the Operations Layer Behind Live at Scale]]></title>
      <description><![CDATA[<p>By: <a href="https://www.linkedin.com/in/brett-axler-11577142/">Brett Axler</a>, <a href="https://www.linkedin.com/in/casper-choffat-005833b/">Casper Choffat</a>, and <a href="https://www.linkedin.com/in/alexlowry355/">Alo Lowry</a></p><p>In the three years since our first Live show, <a href="https://www.netflix.com/title/80167499"><em>Chris Rock: Selective Outrage</em></a>, we have witnessed an incredible expansion of our live content slate and the live operations that support it. From modest beginnings of streaming just one show per month, we are now capable of streaming over nine shows in a single day, reaching tens of millions of concurrent members. This post pulls back the curtain on the Live Operations teams that enable this rapid scale.</p><h3>Humble Beginnings</h3><p>In March 2023, the engineers who built Netflix’s first live streaming pipeline also operated it. There was no dedicated operations team or formal command center. All of our incident response playbooks were written for SVOD, and SLAs were not designed for the speed of live. For the first live shows on the platform, the engineers who designed what is described in <a href="https://netflixtechblog.com/behind-the-streams-live-at-netflix-part-1-d23f917c2f40">earlier parts of this series</a> monitored dashboards on laptops, coordinated over Slack, and troubleshot in real time while millions of members watched.</p><p>The physical setup matched the operational workflows: improvised. Temporary control rooms were put together in conference rooms. For larger events, Netflix rented third-party broadcast facilities, hardware control panels, multiviewers, and communication panels — the kind of infrastructure that established broadcast networks had built over decades. Every show was a team effort. Engineers and leadership at all levels were involved in every event. Each live show, regardless of size, was a massive effort to launch.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*tibjLPNWGd0PP43Y8Idjag.jpeg"><figcaption>Netflix’s Early Live Operations</figcaption></figure><p>Last month, in March 2026, Netflix streamed the World Baseball Classic live to members in Japan. 47 matches over two weeks, with peak concurrent viewership exceeding 17.9 million for a single game, operations running 24/7 from permanent facilities in Los Gatos and Los Angeles, with international coverage extending to Tokyo. In March alone, Netflix launched approximately 70 live events. That is three events shy of the total number Netflix streamed live in all of 2024. The technical systems that make this possible have been covered in detail across this series. What hasn’t been told is the operational story: the people, procedures, and facilities Netflix built to run those systems in real time, under pressure, with no ability to pause or roll back.</p><h3>The Architecture of Live Operations</h3><p><strong>The Architecture of Live Operations: Evolving the Broadcast Operations Center</strong></p><p>When a technology company transitions into live broadcasting, it faces a unique challenge: blending traditional broadcast television practices with massive-scale live-streaming engineering. At the heart of this intersection is the <strong>Broadcast Operations Center (BOC)</strong>.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*tG3zfmPX5ciRafvtnj9phQ.jpeg"><figcaption>The Transmission Operations Center in Los Angeles</figcaption></figure><p>The BOC serves as the critical “cockpit” for live events. It is the physical command center where a fully produced video feed is received directly from a stadium or venue and then handed off to the live streaming infrastructure. Everything from signal ingest, inspection, and conditioning to closed-captioning, graphics insertion, and ad management happens within these walls. By utilizing a hub-and-spoke model with highly redundant architectures, such as dual internet circuits and SMPTE 2022–7 seamless switching technologies, the BOC replaces direct, vulnerable paths from the venue to the live streaming pipeline, making each live event highly repeatable and far less dependent on the quirks of individual event locations.</p><p><strong>Securing the Signal: Reliability from the Venue</strong> Before the BOC can work its magic, we have to guarantee the video and audio feeds actually survive the journey from the production site to our facility. To ensure absolute reliability from the venue, Netflix enforces strict specifications for live signal contribution.</p><p>For any show-critical feed, meaning the primary feed our members will watch live, we require three completely discrete transmission paths. We utilize a strict hierarchy of approved transmission methods, prioritizing dedicated video fiber and single-feed satellite links, followed by dedicated enterprise-grade internet and robust SRT contribution systems.</p><p>We don’t just rely on redundant transport lines; we require full hardware redundancy out of the production truck itself. This includes using separate router line cards and discrete transmission hardware to prevent any single point of failure. Furthermore, every single piece of transmission hardware at the venue must be powered by two discrete power sources, protected by uninterruptible power supply (UPS) batteries, and surge-conditioned.</p><p>Finally, before we ever go live to millions of viewers, our operators execute exhaustive “FACS/FAX” (facilities checks) testing during rehearsals and before every show. This involves running specialized Audio/Video sync tests, latency tests, and quality tests to guarantee perfect audio and video synchronization, validating closed captions, and touring the backup switcher inputs.</p><p><strong>Building the Human Infrastructure:</strong> Building the human operational model to run a facility like the BOC didn’t happen overnight. For a platform scaling from its very first live comedy special to streaming over 400 global events a year, the operational strategy had to undergo a massive, multi-year evolution.</p><p><strong>Phase 1: The “All-Hands” Engineering Era.</strong> In the earliest days of live streaming, there was no dedicated operations team or formal broadcast operations center. The software engineers who wrote the code and built the live-streaming infrastructure were the same people manually operating the events on launch night. Every show was an “all-hands-on-deck” scenario. While this raw, startup-style approach worked for initial milestones, having core developers manually set up and tear down software configurations for every single broadcast was fundamentally incapable of scaling.</p><p><strong>Phase 2: The Shift to Specialized Engineering (SOEs and BOEs).</strong> To separate event execution from core software development, the operational model matured to introduce specialized engineering teams. First, the <strong>Streaming Operations Engineering (SOE)</strong> team was established. These are highly skilled streaming engineers whose sole focus is to configure the full event on the live pipeline and support it during the broadcast. By having SOEs act as the first line of escalation, the core software developers were freed up to focus on building new live-streaming pipeline features.</p><p>However, as the physical broadcast facilities grew, it became clear that supporting the streaming pipeline wasn’t enough; the physical broadcast hardware and facility workflows needed dedicated oversight too. To solve this, <strong>Broadcast Operations Engineers (BOEs)</strong> were introduced to work alongside the SOEs. The BOE acts as the primary escalation point for all physical broadcast facility and hardware issues, overseeing the operation of all shows during a given shift.</p><p><strong>Phase 3: The “Co-Pilot” Control Room Model.</strong> With specialized engineers in place to handle the deep technical infrastructure, the day-to-day operation of the actual video and audio feeds was handed over to dedicated operators. Initially, the Broadcast Control Rooms were structured much like an airplane cockpit.</p><p>This approach utilized a <strong>“first and second captain” workflow</strong>, pairing two Broadcast Control Operators (BCOs) together to run a single event, functioning exactly like a pilot and co-pilot. This collaborative model allowed for intense focus and high-quality execution, making it the ideal setup for running just one or two live events per day. However, as the ambition grew to stream up to 10 concurrent events a day for massive global tournaments, a 1:1 scale of pairing operators simply required too much space and manpower. A new model had to be adopted.</p><p><strong>Phase 4: The Transmission Operations Center (TOC) Fleet Model.</strong> To manage high-density event days and continuous tournament coverage, the workflow was completely reimagined with the launch of the <strong>Transmission Operations Center (TOC) model</strong>. Rather than treating every live broadcast as an isolated launch in its own room, the TOC treats live events like a fleet. It centralizes operations and distinctly separates the traditional broadcast functions from the streaming functions to maximize human efficiency.</p><p>The TOC model divides the labor across three highly specialized, tiered roles:</p><ul><li><strong>Transmission Control Operator (TCO):</strong> The TCO is responsible for managing all inbound signals arriving from the event venues, such as fiber optic, SRT, and satellite feeds. They ensure these incoming feeds meet strict quality, latency, and operational thresholds. Thanks to centralized dashboarding, a single TCO can manage up to <strong>five events concurrently</strong>.</li><li><strong>Streaming Control Operator (SCO):</strong> While the TCO handles what comes <em>in</em>, the SCO manages what goes <em>out</em>. They oversee all outbound feeds, including the streams heading to the live streaming pipeline and any syndication feeds sent to third parties for commercial distribution. Like the TCOs, SCOs can manage up to <strong>five events concurrently</strong>.</li><li><strong>Broadcast Control Operator (BCO):</strong> With the inbound and outbound transmission mechanics handled by the broader TOC, the BCO is able to focus entirely on the creative and qualitative execution of the event. Operating on a <strong>strict 1:1 ratio</strong> (one operator per event), the BCO seamlessly switches between backup inbound feeds if an issue arises, ensures audio and video remain in perfect synchronization, and performs rigorous quality control. They also monitor critical metadata, such as closed captions and digital ad-insertion messages (SCTE), right before the final polished feed is handed into the live streaming pipeline.</li></ul><p><strong>The Big Bet Exception.</strong> While the fleet-style TOC model enables immense concurrency for daily programming, the most critical, high-visibility events, like major holiday football games, utilize a specialized <strong>Big Bet Model</strong>. For these flagship broadcasts, an entire Broadcast Operations Center is dedicated exclusively to a single event. This hyper-focused environment strips away the multi-event ratios, providing operators with advanced instrumentation and dedicated facility engineers to ensure the absolute highest level of reliability for events where failure is simply not an option.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*n3WWOX4zg06kGyTMWE7mcg.png"><figcaption>Operational Workflow at a Glance (Courtesy of Melissa “Mouse” Merencillo)</figcaption></figure><h3>The Live Command Center (LCC)</h3><p>The Live Command Center (LCC) is not an MCR (Master Control Room). Nor is it a traditional Network Operations Center (NOC). The LCC holds the end-to-end view of quality, health metrics, and reliability for every live stream — from signal ingest at the production venue through cloud encoding, CDN delivery, and playback on member devices — and coordinates the human response when any part of that chain breaks.</p><p>What makes this hard is the data and speed requirements. Standard monitoring tools incur propagation delays of minutes. However, during a live stream, a signal degradation that goes undetected for three minutes can affect millions of members before any mitigation begins. The LCC runs a purpose-built observability stack, the Live Control Center, that aggregates telemetry from across the entire pipeline in near real time: concurrent viewer counts, start failure rates, rebuffer ratios, CDN health, encoder status, and signal path health from the contribution feed forward.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*LN1ncCZL_HLrVfgLiBNK0g.png"><figcaption>Live Control Center (Courtesy of Chris Carey)</figcaption></figure><p>During live events, the system ingests up to 38 million events per second. The LCC’s job is to make that volume of data meaningful and actionable for the small team of operators watching it live.</p><p>Two roles staff the LCC leading up to and during live events. LCC Operations Leads are the shift supervisors and incident commanders. They triage anomalies, make escalation decisions, and own the incident response process from detection through resolution.</p><p>Live Technical Launch Managers (TLMs) function as air traffic controllers: they maintain cross-functional context across more than 45 technical, product, and services teams from encoding, CDN, and playback to social media, customer service, and security teams. TLMs start coordinating with these teams months and sometimes years ahead of a live event to ensure escalation paths and playbooks are in place when the LCC needs to translate a CDN engineer’s concern into a product decision at 2am while a game is still in progress. Together, these roles form the operational leadership layer that keeps engineers focused on building rather than watching dashboards.</p><p>The live operations teams rank shows by three categories:</p><ul><li><strong>Low-Profile Events:</strong> These are lightweight, often lack new features, and anticipate low viewership. They are typically managed with a small team of 1–2 operators and automated alerting.</li><li><strong>High-Profile Events:</strong> These are mid-tier events that warrant more attention due to their size, unique features, or anticipated viewership.</li><li><strong>Big Bet Events:</strong> These represent the highest operational weight, such as an NFL game, with massive viewership expectations and special features. They require the full support of the LCC: a fully staffed physical operations room for the entire duration, active incident command structures, and key engineering teams on standby to support their specific product areas.</li></ul><p>In addition to a show’s event category, the TLMs deployed a Live Operational Level (LOL) model that helps engineers determine whether they need to be on standby, live online, or even in the LCC for any given show.</p><p>Based on the show’s event category, special features, expected viewership, and overall risk, non-operational teams are put into one of four categories:</p><p><strong>Red:</strong> Non-operational teams must remain online for the duration of the event. This is most often seen in large boxing matches and sporting events, such as the NFL Christmas Day games.</p><p><strong>Orange:</strong> Non-operational teams are required to check in online ~30 minutes prior to show and are asked to monitor the health of their systems through the first commercial breaks until the LCC releases them to LOL Yellow.</p><p><strong>Yellow:</strong> Non-operational teams are not required to be online, but should be reachable by page in 2 minutes. Special PagerDuty rotations and verifications are in place to ensure these teams are reachable.</p><p><strong>Grey:</strong> Business as usual. Teams will be reached out to by their normal pager rotation if their help is needed during the show.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*sgk7Y5OYAvbfxL7xRxTQbA.png"><figcaption>Visual Representation of LOL Levels (Courtesy of Gemini Nano Banana Pro)</figcaption></figure><p>By tiering events, Netflix ensures that resource allocation is proportionate to operational needs, preventing a continuous “crisis” mentality and allowing our non-operational partners to focus on their day jobs.</p><p>As of April 2026, most engineering teams are Yellow or Grey, with Ops and Site Reliability Engineers making up most of the teams online to support shows, in addition to engineers performing feature tests.</p><h3>Building the Model</h3><p>The first lesson from 2023 was straightforward: what worked for one show a month would not work for ten shows a week. The engineers who built the pipeline were also the ones operating it, which meant the people best positioned to fix problems were also the ones most likely to be paged at 2am. There was no operational layer to absorb that load.</p><p>In 2024, Netflix streamed 72 live events and began building the team that would eventually run them. The first version of the LCC looked nothing like it does today: a cluster of desks, monitors on stands, and laptops running dashboards, set up in the middle of the office. The TLM team was stood up to own cross-functional coordination for live launches and began formalizing the runbooks, event tiering structure, and incident management protocols that would later enable Netflix to scale operations to support hundreds of shows per year.</p><p>By the time Jake Paul vs. Mike Tyson and the first NFL Christmas Games arrived, the LCC had moved into a dedicated conference room, and partnerships with device and labs teams were producing more effective monitoring tools. But the biggest operational lesson of that period came from communications.</p><p>For Tyson/Paul, Netflix had over 300 people online across engineering, product, and business functions. Some people were online because their support was needed, while many others were just excited to be part of it. Coordinating that many people over Slack and Zoom during an active event with 64 million concurrent streams was unmanageable.</p><p>That experience drove the implementation of a <strong>squad model</strong>: defined teams with clear roles, scoped communication channels, and a single escalation path into the LCC. Around the same time, the LCC began integrating with IP-based communications systems, finally bridging the gap between the command center and the Broadcast Operations Center that had been operating largely in a fractured parallel until then.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*07i_xA6sf2_m05Giu3Dqsw.png"><figcaption>Visual Representation of Squad Operations Model (Courtesy of Gemini Nano Banana Pro)</figcaption></figure><p>2025 brought 220 live events and a permanent LCC facility, along with a dedicated operations team, the Live Command Center Operations Leads. With the growing number of shows, TLMs were getting spread thin, spending more than half their week operating shows late into the evening and over weekends, then getting called back into the office at 9 am to lead critical launch meetings. The addition of the LCC Ops Leads resolved the bandwidth issue by separating planning and operations into distinct roles within a single centralized team.</p><p>As the slate continued to grow and large series like the World Baseball Classic and FIFA Women’s World Cup were announced, the vendor-operator model was introduced, creating an elastic workforce that could scale up for large series events without carrying full-time headcount year-round to support peak capacity. <strong>The key enabler was documentation</strong>: standardized runbooks and onboarding materials detailed enough that a trained operator could reach full effectiveness within their first week. WWE RAW became a weekly operation, normalizing what had previously felt exceptional. By early 2026, multi-event days were no longer a test of capacity but had become the expected operating condition.</p><p>The next chapter is international. Netflix has begun standing up regional Live Operations Center coverage to support live events outside North America, with EMEA operations soon running out of London. The model draws on the same runbooks, tooling, and escalation structures developed in Los Gatos, with follow-the-sun shift handoffs connecting EMEA and US teams across time zones. Looking further ahead, Netflix is planning to bring the LCC and BOC under one roof — a single integrated facility that combines broadcast operations and cloud monitoring into a unified space. The physical separation between those two functions has always introduced friction at the seams. Closing it is the logical next step.</p><h3>Operational Principles for Live at Scale</h3><p>Building a live operations discipline means accepting one constraint above all others: you cannot optimize for efficiency before you have built for reliability.</p><p>Netflix designed for quality first: Standardized runbooks, tiered event structures, pre-documented failure modes, so the 50th show runs as smoothly as the fifth. Off-the-shelf monitoring tools with propagation delays don’t meet that bar. The Netflix Live Control Center and Live Control Room platforms exist because observability at live scale is a product decision that demands the same design rigor as the pipeline it monitors, turning millions of telemetry events per second into something a small team can act on in real time. Technical systems and human systems have to scale together, and the most reliable incident response plan is always the one written before anyone needs it.</p><p>The operational model is also a cultural one. Bringing contingent operators into a proprietary tech stack requires deliberate onboarding design. The vendor model only works when documentation is built to be followed confidently by someone new within their first week. <strong>Beyond process, the most durable parts of how Netflix runs live operations reflect something the </strong><a href="https://jobs.netflix.com/culture"><strong>Netflix culture memo</strong></a><strong> makes explicit: the best ideas come from anywhere.</strong> In practice, that means frontline operators catching issues that engineers miss, vendor staff surfacing workflow friction that improves the system for everyone who follows, and a team that treats candid feedback as standard practice rather than an exception. The technology, the slate, and the scale keep changing. The discipline stays current by staying curious and iterating on the tools, the runbooks, and the team.</p><h3>Conclusion: What’s Next</h3><p>With 2026 already off to a successful start in operational scaling, we’re excited to shift our focus to the upcoming launch of our new Live Broadcast Operations Center in Los Angeles and our new Live Operations Center (LOC) in West London. The LOC will initiate Netflix’s follow-the-sun coverage as live content continues to grow with over 400 live events in 2026, including the launch of 24/7 linear free-to-air broadcast channels with TF1 this summer. On the technical front, further development of automated alerting tools and monitoring by exception will continue to reduce operations’ manual workload.</p><p>In 2023, the engineers led the operations. By 2026, they had developed systems that mostly ran themselves, with a dedicated operational team ensuring they operated smoothly for millions of members. The technology behind Netflix’s Live content has been documented throughout this series, but what runs alongside the tech stack is a set of operational principles, rehearsed incident management processes, and monitoring infrastructure that had to be created from scratch and continues to develop.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*7KV_D2VSRlja_fWmHmuOBw.jpeg"></figure><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*sc4doCe3-h3mi7V9trxBbQ.jpeg"></figure><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*dZKV2OIroXQRYA6UZBZAEw.jpeg"></figure><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*w6eGAv10BZWqNBVXNCvqRA.jpeg"></figure><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*y-iRMBW0Ae2YDiDjNzwq-Q.jpeg"></figure><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*dPoOkp75Rjy5CRskFQBoQQ.jpeg"></figure><p>A special thanks to Te-Yuan Huang, Rob Saltiel, Tara Kozuback, Chris Carey, Di Li, Patrick Li, Anne Aaron, and Melissa “Mouse” Merencillo for their support on this article.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=33e2a311c597" width="1" height="1" alt=""><hr><p><a href="https://medium.com/netflix-techblog/the-human-infrastructure-how-netflix-built-the-operations-layer-behind-live-at-scale-33e2a311c597">The Human Infrastructure: How Netflix Built the Operations Layer Behind Live at Scale</a> was originally published in <a href="https://medium.com/netflix-techblog">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://medium.com/netflix-techblog/the-human-infrastructure-how-netflix-built-the-operations-layer-behind-live-at-scale-33e2a311c597</link>
      <guid>https://medium.com/netflix-techblog/the-human-infrastructure-how-netflix-built-the-operations-layer-behind-live-at-scale-33e2a311c597</guid>
      <pubDate>Fri, 17 Apr 2026 17:01:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[The Human Infrastructure: How Netflix Built the Operations Layer Behind Live at Scale]]></title>
      <description><![CDATA[<p>By: <a href="https://www.linkedin.com/in/brett-axler-11577142/">Brett Axler</a>, <a href="https://www.linkedin.com/in/casper-choffat-005833b/">Casper Choffat</a>, and <a href="https://www.linkedin.com/in/alexlowry355/">Alo Lowry</a></p><p>In the three years since our first Live show, <a href="https://www.netflix.com/title/80167499"><em>Chris Rock: Selective Outrage</em></a>, we have witnessed an incredible expansion of our live content slate and the live operations that support it. From modest beginnings of streaming just one show per month, we are now capable of streaming over nine shows in a single day, reaching tens of millions of concurrent members. This post pulls back the curtain on the Live Operations teams that enable this rapid scale.</p><h3>Humble Beginnings</h3><p>In March 2023, the engineers who built Netflix’s first live streaming pipeline also operated it. There was no dedicated operations team or formal command center. All of our incident response playbooks were written for SVOD, and SLAs were not designed for the speed of live. For the first live shows on the platform, the engineers who designed what is described in <a href="https://netflixtechblog.com/behind-the-streams-live-at-netflix-part-1-d23f917c2f40">earlier parts of this series</a> monitored dashboards on laptops, coordinated over Slack, and troubleshot in real time while millions of members watched.</p><p>The physical setup matched the operational workflows: improvised. Temporary control rooms were put together in conference rooms. For larger events, Netflix rented third-party broadcast facilities, hardware control panels, multiviewers, and communication panels — the kind of infrastructure that established broadcast networks had built over decades. Every show was a team effort. Engineers and leadership at all levels were involved in every event. Each live show, regardless of size, was a massive effort to launch.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*tibjLPNWGd0PP43Y8Idjag.jpeg"><figcaption>Netflix’s Early Live Operations</figcaption></figure><p>Last month, in March 2026, Netflix streamed the World Baseball Classic live to members in Japan. 47 matches over two weeks, with peak concurrent viewership exceeding 9.6 million accounts for a single game, operations running 24/7 from permanent facilities in Los Gatos and Los Angeles, with international coverage extending to Tokyo. In March alone, Netflix launched approximately 70 live events. That is three events shy of the total number Netflix streamed live in all of 2024. The technical systems that make this possible have been covered in detail across this series. What hasn’t been told is the operational story: the people, procedures, and facilities Netflix built to run those systems in real time, under pressure, with no ability to pause or roll back.</p><h3>The Architecture of Live Operations</h3><p><strong>The Architecture of Live Operations: Evolving the Broadcast Operations Center</strong></p><p>When a technology company transitions into live broadcasting, it faces a unique challenge: blending traditional broadcast television practices with massive-scale live-streaming engineering. At the heart of this intersection is the <strong>Broadcast Operations Center (BOC)</strong>.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*tG3zfmPX5ciRafvtnj9phQ.jpeg"><figcaption>The Transmission Operations Center in Los Angeles</figcaption></figure><p>The BOC serves as the critical “cockpit” for live events. It is the physical command center where a fully produced video feed is received directly from a stadium or venue and then handed off to the live streaming infrastructure. Everything from signal ingest, inspection, and conditioning to closed-captioning, graphics insertion, and ad management happens within these walls. By utilizing a hub-and-spoke model with highly redundant architectures, such as dual internet circuits and SMPTE 2022–7 seamless switching technologies, the BOC replaces direct, vulnerable paths from the venue to the live streaming pipeline, making each live event highly repeatable and far less dependent on the quirks of individual event locations.</p><p><strong>Securing the Signal: Reliability from the Venue</strong> Before the BOC can work its magic, we have to guarantee the video and audio feeds actually survive the journey from the production site to our facility. To ensure absolute reliability from the venue, Netflix enforces strict specifications for live signal contribution.</p><p>For any show-critical feed, meaning the primary feed our members will watch live, we require three completely discrete transmission paths. We utilize a strict hierarchy of approved transmission methods, prioritizing dedicated video fiber and single-feed satellite links, followed by dedicated enterprise-grade internet and robust SRT contribution systems.</p><p>We don’t just rely on redundant transport lines; we require full hardware redundancy out of the production truck itself. This includes using separate router line cards and discrete transmission hardware to prevent any single point of failure. Furthermore, every single piece of transmission hardware at the venue must be powered by two discrete power sources, protected by uninterruptible power supply (UPS) batteries, and surge-conditioned.</p><p>Finally, before we ever go live to millions of viewers, our operators execute exhaustive “FACS/FAX” (facilities checks) testing during rehearsals and before every show. This involves running specialized Audio/Video sync tests, latency tests, and quality tests to guarantee perfect audio and video synchronization, validating closed captions, and touring the backup switcher inputs.</p><p><strong>Building the Human Infrastructure:</strong> Building the human operational model to run a facility like the BOC didn’t happen overnight. For a platform scaling from its very first live comedy special to streaming over 400 global events a year, the operational strategy had to undergo a massive, multi-year evolution.</p><p><strong>Phase 1: The “All-Hands” Engineering Era.</strong> In the earliest days of live streaming, there was no dedicated operations team or formal broadcast operations center. The software engineers who wrote the code and built the live-streaming infrastructure were the same people manually operating the events on launch night. Every show was an “all-hands-on-deck” scenario. While this raw, startup-style approach worked for initial milestones, having core developers manually set up and tear down software configurations for every single broadcast was fundamentally incapable of scaling.</p><p><strong>Phase 2: The Shift to Specialized Engineering (SOEs and BOEs).</strong> To separate event execution from core software development, the operational model matured to introduce specialized engineering teams. First, the <strong>Streaming Operations Engineering (SOE)</strong> team was established. These are highly skilled streaming engineers whose sole focus is to configure the full event on the live pipeline and support it during the broadcast. By having SOEs act as the first line of escalation, the core software developers were freed up to focus on building new live-streaming pipeline features.</p><p>However, as the physical broadcast facilities grew, it became clear that supporting the streaming pipeline wasn’t enough; the physical broadcast hardware and facility workflows needed dedicated oversight too. To solve this, <strong>Broadcast Operations Engineers (BOEs)</strong> were introduced to work alongside the SOEs. The BOE acts as the primary escalation point for all physical broadcast facility and hardware issues, overseeing the operation of all shows during a given shift.</p><p><strong>Phase 3: The “Co-Pilot” Control Room Model.</strong> With specialized engineers in place to handle the deep technical infrastructure, the day-to-day operation of the actual video and audio feeds was handed over to dedicated operators. Initially, the Broadcast Control Rooms were structured much like an airplane cockpit.</p><p>This approach utilized a <strong>“first and second captain” workflow</strong>, pairing two Broadcast Control Operators (BCOs) together to run a single event, functioning exactly like a pilot and co-pilot. This collaborative model allowed for intense focus and high-quality execution, making it the ideal setup for running just one or two live events per day. However, as the ambition grew to stream up to 10 concurrent events a day for massive global tournaments, a 1:1 scale of pairing operators simply required too much space and manpower. A new model had to be adopted.</p><p><strong>Phase 4: The Transmission Operations Center (TOC) Fleet Model.</strong> To manage high-density event days and continuous tournament coverage, the workflow was completely reimagined with the launch of the <strong>Transmission Operations Center (TOC) model</strong>. Rather than treating every live broadcast as an isolated launch in its own room, the TOC treats live events like a fleet. It centralizes operations and distinctly separates the traditional broadcast functions from the streaming functions to maximize human efficiency.</p><p>The TOC model divides the labor across three highly specialized, tiered roles:</p><ul><li><strong>Transmission Control Operator (TCO):</strong> The TCO is responsible for managing all inbound signals arriving from the event venues, such as fiber optic, SRT, and satellite feeds. They ensure these incoming feeds meet strict quality, latency, and operational thresholds. Thanks to centralized dashboarding, a single TCO can manage up to <strong>five events concurrently</strong>.</li><li><strong>Streaming Control Operator (SCO):</strong> While the TCO handles what comes <em>in</em>, the SCO manages what goes <em>out</em>. They oversee all outbound feeds, including the streams heading to the live streaming pipeline and any syndication feeds sent to third parties for commercial distribution. Like the TCOs, SCOs can manage up to <strong>five events concurrently</strong>.</li><li><strong>Broadcast Control Operator (BCO):</strong> With the inbound and outbound transmission mechanics handled by the broader TOC, the BCO is able to focus entirely on the creative and qualitative execution of the event. Operating on a <strong>strict 1:1 ratio</strong> (one operator per event), the BCO seamlessly switches between backup inbound feeds if an issue arises, ensures audio and video remain in perfect synchronization, and performs rigorous quality control. They also monitor critical metadata, such as closed captions and digital ad-insertion messages (SCTE), right before the final polished feed is handed into the live streaming pipeline.</li></ul><p><strong>The Big Bet Exception.</strong> While the fleet-style TOC model enables immense concurrency for daily programming, the most critical, high-visibility events, like major holiday football games, utilize a specialized <strong>Big Bet Model</strong>. For these flagship broadcasts, an entire Broadcast Operations Center is dedicated exclusively to a single event. This hyper-focused environment strips away the multi-event ratios, providing operators with advanced instrumentation and dedicated facility engineers to ensure the absolute highest level of reliability for events where failure is simply not an option.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*hRrwHlQ0VdhhztIfPvYMyQ.png"><figcaption>Operational Workflow at a Glance (Courtesy of Melissa “Mouse” Merencillo)</figcaption></figure><h3>The Live Command Center (LCC)</h3><p>The Live Command Center (LCC) is not an MCR (Master Control Room). Nor is it a traditional Network Operations Center (NOC). The LCC holds the end-to-end view of quality, health metrics, and reliability for every live stream — from signal ingest at the production venue through cloud encoding, CDN delivery, and playback on member devices — and coordinates the human response when any part of that chain breaks.</p><p>What makes this hard is the data and speed requirements. Standard monitoring tools incur propagation delays of minutes. However, during a live stream, a signal degradation that goes undetected for three minutes can affect millions of members before any mitigation begins. The LCC runs a purpose-built observability stack, the Live Control Center, that aggregates telemetry from across the entire pipeline in near real time: concurrent viewer counts, start failure rates, rebuffer ratios, CDN health, encoder status, and signal path health from the contribution feed forward.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*LN1ncCZL_HLrVfgLiBNK0g.png"><figcaption>Live Control Center (Courtesy of Chris Carey)</figcaption></figure><p>During live events, the system ingests up to 38 million events per second. The LCC’s job is to make that volume of data meaningful and actionable for the small team of operators watching it live.</p><p>Two roles staff the LCC leading up to and during live events. LCC Operations Leads are the shift supervisors and incident commanders. They triage anomalies, make escalation decisions, and own the incident response process from detection through resolution.</p><p>Live Technical Launch Managers (TLMs) function as air traffic controllers: they maintain cross-functional context across more than 45 technical, product, and services teams from encoding, CDN, and playback to social media, customer service, and security teams. TLMs start coordinating with these teams months and sometimes years ahead of a live event to ensure escalation paths and playbooks are in place when the LCC needs to translate a CDN engineer’s concern into a product decision at 2am while a game is still in progress. Together, these roles form the operational leadership layer that keeps engineers focused on building rather than watching dashboards.</p><p>The live operations teams rank shows by three categories:</p><ul><li><strong>Low-Profile Events:</strong> These are lightweight, often lack new features, and anticipate low viewership. They are typically managed with a small team of 1–2 operators and automated alerting.</li><li><strong>High-Profile Events:</strong> These are mid-tier events that warrant more attention due to their size, unique features, or anticipated viewership.</li><li><strong>Big Bet Events:</strong> These represent the highest operational weight, such as an NFL game, with massive viewership expectations and special features. They require the full support of the LCC: a fully staffed physical operations room for the entire duration, active incident command structures, and key engineering teams on standby to support their specific product areas.</li></ul><p>In addition to a show’s event category, the TLMs deployed a Live Operational Level (LOL) model that helps engineers determine whether they need to be on standby, live online, or even in the LCC for any given show.</p><p>Based on the show’s event category, special features, expected viewership, and overall risk, non-operational teams are put into one of four categories:</p><p><strong>Red:</strong> Non-operational teams must remain online for the duration of the event. This is most often seen in large boxing matches and sporting events, such as the NFL Christmas Day games.</p><p><strong>Orange:</strong> Non-operational teams are required to check in online ~30 minutes prior to show and are asked to monitor the health of their systems through the first commercial breaks until the LCC releases them to LOL Yellow.</p><p><strong>Yellow:</strong> Non-operational teams are not required to be online, but should be reachable by page in 2 minutes. Special PagerDuty rotations and verifications are in place to ensure these teams are reachable.</p><p><strong>Grey:</strong> Business as usual. Teams will be reached out to by their normal pager rotation if their help is needed during the show.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*sgk7Y5OYAvbfxL7xRxTQbA.png"><figcaption>Visual Representation of LOL Levels (Courtesy of Gemini Nano Banana Pro)</figcaption></figure><p>By tiering events, Netflix ensures that resource allocation is proportionate to operational needs, preventing a continuous “crisis” mentality and allowing our non-operational partners to focus on their day jobs.</p><p>As of April 2026, most engineering teams are Yellow or Grey, with Ops and Site Reliability Engineers making up most of the teams online to support shows, in addition to engineers performing feature tests.</p><h3>Building the Model</h3><p>The first lesson from 2023 was straightforward: what worked for one show a month would not work for ten shows a week. The engineers who built the pipeline were also the ones operating it, which meant the people best positioned to fix problems were also the ones most likely to be paged at 2am. There was no operational layer to absorb that load.</p><p>In 2024, Netflix streamed 72 live events and began building the team that would eventually run them. The first version of the LCC looked nothing like it does today: a cluster of desks, monitors on stands, and laptops running dashboards, set up in the middle of the office. The TLM team was stood up to own cross-functional coordination for live launches and began formalizing the runbooks, event tiering structure, and incident management protocols that would later enable Netflix to scale operations to support hundreds of shows per year.</p><p>By the time Jake Paul vs. Mike Tyson and the first NFL Christmas Games arrived, the LCC had moved into a dedicated conference room, and partnerships with device and labs teams were producing more effective monitoring tools. But the biggest operational lesson of that period came from communications.</p><p>For Tyson/Paul, Netflix had over 300 people online across engineering, product, and business functions. Some people were online because their support was needed, while many others were just excited to be part of it. Coordinating that many people over Slack and Zoom during an active event with 64 million concurrent streams was unmanageable.</p><p>That experience drove the implementation of a <strong>squad model</strong>: defined teams with clear roles, scoped communication channels, and a single escalation path into the LCC. Around the same time, the LCC began integrating with IP-based communications systems, finally bridging the gap between the command center and the Broadcast Operations Center that had been operating largely in a fractured parallel until then.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*07i_xA6sf2_m05Giu3Dqsw.png"><figcaption>Visual Representation of Squad Operations Model (Courtesy of Gemini Nano Banana Pro)</figcaption></figure><p>2025 brought 220 live events and a permanent LCC facility, along with a dedicated operations team, the Live Command Center Operations Leads. With the growing number of shows, TLMs were getting spread thin, spending more than half their week operating shows late into the evening and over weekends, then getting called back into the office at 9 am to lead critical launch meetings. The addition of the LCC Ops Leads resolved the bandwidth issue by separating planning and operations into distinct roles within a single centralized team.</p><p>As the slate continued to grow and large series like the World Baseball Classic and FIFA Women’s World Cup were announced, the vendor-operator model was introduced, creating an elastic workforce that could scale up for large series events without carrying full-time headcount year-round to support peak capacity. <strong>The key enabler was documentation</strong>: standardized runbooks and onboarding materials detailed enough that a trained operator could reach full effectiveness within their first week. WWE RAW became a weekly operation, normalizing what had previously felt exceptional. By early 2026, multi-event days were no longer a test of capacity but had become the expected operating condition.</p><p>The next chapter is international. Netflix has begun standing up regional Live Operations Center coverage to support live events outside North America, with EMEA operations soon running out of London. The model draws on the same runbooks, tooling, and escalation structures developed in Los Gatos, with follow-the-sun shift handoffs connecting EMEA and US teams across time zones. Looking further ahead, Netflix is planning to bring the LCC and BOC under one roof — a single integrated facility that combines broadcast operations and cloud monitoring into a unified space. The physical separation between those two functions has always introduced friction at the seams. Closing it is the logical next step.</p><h3>Operational Principles for Live at Scale</h3><p>Building a live operations discipline means accepting one constraint above all others: you cannot optimize for efficiency before you have built for reliability.</p><p>Netflix designed for quality first: Standardized runbooks, tiered event structures, pre-documented failure modes, so the 50th show runs as smoothly as the fifth. Off-the-shelf monitoring tools with propagation delays don’t meet that bar. The Netflix Live Control Center and Live Control Room platforms exist because observability at live scale is a product decision that demands the same design rigor as the pipeline it monitors, turning millions of telemetry events per second into something a small team can act on in real time. Technical systems and human systems have to scale together, and the most reliable incident response plan is always the one written before anyone needs it.</p><p>The operational model is also a cultural one. Bringing contingent operators into a proprietary tech stack requires deliberate onboarding design. The vendor model only works when documentation is built to be followed confidently by someone new within their first week. <strong>Beyond process, the most durable parts of how Netflix runs live operations reflect something the </strong><a href="https://jobs.netflix.com/culture"><strong>Netflix culture memo</strong></a><strong> makes explicit: the best ideas come from anywhere.</strong> In practice, that means frontline operators catching issues that engineers miss, vendor staff surfacing workflow friction that improves the system for everyone who follows, and a team that treats candid feedback as standard practice rather than an exception. The technology, the slate, and the scale keep changing. The discipline stays current by staying curious and iterating on the tools, the runbooks, and the team.</p><h3>Conclusion: What’s Next</h3><p>With 2026 already off to a successful start in operational scaling, we’re excited to shift our focus to the upcoming launch of our new Live Broadcast Operations Center in Los Angeles and our new Live Operations Center (LOC) in West London. The LOC will initiate Netflix’s follow-the-sun coverage as live content continues to grow with over 400 live events in 2026, including the launch of 24/7 linear free-to-air broadcast channels with TF1 this summer. On the technical front, further development of automated alerting tools and monitoring by exception will continue to reduce operations’ manual workload.</p><p>In 2023, the engineers led the operations. By 2026, they had developed systems that mostly ran themselves, with a dedicated operational team ensuring they operated smoothly for millions of members. The technology behind Netflix’s Live content has been documented throughout this series, but what runs alongside the tech stack is a set of operational principles, rehearsed incident management processes, and monitoring infrastructure that had to be created from scratch and continues to develop.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*7KV_D2VSRlja_fWmHmuOBw.jpeg"></figure><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*sc4doCe3-h3mi7V9trxBbQ.jpeg"></figure><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*dZKV2OIroXQRYA6UZBZAEw.jpeg"></figure><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*w6eGAv10BZWqNBVXNCvqRA.jpeg"></figure><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*y-iRMBW0Ae2YDiDjNzwq-Q.jpeg"></figure><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*dPoOkp75Rjy5CRskFQBoQQ.jpeg"></figure><p>A special thanks to Te-Yuan Huang, Rob Saltiel, Tara Kozuback, Chris Carey, Di Li, Patrick Li, Anne Aaron, and Melissa “Mouse” Merencillo for their support on this article.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=33e2a311c597" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/the-human-infrastructure-how-netflix-built-the-operations-layer-behind-live-at-scale-33e2a311c597">The Human Infrastructure: How Netflix Built the Operations Layer Behind Live at Scale</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/the-human-infrastructure-how-netflix-built-the-operations-layer-behind-live-at-scale-33e2a311c597</link>
      <guid>https://netflixtechblog.com/the-human-infrastructure-how-netflix-built-the-operations-layer-behind-live-at-scale-33e2a311c597</guid>
      <pubDate>Fri, 17 Apr 2026 17:01:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Evaluating Netflix Show Synopses with LLM-as-a-Judge]]></title>
      <description><![CDATA[<p>by <a href="https://www.linkedin.com/in/gabrielaalessio/">Gabriela Alessio</a>, <a href="https://www.linkedin.com/in/cameronntaylor/">Cameron Taylor</a>, and <a href="https://www.linkedin.com/in/cwolferesearch/">Cameron R. Wolfe</a></p><h3>Introduction</h3><p>When members log into Netflix, one of the hardest choices is what to watch. The challenge isn’t a lack of options — <em>there are thousands of titles</em> — but finding the most intriguing one is complex and deeply personal. To help, we surface <a href="https://netflixtechblog.com/artwork-personalization-c589f074ad76">personalized promotional assets</a>, especially the show synopsis — <em>a brief description highlighting key plot elements, with cues like genre or talent</em>.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*A2j9Xni86SVnF_VyjFNwfw.png"></figure><p>Strong synopses help members scan, understand, and choose. Poor synopses frustrate, mislead, and drive abandonment. Ensuring high-quality synopses is essential, but scaling quality validation is hard. We host hundreds of thousands of synopses, usually with multiple variants per show. We need to ensure quality at scale so every member gets a consistently great experience every time they read a synopsis. This approach helps us scale high‑quality synopsis coverage for our rapidly expanding catalog, enabling greater speed and coverage without sacrificing quality.</p><p>This report outlines our LLM-based approach for evaluating synopsis quality. Using recent advances in agents, reasoning, and LLM-as-a-Judge, we score four key synopsis quality dimensions, achieving 85%+ agreement with creative writers. Additionally, we show that higher LLM judge quality is correlated with key streaming metrics, <em>allowing us to proactively identify and fix impactful issues weeks or months before a show debuts on Netflix</em>.</p><h3>The Making of a “Good” Synopsis</h3><p>Writing high-quality synopses requires creative expertise. Our expert creative leads are best positioned to craft the creative approaches and define quality standards. However, AI can help us consistently evaluate these expert-driven quality criteria at scale. Synopsis quality at Netflix, which our system aims to predict, is viewed along two dimensions:</p><ol><li><em>Creative Quality</em>: members of our creative writing team assess synopsis quality according to our internal writing guidelines and rubrics.</li><li><em>Member Implicit Feedback</em>: we measure the relative impact of a particular show synopsis on core streaming metrics.</li></ol><p>These two definitions of quality capture distinct and important aspects of quality, one focused upon creative excellence and the other upon utility to members.</p><h4>Creative Quality</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*KzbYGovn903y_ZGysqdnYw.png"></figure><p>For this project, we evaluate synopses against a subset of our creative writing quality rubric — <em>the same criteria to which human writers would adhere</em>. These quality rubrics change over time as quality standards evolve. Given Netflix’s distinctive voice and elevated editorial standards, the quality bar is high. Each criterion has extensive guidelines with examples across regions, genres, and synopsis types.</p><p><strong>Human evaluation.</strong> We began by partnering with a group of creative writing experts to iteratively refine our definition of creative quality. We initially labeled ~1,000 diverse synopses, where three expert writers scored each against the criteria and explained their ratings. Due to the subjectivity of the task, early instance-level agreement was low. To reach a better consensus, we conducted calibration rounds (~50 synopses per round), surfaced disagreements, and evolved our quality scoring guidelines. Key interventions that were found to improve agreement include:</p><ul><li>Using binary scores (instead of 1–4 Likert scores).</li><li>Allowing writers to reference past examples.</li><li>Maintaining a searchable taxonomy of common errors.</li></ul><p><strong>Golden evaluation data. </strong>After eight calibration rounds, writer agreement reached ~80%. To further stabilize labels, we used a model-in-the-loop consensus where:</p><ul><li>Multiple writers score each synopsis.</li><li>An LLM, guided by the rubric, aggregates to a final label.</li><li>Writers review cases with substantial disagreement.</li></ul><p>The result is a golden set of ~600 synopses with binary, criteria-level scores and explanations — <em>our North Star for aligning an LLM judge with expert opinion</em>.</p><h4>Member Implicit Feedback</h4><p>Netflix gauges implicit member feedback on a synopsis with two metrics:</p><ol><li><em>Take Fraction</em>: how often members who see a title’s synopsis choose to start watching it.</li><li><em>Abandonment Rate</em>: how often members start a title but stop watching soon after.</li></ol><p>Higher take fraction indicates more choosing, while lower abandonment suggests authentic, non-misleading presentation. Both of these metrics have been validated via A/B testing to serve as short-term behavioral proxies for long-term member retention. As part of evaluating our system, we also study the ability of LLM-derived quality scores to predict short-term engagement metrics. This step confirms that our scores capture behaviorally meaningful signals and assesses our ability to forecast member response to a given synopsis.</p><h3>Scaling Quality Scoring with LLM-as-a-Judge</h3><p>We begin our experiments by creating simple, per-criteria prompts that:</p><ol><li>Supply criterion-specific show metadata.</li><li>Summarize the relevant quality guidelines.</li><li>Use <a href="https://arxiv.org/abs/2205.11916">zero-shot chain-of-thought prompting</a> to elicit an explanation.</li><li>Request a binary decision for the synopsis.</li></ol><p>Using a single prompt to evaluate all quality criteria is found to overload the LLM and yields poor performance — <em>dedicated judges for each criteria perform better</em>. Because criteria are unique, each task has its own setup, but there are some shared components:</p><ul><li>We use the same LLM for all criteria.</li><li>The judge always outputs an explanation before its final score.</li><li>Final scores are binary.</li></ul><p>Due to our use of binary scoring, judges can be evaluated with simple accuracy metrics over the golden dataset. Next, we summarize the experiments that led to our final system.</p><p><strong>Prompt optimization.</strong> Because LLMs are sensitive to prompt phrasing, we apply <a href="https://arxiv.org/abs/2305.03495">Automatic Prompt Optimization (APO)</a> over a ~300-sample dev set. Scoring guidelines are provided as additional context to the prompt optimizer. After APO, we manually refine candidate prompts with the help of an LLM, yielding initial prompts with accuracies shown below. These prompts work well for some criteria (e.g., precision) but poorly for others (e.g., clarity), highlighting criterion-specific nuances.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*voqhaot2_6G67DSzH0rFzQ.png"></figure><p><strong>Improved reasoning.</strong> Many failures of our initial system arise due to a lack of accurate reasoning through highly-subjective evaluation examples. To improve reasoning accuracy, we leverage two forms of inference-time scaling:</p><ul><li><em>Longer rationales</em>: increase the length of the rationale or explanation generated by the LLM prior to producing a final score.</li><li><em>Consensus scoring</em>: sample several outputs from the LLM and aggregate their scores to produce the final result.</li></ul><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*aEMhMA1OHAGkrvQ0Ljhesg.png"></figure><p><strong>Tiered rationales.</strong> Using tone as an example, we tested whether longer rationales are helpful by defining three rationale length tiers (shown above) and comparing their accuracies. Accuracy rises with longer rationales but returns are diminishing. Medium rationales noticeably outperform short ones, while long rationales offer only a slight additional gain; see below.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*ER-gjiA0IhaYFLr1EchOuA.png"></figure><p>Longer rationales improve performance but degrade human-readability, which is problematic given that explanations are key pieces of evidence for creative experts. As a solution, we adopt tiered rationales: <em>the judge reasons at any length but concisely summarizes its reasoning process prior to the final score. </em>Tiered rationales preserve the benefits of extended reasoning, make outputs easier to inspect, and even benefit scoring accuracy. For example, our tone evaluator improves from 86.55% to 87.85% binary accuracy when using tiered rationales.</p><p><strong>Consensus scoring.</strong> We can also allocate more inference-time compute by sampling multiple outputs per synopsis and aggregating their scores. We aggregate via a rounded average to ensure that the final score remains binary. For tone and clarity criteria with tiered rationales, 5× consensus scoring yields a clear accuracy boost as shown below.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*U7YSyivmXC_M_kw51-tl-g.png"></figure><p>Consensus scoring on the precision evaluator, which uses a vanilla (short) chain-of-thought, yields no benefit. As an explanation, we notice that longer rationales increase variance in scores across multiple outputs, while short rationales yield consistent scores. Consensus may be most useful for evaluators with longer rationales, where it helps to stabilize score variance. When shorter rationales are used, all scores tend to be the same, making consensus less meaningful.</p><p><strong>What about reasoning models?</strong> While our setup elicits reasoning from a standard LLM, we also explored quality scoring with true reasoning models (i.e., models that generate long reasoning trajectories prior to final output). For tone, using a reasoning model with 5× consensus yields improving accuracy with increasing reasoning effort, even outperforming tiered rationales at the highest reasoning effort; see below. However, we skip reasoning models in our final system, as they significantly increase inference costs for only a marginal performance gain.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*D9Ph4ASfjx5pzdfx5EnODQ.png"></figure><p><strong>Agents-as-a-Judge for factuality.</strong> Synopses have four common types of factuality errors:</p><ol><li>Incorrect plot information.</li><li>Incorrect metadata (e.g., genre, location, release date).</li><li>Incorrect on- or off-screen talent.</li><li>Incorrect award information.</li></ol><p>Detecting these factuality errors requires comparing the synopsis to ground-truth context, where necessary context varies per criteria. For example, plot information requires a plot summary or script, while award information needs a list of awards. As we have learned, simplicity drives reliability: <em>too much context or too many criteria harms accuracy</em>. Motivated by this idea, we adopt factuality agents, where each agent evaluates one narrow aspect of factuality.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*oXnk59qsQPASZAgTnPi0HA.png"></figure><p>An agent receives context tailored to one facet of factuality and produces both a rationale and a binary factuality score. The final score of the Agents-as-a-Judge system is the minimum factuality score across agents — <em>any failed aspect yields an overall fail</em>. All rationales are fed to an LLM aggregator to produce a combined rationale to accompany the final score. As shown below, leveraging factuality agents significantly benefits scoring accuracy. Further benefits are achieved by using tiered rationales and consensus scoring within each agent.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*_aZe60Sa1dwWvPKIH7YRgQ.png"></figure><p><strong>Final system. </strong>In summary, our automatic evaluation system uses a combination of standard LLM-as-a-Judge, tiered rationales, consensus scoring, and Agents-as-a-Judge to maximize binary scoring accuracy for each criteria. A summary of the techniques used for each criteria and the associated binary scoring accuracy is provided below.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*qZmJgN46QRLYpEhtlJutEQ.png"></figure><h3>Member Validation of LLM-as-a-Judge</h3><p>Beyond expert agreement, we also study how LLM-as-a-Judge scores relate to member behavior. This analysis serves two goals:</p><ul><li>Further validating LLM-judge accuracy.</li><li>Linking creative quality to member-perceived quality.</li></ul><p>Framed as predictors of member outcomes, LLM judges help us assess how promotional assets affect viewing and determine which creative attributes matter most to members discovering content they enjoy. To perform this analysis, we take advantage of the fact that most shows have multiple, personalized synopses (i.e., a synopsis “suite”). Using this suite, we can measure the causal effect of synopsis selection on metrics like take fraction and abandonment rate.</p><p><strong>Our methodology. </strong>We correlate synopsis performance (take fraction or abandonment) with LLM quality scores. Specifically, within each show s, we relate changes in a synopsis’s LLM score to changes in its performance, normalizing by the show-level standard deviation and clustering standard errors by show; see below.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*4-wIexhYKeqcMyf-a0TqIg.png"></figure><p>β captures the average association between within-show changes in LLM score and changes in performance. While we don’t have clean, experimental variation in LLM scores, this analysis still validates predictive value and practical utility.</p><p><strong>Member-focused results.</strong> We report correlations for individual LLM criteria and a “Weighted Score” that combines all criteria to reduce noise and maximize signal from behavioral data. As shown below, results show promising prediction of take fraction and abandonment. Precision and clarity are especially predictive, and the weighted score provides a statistically useful signal of higher take and lower abandonment. In short, LLM evaluators capture factors that matter to members, making them a valuable tool for monitoring synopsis quality and engagement.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*HnRJo-DOrH9ifgekC5xq6A.png"></figure><h3>Closing Remarks</h3><p>The LLM-as-a-Judge system used to evaluate show synopses at Netflix is the result of extensive experimentation grounded in both creative expertise and member outcomes. Building an automatic evaluation system that works reliably in practice is hard, and the approach we have described reflects countless lessons learned through iteration to improve accuracy and scalability. We have validated the system extensively with human evaluation at both the system and component levels, and we have shown that its outputs correlate with key streaming metrics. As a result, we are confident that it captures the dimensions of synopsis quality that matter most — both creatively and from the member perspective — which has driven its widespread adoption in the Netflix synopsis authoring workflow.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=6269251e6f28" width="1" height="1" alt=""><hr><p><a href="https://medium.com/netflix-techblog/evaluating-netflix-show-synopses-with-llm-as-a-judge-6269251e6f28">Evaluating Netflix Show Synopses with LLM-as-a-Judge</a> was originally published in <a href="https://medium.com/netflix-techblog">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://medium.com/netflix-techblog/evaluating-netflix-show-synopses-with-llm-as-a-judge-6269251e6f28</link>
      <guid>https://medium.com/netflix-techblog/evaluating-netflix-show-synopses-with-llm-as-a-judge-6269251e6f28</guid>
      <pubDate>Fri, 10 Apr 2026 18:26:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Evaluating Netflix Show Synopses with LLM-as-a-Judge]]></title>
      <description><![CDATA[<p>by <a href="https://www.linkedin.com/in/gabrielaalessio/">Gabriela Alessio</a>, <a href="https://www.linkedin.com/in/cameronntaylor/">Cameron Taylor</a>, and <a href="https://www.linkedin.com/in/cwolferesearch/">Cameron R. Wolfe</a></p><h3>Introduction</h3><p>When members log into Netflix, one of the hardest choices is what to watch. The challenge isn’t a lack of options — <em>there are thousands of titles</em> — but finding the most intriguing one is complex and deeply personal. To help, we surface <a href="https://netflixtechblog.com/artwork-personalization-c589f074ad76">personalized promotional assets</a>, especially the show synopsis — <em>a brief description highlighting key plot elements, with cues like genre or talent</em>.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*A2j9Xni86SVnF_VyjFNwfw.png"></figure><p>Strong synopses help members scan, understand, and choose. Poor synopses frustrate, mislead, and drive abandonment. Ensuring high-quality synopses is essential, but scaling quality validation is hard. We host hundreds of thousands of synopses, usually with multiple variants per show. We need to ensure quality at scale so every member gets a consistently great experience every time they read a synopsis. This approach helps us scale high‑quality synopsis coverage for our rapidly expanding catalog, enabling greater speed and coverage without sacrificing quality.</p><p>This report outlines our LLM-based approach for evaluating synopsis quality. Using recent advances in agents, reasoning, and LLM-as-a-Judge, we score four key synopsis quality dimensions, achieving 85%+ agreement with creative writers. Additionally, we show that higher LLM judge quality is correlated with key streaming metrics, <em>allowing us to proactively identify and fix impactful issues weeks or months before a show debuts on Netflix</em>.</p><h3>The Making of a “Good” Synopsis</h3><p>Writing high-quality synopses requires creative expertise. Our expert creative leads are best positioned to craft the creative approaches and define quality standards. However, AI can help us consistently evaluate these expert-driven quality criteria at scale. Synopsis quality at Netflix, which our system aims to predict, is viewed along two dimensions:</p><ol><li><em>Creative Quality</em>: members of our creative writing team assess synopsis quality according to our internal writing guidelines and rubrics.</li><li><em>Member Implicit Feedback</em>: we measure the relative impact of a particular show synopsis on core streaming metrics.</li></ol><p>These two definitions of quality capture distinct and important aspects of quality, one focused upon creative excellence and the other upon utility to members.</p><h4>Creative Quality</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*KzbYGovn903y_ZGysqdnYw.png"></figure><p>For this project, we evaluate synopses against a subset of our creative writing quality rubric — <em>the same criteria to which human writers would adhere</em>. These quality rubrics change over time, and more details on the current quality standards can be found in our <a href="https://partnerhelp.netflixstudios.com/hc/en-us/articles/33377468460435-Editorial-Style-Guide">Editorial Style Guide</a> and <a href="https://partnerhelp.netflixstudios.com/hc/en-us/articles/33377400172563-Technical-Style-Guide">Technical Style Guide</a>. Given Netflix’s distinctive voice and elevated editorial standards, the quality bar is high. Each criterion has extensive guidelines with examples across regions, genres, and synopsis types.</p><p><strong>Human evaluation.</strong> We began by partnering with a group of creative writing experts to iteratively refine our definition of creative quality. We initially labeled ~1,000 diverse synopses, where three expert writers scored each against the criteria and explained their ratings. Due to the subjectivity of the task, early instance-level agreement was low. To reach a better consensus, we conducted calibration rounds (~50 synopses per round), surfaced disagreements, and evolved our quality scoring guidelines. Key interventions that were found to improve agreement include:</p><ul><li>Using binary scores (instead of 1–4 Likert scores).</li><li>Allowing writers to reference past examples.</li><li>Maintaining a searchable taxonomy of common errors.</li></ul><p><strong>Golden evaluation data. </strong>After eight calibration rounds, writer agreement reached ~80%. To further stabilize labels, we used a model-in-the-loop consensus where:</p><ul><li>Multiple writers score each synopsis.</li><li>An LLM, guided by the rubric, aggregates to a final label.</li><li>Writers review cases with substantial disagreement.</li></ul><p>The result is a golden set of ~600 synopses with binary, criteria-level scores and explanations — <em>our North Star for aligning an LLM judge with expert opinion</em>.</p><h4>Member Implicit Feedback</h4><p>Netflix gauges implicit member feedback on a synopsis with two metrics:</p><ol><li><em>Take Fraction</em>: how often members who see a title’s synopsis choose to start watching it.</li><li><em>Abandonment Rate</em>: how often members start a title but stop watching soon after.</li></ol><p>Higher take fraction indicates more choosing, while lower abandonment suggests authentic, non-misleading presentation. Both of these metrics have been validated via A/B testing to serve as short-term behavioral proxies for long-term member retention. As part of evaluating our system, we also study the ability of LLM-derived quality scores to predict short-term engagement metrics. This step confirms that our scores capture behaviorally meaningful signals and assesses our ability to forecast member response to a given synopsis.</p><h3>Scaling Quality Scoring with LLM-as-a-Judge</h3><p>We begin our experiments by creating simple, per-criteria prompts that:</p><ol><li>Supply criterion-specific show metadata.</li><li>Summarize the relevant quality guidelines.</li><li>Use <a href="https://arxiv.org/abs/2205.11916">zero-shot chain-of-thought prompting</a> to elicit an explanation.</li><li>Request a binary decision for the synopsis.</li></ol><p>Using a single prompt to evaluate all quality criteria is found to overload the LLM and yields poor performance — <em>dedicated judges for each criteria perform better</em>. Because criteria are unique, each task has its own setup, but there are some shared components:</p><ul><li>We use the same LLM for all criteria.</li><li>The judge always outputs an explanation before its final score.</li><li>Final scores are binary.</li></ul><p>Due to our use of binary scoring, judges can be evaluated with simple accuracy metrics over the golden dataset. Next, we summarize the experiments that led to our final system.</p><p><strong>Prompt optimization.</strong> Because LLMs are sensitive to prompt phrasing, we apply <a href="https://arxiv.org/abs/2305.03495">Automatic Prompt Optimization (APO)</a> over a ~300-sample dev set. Scoring guidelines are provided as additional context to the prompt optimizer. After APO, we manually refine candidate prompts with the help of an LLM, yielding initial prompts with accuracies shown below. These prompts work well for some criteria (e.g., precision) but poorly for others (e.g., clarity), highlighting criterion-specific nuances.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*voqhaot2_6G67DSzH0rFzQ.png"></figure><p><strong>Improved reasoning.</strong> Many failures of our initial system arise due to a lack of accurate reasoning through highly-subjective evaluation examples. To improve reasoning accuracy, we leverage two forms of inference-time scaling:</p><ul><li><em>Longer rationales</em>: increase the length of the rationale or explanation generated by the LLM prior to producing a final score.</li><li><em>Consensus scoring</em>: sample several outputs from the LLM and aggregate their scores to produce the final result.</li></ul><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*aEMhMA1OHAGkrvQ0Ljhesg.png"></figure><p><strong>Tiered rationales.</strong> Using tone as an example, we tested whether longer rationales are helpful by defining three rationale length tiers (shown above) and comparing their accuracies. Accuracy rises with longer rationales but returns are diminishing. Medium rationales noticeably outperform short ones, while long rationales offer only a slight additional gain; see below.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*ER-gjiA0IhaYFLr1EchOuA.png"></figure><p>Longer rationales improve performance but degrade human-readability, which is problematic given that explanations are key pieces of evidence for creative experts. As a solution, we adopt tiered rationales: <em>the judge reasons at any length but concisely summarizes its reasoning process prior to the final score. </em>Tiered rationales preserve the benefits of extended reasoning, make outputs easier to inspect, and even benefit scoring accuracy. For example, our tone evaluator improves from 86.55% to 87.85% binary accuracy when using tiered rationales.</p><p><strong>Consensus scoring.</strong> We can also allocate more inference-time compute by sampling multiple outputs per synopsis and aggregating their scores. We aggregate via a rounded average to ensure that the final score remains binary. For tone and clarity criteria with tiered rationales, 5× consensus scoring yields a clear accuracy boost as shown below.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*U7YSyivmXC_M_kw51-tl-g.png"></figure><p>Consensus scoring on the precision evaluator, which uses a vanilla (short) chain-of-thought, yields no benefit. As an explanation, we notice that longer rationales increase variance in scores across multiple outputs, while short rationales yield consistent scores. Consensus may be most useful for evaluators with longer rationales, where it helps to stabilize score variance. When shorter rationales are used, all scores tend to be the same, making consensus less meaningful.</p><p><strong>What about reasoning models?</strong> While our setup elicits reasoning from a standard LLM, we also explored quality scoring with true reasoning models (i.e., models that generate long reasoning trajectories prior to final output). For tone, using a reasoning model with 5× consensus yields improving accuracy with increasing reasoning effort, even outperforming tiered rationales at the highest reasoning effort; see below. However, we skip reasoning models in our final system, as they significantly increase inference costs for only a marginal performance gain.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*D9Ph4ASfjx5pzdfx5EnODQ.png"></figure><p><strong>Agents-as-a-Judge for factuality.</strong> Synopses have four common types of factuality errors:</p><ol><li>Incorrect plot information.</li><li>Incorrect metadata (e.g., genre, location, release date).</li><li>Incorrect on- or off-screen talent.</li><li>Incorrect award information.</li></ol><p>Detecting these factuality errors requires comparing the synopsis to ground-truth context, where necessary context varies per criteria. For example, plot information requires a plot summary or script, while award information needs a list of awards. As we have learned, simplicity drives reliability: <em>too much context or too many criteria harms accuracy</em>. Motivated by this idea, we adopt factuality agents, where each agent evaluates one narrow aspect of factuality.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*oXnk59qsQPASZAgTnPi0HA.png"></figure><p>An agent receives context tailored to one facet of factuality and produces both a rationale and a binary factuality score. The final score of the Agents-as-a-Judge system is the minimum factuality score across agents — <em>any failed aspect yields an overall fail</em>. All rationales are fed to an LLM aggregator to produce a combined rationale to accompany the final score. As shown below, leveraging factuality agents significantly benefits scoring accuracy. Further benefits are achieved by using tiered rationales and consensus scoring within each agent.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*_aZe60Sa1dwWvPKIH7YRgQ.png"></figure><p><strong>Final system. </strong>In summary, our automatic evaluation system uses a combination of standard LLM-as-a-Judge, tiered rationales, consensus scoring, and Agents-as-a-Judge to maximize binary scoring accuracy for each criteria. A summary of the techniques used for each criteria and the associated binary scoring accuracy is provided below.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*qZmJgN46QRLYpEhtlJutEQ.png"></figure><h3>Member Validation of LLM-as-a-Judge</h3><p>Beyond expert agreement, we also study how LLM-as-a-Judge scores relate to member behavior. This analysis serves two goals:</p><ul><li>Further validating LLM-judge accuracy.</li><li>Linking creative quality to member-perceived quality.</li></ul><p>Framed as predictors of member outcomes, LLM judges help us assess how promotional assets affect viewing and determine which creative attributes matter most to members discovering content they enjoy. To perform this analysis, we take advantage of the fact that most shows have multiple, personalized synopses (i.e., a synopsis “suite”). Using this suite, we can measure the causal effect of synopsis selection on metrics like take fraction and abandonment rate.</p><p><strong>Our methodology. </strong>We correlate synopsis performance (take fraction or abandonment) with LLM quality scores. Specifically, within each show s, we relate changes in a synopsis’s LLM score to changes in its performance, normalizing by the show-level standard deviation and clustering standard errors by show; see below.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*4-wIexhYKeqcMyf-a0TqIg.png"></figure><p>β captures the average association between within-show changes in LLM score and changes in performance. While we don’t have clean, experimental variation in LLM scores, this analysis still validates predictive value and practical utility.</p><p><strong>Member-focused results.</strong> We report correlations for individual LLM criteria and a “Weighted Score” that combines all criteria to reduce noise and maximize signal from behavioral data. As shown below, results show promising prediction of take fraction and abandonment. Precision and clarity are especially predictive, and the weighted score provides a statistically useful signal of higher take and lower abandonment. In short, LLM evaluators capture factors that matter to members, making them a valuable tool for monitoring synopsis quality and engagement.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*HnRJo-DOrH9ifgekC5xq6A.png"></figure><h3>Closing Remarks</h3><p>The LLM-as-a-Judge system used to evaluate show synopses at Netflix is the result of extensive experimentation grounded in both creative expertise and member outcomes. Building an automatic evaluation system that works reliably in practice is hard, and the approach we have described reflects countless lessons learned through iteration to improve accuracy and scalability. We have validated the system extensively with human evaluation at both the system and component levels, and we have shown that its outputs correlate with key streaming metrics. As a result, we are confident that it captures the dimensions of synopsis quality that matter most — both creatively and from the member perspective — which has driven its widespread adoption in the Netflix synopsis authoring workflow.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=6269251e6f28" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/evaluating-netflix-show-synopses-with-llm-as-a-judge-6269251e6f28">Evaluating Netflix Show Synopses with LLM-as-a-Judge</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/evaluating-netflix-show-synopses-with-llm-as-a-judge-6269251e6f28</link>
      <guid>https://netflixtechblog.com/evaluating-netflix-show-synopses-with-llm-as-a-judge-6269251e6f28</guid>
      <pubDate>Fri, 10 Apr 2026 18:26:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Stop Answering the Same Question Twice: Interval-Aware Caching for Druid at Netflix Scale]]></title>
      <description><![CDATA[<p><em>By </em><a href="https://www.linkedin.com/in/sykesb/"><em>Ben Sykes</em></a></p><p>In a <a href="https://netflixtechblog.com/how-netflix-uses-druid-for-real-time-insights-to-ensure-a-high-quality-experience-19e1e8568d06">previous post</a>, we described how Netflix uses Apache Druid to ingest millions of events per second and query trillions of rows, providing the real-time insights needed to ensure a high-quality experience for our members. Since that post, our scale has grown considerably.</p><p>With our database holding over 10 trillion rows and regularly ingesting up to 15 million events per second, the value of our real-time data is undeniable. But this massive scale introduced a new challenge: queries. The live show monitoring, dashboards, automated alerting, canary analysis, and A/B test monitoring that are built on top of Druid became so heavily relied upon that the repetitive query load started to become a scaling concern in itself.</p><p>This post describes an experimental caching layer we built to address this problem, and the trade-offs we chose to accept.</p><h3><strong>The Problem</strong></h3><p>Our internal dashboards are heavily used for real-time monitoring, especially during high-profile live shows or global launches. A typical dashboard has 10+ charts, each triggering one or more Druid queries; one popular dashboard with 26 charts and stats generates 64 queries per load. When dozens of engineers view the same dashboards and metrics for the same event, the query volume quickly becomes unmanageable.</p><p>Take the popular dashboard above: 64 queries per load, refreshing every 10 seconds, viewed by 30 people. That’s 192 queries per second from one dashboard, mostly for nearly identical data. We still need Druid capacity for automated alerting, canary analysis, and ad-hoc queries. And because these dashboards request a rolling last-few-hours window, each refresh changes slightly as the time range advances.</p><p>Druid’s built-in caches are effective. Both the full-result cache and the per-segment cache. But neither is designed to handle the continuous, overlapping time-window shifts inherent to rolling-window dashboards. The full-result cache misses for two reasons.</p><ul><li>If the time window shifts even slightly, the query is different, so it’s a cache miss.</li><li>Druid deliberately refuses to cache results that involve realtime segments (those still being indexed), because it values deterministic, stable cache results and query correctness over a higher cache hit rate.</li></ul><p>The per-segment cache does help avoid redundant scans on historical nodes, but we still need to collect those cached segment results from each data node and merge them in the brokers with data from the realtime nodes for every query.</p><p>During major shows, rolling-window dashboards can generate a flood of near-duplicate queries that Druid’s caches mostly miss, creating heavy redundant load. At our scale, solving this by simply adding more hardware is prohibitively expensive.</p><p>We needed a smarter approach.</p><h3>The Insight</h3><p>When a dashboard requests the last 3 hours of data, the vast majority of that data, everything except the most recent few minutes, is already settled. The data from 2 hours ago won’t change.</p><p>What if we could remember the older portions of the result and only ask Druid for the part that’s actually new?</p><p>This is the core idea behind a new caching service that understands the structure of Druid queries and serves previously-seen results from cache while fetching only the freshest portion from Druid.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*8bFAtFl8Z5pwEyoOgK54ag.png"></figure><h3><strong>A Deliberate Trade-Off</strong></h3><p>Before diving into the implementation, it’s worth being explicit about the trade-off we’re making. Caching query results introduces some staleness, specifically, up to 5 seconds for the newest data. This is acceptable for most of our operational dashboards, which refresh every 10 to 30 seconds. In practice, many of our queries already set an end time of now-1m or now-5s to avoid the “flappy tail” that can occur with currently-arriving data.</p><p>Since our end-to-end data pipeline latency is typically under 5 seconds at P90, a 5-second cache TTL on the freshest data introduces negligible additional staleness on top of what’s already inherent in the system. We decided it was better to accept this small amount of staleness in exchange for significantly lower query load on Druid. But a 5s cache on its own is not very useful.</p><h3>Exponential TTLs</h3><p>Not all data points are equally trustworthy. In real-time analytics, there’s a well-known late-arriving data problem. Events can arrive out of order or be delayed in the ingestion pipeline. A data point from 30 seconds ago might still change as late-arriving events trickle in. A data point from 30 minutes ago is almost certainly final.</p><p>We use this observation to set cache TTLs that increase exponentially with the age of the data. Data less than 2 minutes old gets a minimum TTL of 5 seconds. After that, the TTL doubles for each additional minute of age: 10 seconds at 2 minutes old, 20 seconds at 3 minutes, 40 seconds at 4 minutes, and so on, up to a maximum TTL of 1 hour.</p><p>The effect is that fresh data cycles through the cache rapidly, so any corrections from late-arriving events in the most recent couple of minutes are picked up quickly. Older data lingers much longer, because our confidence in its accuracy grows with time.</p><p>For a 3-hour rolling window, the exponential TTL ensures the vast majority of the query is served from the cache, leaving Druid to only scan the most recent, unsettled data.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*oyLRzcOO7otVniQM6khD6g.png"></figure><h3><strong>Bucketing</strong></h3><p>If we were to use a single-level cache key for the query and interval, similar to Druid’s existing result-level cache, we wouldn’t be able to extract only the relevant time range from cached results. A shifted window means a different key, which means a cache miss.</p><p>Instead, we use a map-of-maps. The top-level key is the query hash without the time interval; the inner keys are timestamps bucketed to the query granularity (or 1 minute, whichever is larger) and encoded as big-endian bytes so lexicographic order matches time. This enables efficient range scans; fetching all cached buckets between times A and B for a query hash. A 3-hour query at 1-minute granularity becomes 180 independent cached buckets, each with its own TTL; when the window shifts (e.g., 30 seconds later), we reuse most buckets from cache and only query Druid for the new data.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*6DL-4Lpu_CqfZBVmF2rNvg.png"></figure><h3><strong>How It Works</strong></h3><p>Today, the cache runs as an external service integrated transparently by intercepting requests at the Druid Router and redirecting them to the cache. If the cache fully satisfies a request, it returns the result; otherwise it shrinks the time interval to the uncached portion and calls back into the Router, bypassing the redirect to query Druid normally. Non-cached requests (e.g., metadata queries or queries without time group-bys) pass straight through to Druid unchanged.</p><p>This intercepting proxy design allows us to enable or disable caching without any client changes and is a key to its adoption. We see this setup as temporary while we work out a way to better integrate this capability into Druid more natively.</p><p>When a cacheable query arrives, those that are grouping-by time (timeseries, groupBy), the cache performs the following steps.</p><p><strong>Parsing and Hashing.</strong> We parse each incoming query to extract the time interval, granularity, and structure, then compute a SHA-256 hash of the query with the time interval and parts of the context removed. That hash is the cache key: it encodes <em>what</em> is being asked (datasource, filters, aggregations, granularity) but not <em>when</em>, so the same logical query over different overlapping time windows maps to the same cache entry. There are some context properties that can alter the response structure or contents, so these are included in the cache-key.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*QaeamWjRYGEQkW6t-XePJw.png"></figure><p><strong>Cache Lookup.</strong> Using the cache key, we fetch cached points within the requested range, but only if they’re contiguous from the start. Because bucket TTLs can expire unevenly, gaps can appear; when we hit a gap, we stop and fetch all newer data from Druid. This guarantees a complete, unbroken result set while sending at most one Druid query, rather than “filling gaps” with multiple small, fragmented queries that would increase Druid load.</p><p><strong>Fetching the Missing Tail.</strong> On a partial cache hit (e.g., 2h 50m of a 3h window), we rebuild the query with a narrowed interval for the missing 10 minutes and send only that to Druid. Since Druid then scans just the recent segments for a small time range, the query is usually faster and cheaper than the original.</p><p><strong>Combining.</strong> The cached data and fresh data are concatenated, sorted by timestamp, and returned to the client. From the client’s perspective, the response looks identical to what Druid would have returned, same JSON format, same fields.</p><p><strong>Asynchronous Caching.</strong> The fresh data from Druid is parsed into individual time-granularity buckets and written back to the cache asynchronously, so we don’t add latency to the response path.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*k3uoqCFgzbRZEuiNo_Rlzw.png"></figure><h3>Negative Caching</h3><p>Some metrics are sparse. Certain time buckets may genuinely have no data. Without special handling, the cache would treat these empty buckets as gaps and re-query Druid for them every time.</p><p>We handle this by caching empty sentinel values for time buckets where Druid returned no data. Our gap-detection logic recognizes these empty entries as valid cached data rather than missing data, preventing needless re-queries for naturally sparse metrics.</p><p>However, we’re careful not to negative-cache trailing empty buckets. If a query returns data up to minute 45 and nothing after, we only cache empty entries for gaps <em>between</em> data points, not after the last one. This avoids incorrectly caching “no data” for time periods where events simply haven’t arrived yet, which would exacerbate the chart delays of late arriving data.</p><h3>The Storage Layer</h3><p>For the backing store, we use Netflix’s <a href="https://netflixtechblog.com/introducing-netflixs-key-value-data-abstraction-layer-1ea8a0a11b30">Key-Value Data Abstraction Layer (KVDAL)</a>, backed by Cassandra. KVDAL provides a two-level map abstraction, a natural fit for our needs. The outer key is the query hash, and the inner keys are timestamps. Crucially, KVDAL supports independent TTLs on each inner key-value pair, eliminating the need for us to manage cache eviction manually.</p><p>This two-level structure gives us efficient range queries over the inner keys, which is exactly what we need for partial cache lookups: “give me all cached buckets between time A and time B for query hash X.”</p><h3><strong>Results</strong></h3><p>The biggest win is during high-volume events (e.g., live shows): when many users view the same dashboards, the cache serves most identical queries as full hits, so the query rate reaching Druid is essentially the same with 1 viewer or 100. The scaling bottleneck moves from Druid’s query capacity to the much cheaper-to-scale cache, and with ~5.5 ms P90 cache responses, dashboards load faster for everyone.</p><p>On a typical day, 82% of real user queries get at least a partial cache hit, and 84% of result data is served from cache. As a result, the queries that reach Druid scan much narrower time ranges, touching fewer segments and processing less data, freeing Druid to focus on aggregating the newest data instead of repeatedly re-querying historical segments.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*S2tG3E2KlGIujSaEg4oetA.png"></figure><p>An experiment validated this, showing about a 33% drop in queries to Druid and a 66% improvement in overall P90 query times. It also cut result bytes and segments queried, and in some cases, enabling the cache reduced result bytes by more than 14x. Caveat: the size of these gains depends heavily on how similar and repetitive the query workload is.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/592/1*B-j0CnKKr5caFQRI8zVZ1A.png"></figure><h3>Looking Ahead</h3><p>This caching layer is still experimental, but results are promising and we’re exploring next steps. We’ve added partial support for templated SQL so dashboard tools can benefit without writing native Druid queries.</p><p>Longer term, we’d like interval-aware caching to be built into Druid: an external proxy adds infrastructure to manage, extra network hops, and workarounds (like SQL templating) to extract intervals. Implemented inside Druid, it could be more efficient, with direct access to the query planner and segment metadata, and benefit the broader community without custom infrastructure. We’d likely ship it as an opt-in, configurable, result-level cache in the Brokers, with metrics to tune TTLs and measure effectiveness. Please leave a comment if you have a use-case that could benefit from this feature.</p><p>More broadly, this strategy, splitting time-series results into independently cached, granularity-aligned buckets with age-based exponential TTLs, isn’t Druid-specific and could apply to any time-series database with frequent overlapping-window queries.</p><h3>Summary</h3><p>As more Netflix teams rely on real-time analytics, query volume grows too. Dashboards are essential at our scale, but their popularity can become a scaling bottleneck. By inserting an intelligent cache between dashboards and Druid, one that understands query structure, breaks results into granularity-aligned buckets, and trades a small amount of staleness for much lower Druid load, we’ve increased query capacity without scaling infrastructure proportionally, and hope to deliver these benefits to the Druid community soon as a built-in Druid feature.</p><p>Sometimes the best way to handle a flood of queries is to stop answering the same question twice.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=22fadc9b840e" width="1" height="1" alt=""><hr><p><a href="https://medium.com/netflix-techblog/stop-answering-the-same-question-twice-interval-aware-caching-for-druid-at-netflix-scale-22fadc9b840e">Stop Answering the Same Question Twice: Interval-Aware Caching for Druid at Netflix Scale</a> was originally published in <a href="https://medium.com/netflix-techblog">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://medium.com/netflix-techblog/stop-answering-the-same-question-twice-interval-aware-caching-for-druid-at-netflix-scale-22fadc9b840e</link>
      <guid>https://medium.com/netflix-techblog/stop-answering-the-same-question-twice-interval-aware-caching-for-druid-at-netflix-scale-22fadc9b840e</guid>
      <pubDate>Tue, 07 Apr 2026 00:15:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Stop Answering the Same Question Twice: Interval-Aware Caching for Druid at Netflix Scale]]></title>
      <description><![CDATA[<p><em>By </em><a href="https://www.linkedin.com/in/sykesb/"><em>Ben Sykes</em></a></p><p>In a <a href="https://netflixtechblog.com/how-netflix-uses-druid-for-real-time-insights-to-ensure-a-high-quality-experience-19e1e8568d06">previous post</a>, we described how Netflix uses Apache Druid to ingest millions of events per second and query trillions of rows, providing the real-time insights needed to ensure a high-quality experience for our members. Since that post, our scale has grown considerably.</p><p>With our database holding over 10 trillion rows and regularly ingesting up to 15 million events per second, the value of our real-time data is undeniable. But this massive scale introduced a new challenge: queries. The live show monitoring, dashboards, automated alerting, canary analysis, and A/B test monitoring that are built on top of Druid became so heavily relied upon that the repetitive query load started to become a scaling concern in itself.</p><p>This post describes an experimental caching layer we built to address this problem, and the trade-offs we chose to accept.</p><h3><strong>The Problem</strong></h3><p>Our internal dashboards are heavily used for real-time monitoring, especially during high-profile live shows or global launches. A typical dashboard has 10+ charts, each triggering one or more Druid queries; one popular dashboard with 26 charts and stats generates 64 queries per load. When dozens of engineers view the same dashboards and metrics for the same event, the query volume quickly becomes unmanageable.</p><p>Take the popular dashboard above: 64 queries per load, refreshing every 10 seconds, viewed by 30 people. That’s 192 queries per second from one dashboard, mostly for nearly identical data. We still need Druid capacity for automated alerting, canary analysis, and ad-hoc queries. And because these dashboards request a rolling last-few-hours window, each refresh changes slightly as the time range advances.</p><p>Druid’s built-in caches are effective. Both the full-result cache and the per-segment cache. But neither is designed to handle the continuous, overlapping time-window shifts inherent to rolling-window dashboards. The full-result cache misses for two reasons.</p><ul><li>If the time window shifts even slightly, the query is different, so it’s a cache miss.</li><li>Druid deliberately refuses to cache results that involve realtime segments (those still being indexed), because it values deterministic, stable cache results and query correctness over a higher cache hit rate.</li></ul><p>The per-segment cache does help avoid redundant scans on historical nodes, but we still need to collect those cached segment results from each data node and merge them in the brokers with data from the realtime nodes for every query.</p><p>During major shows, rolling-window dashboards can generate a flood of near-duplicate queries that Druid’s caches mostly miss, creating heavy redundant load. At our scale, solving this by simply adding more hardware is prohibitively expensive.</p><p>We needed a smarter approach.</p><h3>The Insight</h3><p>When a dashboard requests the last 3 hours of data, the vast majority of that data, everything except the most recent few minutes, is already settled. The data from 2 hours ago won’t change.</p><p>What if we could remember the older portions of the result and only ask Druid for the part that’s actually new?</p><p>This is the core idea behind a new caching service that understands the structure of Druid queries and serves previously-seen results from cache while fetching only the freshest portion from Druid.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*8bFAtFl8Z5pwEyoOgK54ag.png"></figure><h3><strong>A Deliberate Trade-Off</strong></h3><p>Before diving into the implementation, it’s worth being explicit about the trade-off we’re making. Caching query results introduces some staleness, specifically, up to 5 seconds for the newest data. This is acceptable for most of our operational dashboards, which refresh every 10 to 30 seconds. In practice, many of our queries already set an end time of now-1m or now-5s to avoid the “flappy tail” that can occur with currently-arriving data.</p><p>Since our end-to-end data pipeline latency is typically under 5 seconds at P90, a 5-second cache TTL on the freshest data introduces negligible additional staleness on top of what’s already inherent in the system. We decided it was better to accept this small amount of staleness in exchange for significantly lower query load on Druid. But a 5s cache on its own is not very useful.</p><h3>Exponential TTLs</h3><p>Not all data points are equally trustworthy. In real-time analytics, there’s a well-known late-arriving data problem. Events can arrive out of order or be delayed in the ingestion pipeline. A data point from 30 seconds ago might still change as late-arriving events trickle in. A data point from 30 minutes ago is almost certainly final.</p><p>We use this observation to set cache TTLs that increase exponentially with the age of the data. Data less than 2 minutes old gets a minimum TTL of 5 seconds. After that, the TTL doubles for each additional minute of age: 10 seconds at 2 minutes old, 20 seconds at 3 minutes, 40 seconds at 4 minutes, and so on, up to a maximum TTL of 1 hour.</p><p>The effect is that fresh data cycles through the cache rapidly, so any corrections from late-arriving events in the most recent couple of minutes are picked up quickly. Older data lingers much longer, because our confidence in its accuracy grows with time.</p><p>For a 3-hour rolling window, the exponential TTL ensures the vast majority of the query is served from the cache, leaving Druid to only scan the most recent, unsettled data.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*oyLRzcOO7otVniQM6khD6g.png"></figure><h3><strong>Bucketing</strong></h3><p>If we were to use a single-level cache key for the query and interval, similar to Druid’s existing result-level cache, we wouldn’t be able to extract only the relevant time range from cached results. A shifted window means a different key, which means a cache miss.</p><p>Instead, we use a map-of-maps. The top-level key is the query hash without the time interval; the inner keys are timestamps bucketed to the query granularity (or 1 minute, whichever is larger) and encoded as big-endian bytes so lexicographic order matches time. This enables efficient range scans; fetching all cached buckets between times A and B for a query hash. A 3-hour query at 1-minute granularity becomes 180 independent cached buckets, each with its own TTL; when the window shifts (e.g., 30 seconds later), we reuse most buckets from cache and only query Druid for the new data.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*6DL-4Lpu_CqfZBVmF2rNvg.png"></figure><h3><strong>How It Works</strong></h3><p>Today, the cache runs as an external service integrated transparently by intercepting requests at the Druid Router and redirecting them to the cache. If the cache fully satisfies a request, it returns the result; otherwise it shrinks the time interval to the uncached portion and calls back into the Router, bypassing the redirect to query Druid normally. Non-cached requests (e.g., metadata queries or queries without time group-bys) pass straight through to Druid unchanged.</p><p>This intercepting proxy design allows us to enable or disable caching without any client changes and is a key to its adoption. We see this setup as temporary while we work out a way to better integrate this capability into Druid more natively.</p><p>When a cacheable query arrives, those that are grouping-by time (timeseries, groupBy), the cache performs the following steps.</p><p><strong>Parsing and Hashing.</strong> We parse each incoming query to extract the time interval, granularity, and structure, then compute a SHA-256 hash of the query with the time interval and parts of the context removed. That hash is the cache key: it encodes <em>what</em> is being asked (datasource, filters, aggregations, granularity) but not <em>when</em>, so the same logical query over different overlapping time windows maps to the same cache entry. There are some context properties that can alter the response structure or contents, so these are included in the cache-key.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*QaeamWjRYGEQkW6t-XePJw.png"></figure><p><strong>Cache Lookup.</strong> Using the cache key, we fetch cached points within the requested range, but only if they’re contiguous from the start. Because bucket TTLs can expire unevenly, gaps can appear; when we hit a gap, we stop and fetch all newer data from Druid. This guarantees a complete, unbroken result set while sending at most one Druid query, rather than “filling gaps” with multiple small, fragmented queries that would increase Druid load.</p><p><strong>Fetching the Missing Tail.</strong> On a partial cache hit (e.g., 2h 50m of a 3h window), we rebuild the query with a narrowed interval for the missing 10 minutes and send only that to Druid. Since Druid then scans just the recent segments for a small time range, the query is usually faster and cheaper than the original.</p><p><strong>Combining.</strong> The cached data and fresh data are concatenated, sorted by timestamp, and returned to the client. From the client’s perspective, the response looks identical to what Druid would have returned, same JSON format, same fields.</p><p><strong>Asynchronous Caching.</strong> The fresh data from Druid is parsed into individual time-granularity buckets and written back to the cache asynchronously, so we don’t add latency to the response path.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*k3uoqCFgzbRZEuiNo_Rlzw.png"></figure><h3>Negative Caching</h3><p>Some metrics are sparse. Certain time buckets may genuinely have no data. Without special handling, the cache would treat these empty buckets as gaps and re-query Druid for them every time.</p><p>We handle this by caching empty sentinel values for time buckets where Druid returned no data. Our gap-detection logic recognizes these empty entries as valid cached data rather than missing data, preventing needless re-queries for naturally sparse metrics.</p><p>However, we’re careful not to negative-cache trailing empty buckets. If a query returns data up to minute 45 and nothing after, we only cache empty entries for gaps <em>between</em> data points, not after the last one. This avoids incorrectly caching “no data” for time periods where events simply haven’t arrived yet, which would exacerbate the chart delays of late arriving data.</p><h3>The Storage Layer</h3><p>For the backing store, we use Netflix’s <a href="https://netflixtechblog.com/introducing-netflixs-key-value-data-abstraction-layer-1ea8a0a11b30">Key-Value Data Abstraction Layer (KVDAL)</a>, backed by Cassandra. KVDAL provides a two-level map abstraction, a natural fit for our needs. The outer key is the query hash, and the inner keys are timestamps. Crucially, KVDAL supports independent TTLs on each inner key-value pair, eliminating the need for us to manage cache eviction manually.</p><p>This two-level structure gives us efficient range queries over the inner keys, which is exactly what we need for partial cache lookups: “give me all cached buckets between time A and time B for query hash X.”</p><h3><strong>Results</strong></h3><p>The biggest win is during high-volume events (e.g., live shows): when many users view the same dashboards, the cache serves most identical queries as full hits, so the query rate reaching Druid is essentially the same with 1 viewer or 100. The scaling bottleneck moves from Druid’s query capacity to the much cheaper-to-scale cache, and with ~5.5 ms P90 cache responses, dashboards load faster for everyone.</p><p>On a typical day, 82% of real user queries get at least a partial cache hit, and 84% of result data is served from cache. As a result, the queries that reach Druid scan much narrower time ranges, touching fewer segments and processing less data, freeing Druid to focus on aggregating the newest data instead of repeatedly re-querying historical segments.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*S2tG3E2KlGIujSaEg4oetA.png"></figure><p>An experiment validated this, showing about a 33% drop in queries to Druid and a 66% improvement in overall P90 query times. It also cut result bytes and segments queried, and in some cases, enabling the cache reduced result bytes by more than 14x. Caveat: the size of these gains depends heavily on how similar and repetitive the query workload is.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/592/1*B-j0CnKKr5caFQRI8zVZ1A.png"></figure><h3>Looking Ahead</h3><p>This caching layer is still experimental, but results are promising and we’re exploring next steps. We’ve added partial support for templated SQL so dashboard tools can benefit without writing native Druid queries.</p><p>Longer term, we’d like interval-aware caching to be built into Druid: an external proxy adds infrastructure to manage, extra network hops, and workarounds (like SQL templating) to extract intervals. Implemented inside Druid, it could be more efficient, with direct access to the query planner and segment metadata, and benefit the broader community without custom infrastructure. We’d likely ship it as an opt-in, configurable, result-level cache in the Brokers, with metrics to tune TTLs and measure effectiveness. Please leave a comment if you have a use-case that could benefit from this feature.</p><p>More broadly, this strategy, splitting time-series results into independently cached, granularity-aligned buckets with age-based exponential TTLs, isn’t Druid-specific and could apply to any time-series database with frequent overlapping-window queries.</p><h3>Summary</h3><p>As more Netflix teams rely on real-time analytics, query volume grows too. Dashboards are essential at our scale, but their popularity can become a scaling bottleneck. By inserting an intelligent cache between dashboards and Druid, one that understands query structure, breaks results into granularity-aligned buckets, and trades a small amount of staleness for much lower Druid load, we’ve increased query capacity without scaling infrastructure proportionally, and hope to deliver these benefits to the Druid community soon as a built-in Druid feature.</p><p>Sometimes the best way to handle a flood of queries is to stop answering the same question twice.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=22fadc9b840e" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/stop-answering-the-same-question-twice-interval-aware-caching-for-druid-at-netflix-scale-22fadc9b840e">Stop Answering the Same Question Twice: Interval-Aware Caching for Druid at Netflix Scale</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/stop-answering-the-same-question-twice-interval-aware-caching-for-druid-at-netflix-scale-22fadc9b840e</link>
      <guid>https://netflixtechblog.com/stop-answering-the-same-question-twice-interval-aware-caching-for-druid-at-netflix-scale-22fadc9b840e</guid>
      <pubDate>Tue, 07 Apr 2026 00:15:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Powering Multimodal Intelligence for Video Search]]></title>
      <description><![CDATA[<h3>Synchronizing the Senses: Powering Multimodal Intelligence for Video Search</h3><p><em>By</em>: <a href="https://www.linkedin.com/in/meenakshijindal/">Meenakshi Jindal</a> and <a href="https://www.linkedin.com/in/~munya/">Munya Marazanye</a></p><p>Today’s filmmakers capture more footage than ever to maximize their creative options, often generating hundreds, if not thousands, of hours of raw material per season or franchise. Extracting the vital moments needed to craft compelling storylines from this sheer volume of media is a notoriously slow and punishing process. When editorial teams cannot surface these key moments quickly, creative momentum stalls and severe fatigue sets in.</p><p>Meanwhile, the broader search landscape is undergoing a profound transformation. We are moving beyond simple keyword matching toward AI-driven systems capable of understanding deep context and intent. Yet, while these advances have revolutionized text and image retrieval, searching through video, the richest medium for storytelling, remains a daunting “needle in a haystack” challenge.</p><p>The solution to this bottleneck cannot rely on a single algorithm. Instead, it demands orchestrating an expansive ensemble of specialized models: tools that identify specific characters, map visual environments, and parse nuanced dialogue. The ultimate challenge lies in unifying these heterogeneous signals, textual labels, and high-dimensional vectors into a cohesive, real-time intelligence. One that cuts through the noise and responds to complex queries at the speed of thought, truly empowering the creative process.</p><h3>Why Video Search is Deceptively Complex</h3><p>Since video is a multi-layered medium, building an effective search engine required us to overcome significant technical bottlenecks. Multi-modal search is exponentially more complex than traditional indexing: it demands the unification of outputs from multiple specialized models, each analyzing a different facet of the content to generate its own distinct metadata. The ultimate challenge lies in harmonizing these heterogeneous data streams to support rich, multi-dimensional queries in real time.</p><ol><li><strong>Unifying the Timeline</strong></li></ol><p>To ensure critical moments aren’t lost across scene boundaries, each model segments the video into overlapping intervals. The resulting metadata varies wildly, ranging from discrete text-based object labels to dense vector embeddings. Synchronizing these disjointed, multi-modal timelines into a unified chronological map presents a massive computational hurdle.</p><p><strong>2. Processing at Scale</strong></p><p>A standard 2,000-hour production archive can contain over 216 million frames. When processed through an ensemble of specialized models, this baseline explodes into billions of multi-layered data points. Storing, aligning, and intersecting this staggering volume of records while maintaining sub-second query latency far exceeds the capabilities of traditional database architectures.</p><p><strong>3. Surfacing the Best Moments</strong></p><p>Surface-level mathematical similarity is not enough to identify the most relevant clip. Because continuous shots naturally generate thousands of visually redundant candidates, the system must dynamically cluster and deduplicate results to surface the singular best match for a given scene. To achieve this, effective ranking relies on a sophisticated hybrid scoring engine that weighs symbolic text matches against semantic vector embeddings, ensuring both precision and interpretability.</p><p><strong>4. Zero-Friction Search</strong></p><p>For filmmakers, search is a stream-of-consciousness process, and a ten-second delay can disrupt the creative flow. Because sequential scanning of raw footage is fundamentally unscalable, our architecture is built to navigate and correlate billions of vectors and metadata records efficiently, operating at the speed of thought.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*wVRT7XY4C9bzNHE-lk-ViA.png"><figcaption><em>Figure 1: Unified Multimodal Result Processing</em></figcaption></figure><h3>The Ingestion and Fusion Pipeline</h3><p>To ensure system resilience and scalability, the transition from raw model output to searchable intelligence follows a decoupled, three-stage process:</p><h4>1. Transactional Persistence</h4><p>Raw annotations are ingested via high-availability <a href="https://netflixtechblog.com/data-ingestion-pipeline-with-operation-management-3c5c638740a8">pipelines</a> and stored in our <a href="https://netflixtechblog.com/scalable-annotation-service-marken-f5ba9266d428">annotation service</a>, which leverages <a href="https://cassandra.apache.org/_/index.html">Apache Cassandra</a> for distributed storage. This stage strictly prioritizes data integrity and high-speed write throughput, guaranteeing that every piece of model output is safely captured.</p><pre>{<br>  "type": "SCENE_SEARCH",<br>  "time_range": {<br>    "start_time_ns": 4000000000,<br>    "end_time_ns": 9000000000<br>  },<br>  "embedding_vector": [<br>    -0.036, -0.33, -0.29 ...<br>  ],<br>  "label": "kitchen",<br>  "confidence_score": 0.72<br>}</pre><p><em>Figure 2: Sample Scene Search Model Annotation Output</em></p><h4>2. Offline Data Fusion</h4><p>Once the annotation service securely persists the raw data, the system publishes an event via <a href="https://kafka.apache.org/">Apache Kafka</a> to trigger an asynchronous processing job. Serving as the architecture’s central logic layer, this offline pipeline handles the heavy computational lifting out-of-band. It performs precise temporal intersections, fusing overlapping annotations from disparate models into cohesive, unified records that empower complex, multi-dimensional queries.</p><p>Cleanly decoupling these intensive processing tasks from the ingestion pipeline guarantees that complex data intersections never bottleneck real-time intake. As a result, the system maintains maximum uptime and peak responsiveness, even when processing the massive scale of the Netflix media catalog.</p><p><strong>Temporal Bucketing and Intersection</strong></p><p>To achieve this intersection at scale, the offline pipeline normalizes disparate model outputs by mapping them into fixed-size temporal buckets (one-second intervals). This discretization process unfolds in three steps:</p><ul><li><strong>Bucket Mapping:</strong> Continuous detections are segmented into discrete intervals. For example, if a model detects a character (“Joey”) from seconds 2 through 8, the pipeline maps this continuous span of frames into seven distinct one-second buckets.</li><li><strong>Annotation Intersection:</strong> When multiple models generate annotations for the exact same temporal bucket, such as character recognition <em>“Joey”</em> and scene detection <em>“kitchen” </em>overlapping in second 4, the system fuses them into a single, comprehensive record.</li><li><strong>Optimized Persistence:</strong> These newly enriched records are written back to Cassandra as distinct entities. This creates a highly optimized, second-by-second index of multi-modal intersections, perfectly associating every fused annotation with its source asset.</li></ul><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*ft36HZpYbqUBfhP8p0iShw.png"><figcaption><em>Figure 3: Temporal Data Fusion with Fixed-Size Time Buckets</em></figcaption></figure><p>The following record shows the overlap of the character<em> “Joey”</em> and scene <em>“kitchen” </em>annotations during a 4 to 5 second window in a video asset:</p><pre>{<br>  "associated_ids": {<br>    "MOVIE_ID": "81686010",<br>    "ASSET_ID": "01325120–7482–11ef-b66f-0eb58bc8a0ad"<br>  },<br>  "time_bucket_start_ns": 4000000000,<br>  "time_bucket_end_ns": 5000000000,<br>  "source_annotations": [<br>    {<br>      "annotation_id": "7f5959b4–5ec7–11f0-b475–122953903c43",<br>      "annotation_type": "CHARACTER_SEARCH",<br>      "label": "Joey",<br>      "time_range": {<br>        "start_time_ns": 2000000000,<br>        "end_time_ns": 8000000000<br>      }<br>    },<br>    {<br>      "annotation_id": "c9d59338–842c-11f0–91de-12433798cf4d",<br>      "annotation_type": "SCENE_SEARCH",<br>      "time_range": {<br>        "start_time_ns": 4000000000,<br>        "end_time_ns": 9000000000<br>      },<br>      "label": "kitchen",<br>      "embedding_vector": [<br>        0.9001, 0.00123 ....<br>      ]<br>    }<br>  ]<br>}</pre><p><em>Figure 4: Sample Intersection Record For Character + Scene Search</em></p><h4>3. Indexing for Real Time Search</h4><p>Once the enriched temporal buckets are securely persisted in Cassandra, a subsequent event triggers their ingestion into Elasticsearch.</p><p>To guarantee absolute data consistency, the pipeline executes upsert operations using a composite key <em>(asset ID + time bucket)</em> as the unique document identifier. If a temporal bucket already exists for a specific second of video, perhaps populated by an earlier model run, the system intelligently updates the existing record rather than generating a duplicate. This mechanism establishes a single, unified source of truth for every second of footage.</p><p>Architecturally, the pipeline structures each temporal bucket as a nested document. The root level captures the overarching asset context, while associated child documents house the specific, multi-modal annotation data. This hierarchical data model is precisely what empowers users to execute highly efficient, cross-annotation queries at scale.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*Mt1fbK8AvwgNk-LY55wFqg.png"><figcaption><em>Figure 5: Simplified Elasticsearch Document Structure</em></figcaption></figure><h3>Multimodal Discovery and Result Ranking</h3><p>The search service provides a high-performance interface for real-time discovery across the global Netflix catalog. Upon receiving a user request, the system immediately initiates a query preprocessing phase, generating a structured execution plan through three core steps:</p><ul><li><strong>Query Type Detection:</strong> Dynamically categorizes the incoming request to route it down the most efficient retrieval path.</li><li><strong>Filter Extraction:</strong> Isolates specific semantic constraints such as character names, physical objects, or environmental contexts to rapidly narrow the candidate pool.</li><li><strong>Vector Transformation:</strong> Converts raw text into high-dimensional, model-specific embeddings to enable deep, context-aware semantic matching.</li></ul><p>Once generated, the system compiles this structured plan into a highly optimized Elasticsearch query, executing it directly against the pre-fused temporal buckets to deliver instantaneous, frame-accurate results.</p><h4>Fine-Tuning Semantic Search</h4><p>To support the diverse workflows of different production teams, the system provides fine-grained control over search behavior through configurable parameters:</p><ul><li><strong>Exact vs. Approximate Search:</strong> Users can toggle between exact k-Nearest Neighbors (k-NN) for uncompromising precision, and Approximate Nearest Neighbor (ANN) algorithms (such as <a href="https://en.wikipedia.org/wiki/Hierarchical_navigable_small_world">HNSW</a>) to maintain blazing speed when querying massive datasets.</li><li><strong>Dynamic Similarity Metrics:</strong> The system supports multiple distance calculations, including <a href="https://en.wikipedia.org/wiki/Cosine_similarity">cosine similarity</a> and <a href="https://en.wikipedia.org/wiki/Euclidean_distance">Euclidean distance</a>. Because different models shape their high-dimensional vector spaces distinctly based on their underlying training architectures, the flexibility to swap metrics ensures that mathematical closeness perfectly translates to true semantic relevance.</li><li><strong>Confidence Thresholding:</strong> By establishing strict minimum score boundaries for results, users can actively prune the “long tail” of low-probability matches. This aggressively filters out visual noise, guaranteeing that creative teams are not distracted and only review results that meet a rigorous standard of mathematical similarity.</li></ul><h4>Textual Analysis &amp; Linguistic Precision</h4><p>To handle the deep nuances of dialogue-heavy searches, such as isolating a character’s exact catchphrase amidst thousands of hours of speech, we implement a sophisticated text analysis strategy within Elasticsearch. This ensures that conversational context is captured and indexed accurately.</p><ul><li><strong>Phrase &amp; Proximity Matching:</strong> To respect the narrative weight of specific lines (e.g., “Friends don’t lie” in Stranger Things), we leverage <a href="https://www.elastic.co/docs/reference/query-languages/query-dsl/query-dsl-match-query-phrase">match-phrase</a> queries with a configurable <em>slop</em> parameter. This guarantees the system retrieves the correct scene even if the user’s memory slightly deviates from the exact transcription.</li><li><strong>N-Gram Analysis for Partial Discovery:</strong> Because video search is inherently exploratory, we utilize edge <a href="https://en.wikipedia.org/wiki/N-gram">N-gram</a> tokenizers to support “search-as-you-type” functionality. By actively indexing dialogue and metadata substrings, the system surfaces frame-accurate results the moment an editor begins typing, drastically reducing cognitive load.</li><li><strong>Tokenization and Linguistic Stemming:</strong> To seamlessly support the global scale of the Netflix catalog, our analysis chain applies sophisticated <a href="https://en.wikipedia.org/wiki/Stemming">stemming</a> across multiple languages. This ensures a query for “running” automatically intersects with scenes tagged with “run” or “ran,” collapsing grammatical variations into a single, unified search intent.</li><li><strong>Levenshtein Fuzzy Matching:</strong> To account for transcription anomalies or phonetic misspellings, we incorporate fuzzy search capabilities based on <a href="https://en.wikipedia.org/wiki/Levenshtein_distance">Levenshtein</a> distance algorithms. This intelligent soft-matching approach ensures that high-value shots are never lost to minor data-entry errors or imperfect queries.</li></ul><h4>Aggregations and Flexible Grouping</h4><p>The architecture operates at immense scale, seamlessly executing queries within a single title or across thousands of assets simultaneously. To combat result fatigue, the system leverages custom aggregations to intelligently cluster and group outputs based on specific parameters, such as isolating the top 5 most relevant clips of an actor per episode. This guarantees a diverse, highly representative return set, preventing any single asset from dominating the search results.</p><h4>Search Response Curation</h4><p>While temporal buckets are the internal mechanism for search efficiency, the system post-processes Elasticsearch results to reconstruct original time boundaries. The reconstruction process ensures results reflect narrative scene context rather than arbitrary intervals. Depending on the query intent, the system generates results based on two logic types:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*JamOz6AmYyLfsCOBtUY3Bg.png"><figcaption><em>Figure 6: Depiction of Temporal Union vs Intersection</em></figcaption></figure><ul><li><strong>Union:</strong> Returns the full span of all matching annotations <em>(3–8 sec),</em> which prioritizes breadth, capturing any instance where a specified feature occurs.</li><li><strong>Intersection:</strong> Returns only the exact overlapping duration of matching signals <em>(4–6 sec).</em> The intersection logic focuses on co-occurrence, isolating moments when multiple criteria align.</li></ul><pre>{<br>  "entity_id": {<br>    "entity_type": "ASSET",<br>    "id": "1bba97a1–3562–4426–9cd2-dfbacddcb97b"<br>  },<br>  "range_intervals": [<br>    {<br>      "intersection_time_range": {<br>        "start_time_ns": 4000000000,<br>        "end_time_ns": 8000000000<br>      },<br>      "union_time_range": {<br>        "start_time_ns": 2000000000,<br>        "end_time_ns": 9000000000<br>      },<br>      "source_annotations": [<br>        {<br>          "annotation_id": "fc1525d0–93a7–11ef-9344–1239fc3a8917",<br>          "annotation_type": "SCENE_SEARCH",<br>          "metadata": {<br>            "label": "kitchen"<br>          }<br>        },<br>        {<br>          "annotation_id": "5974fb01–93b0–11ef-9344–1239fc3a8917",<br>          "annotation_type": "CHARACTER_SEARCH",<br>          "metadata": {<br>            "character_name": [<br>              "Joey"<br>            ]<br>          }<br>        }<br>      ]<br>    }<br>  ]<br>}</pre><p><em>Figure 7: Sample Query Response</em></p><h3><strong>Future Extensions</strong></h3><p>While our current architecture establishes a highly resilient and scalable foundation, it represents only the first phase of our multi-modal search vision. To continuously close the gap between human intuition and machine retrieval, our roadmap focuses on three core evolutions:</p><ul><li><strong>Natural Language Discovery:</strong> Transitioning from structured JSON payloads to fluid, conversational interfaces (e.g., <em>“Find the best tracking shots of Tom Holland running on a roof”</em>). This will abstract away underlying query complexity, allowing creatives to interact with the archive organically.</li><li><strong>Adaptive Ranking:</strong> Implementing machine learning feedback loops to dynamically refine scoring algorithms. By continuously analyzing how editorial teams interact with and select clips, the system will self-tune its mathematical definition of semantic relevance over time.</li><li><strong>Domain-Specific Personalization:</strong> Dynamically calibrating search weights and retrieval behaviors to match the exact context of the user. The platform will tailor its results depending on whether a team is cutting high-action marketing trailers, editing narrative scenes, or conducting deep archival research.</li></ul><p>Ultimately, these advancements will elevate the platform from a highly optimized search engine into an intelligent creative partner, fully equipped to navigate the ever-growing complexity and scale of global video media.</p><h3><strong>Acknowledgements</strong></h3><p>We would like to extend our gratitude to the following teams and individuals whose expertise and collaboration were instrumental in the development of this system:</p><ul><li><strong>Data Science Engineering:</strong> <a href="https://www.linkedin.com/in/nagendrak/">Nagendra Kamath</a>, <a href="https://www.linkedin.com/in/chao-pan-02791824/">Chao Pan</a>, <a href="mailto:prachees@netflix.com">Prachee Sharma</a>, <a href="https://www.linkedin.com/in/ying-liao-nyu/">Ying Liao</a> and <a href="mailto:csoo@netflix.com">Carolyn Soo</a> for the critical media model insights that informed our architectural design.</li><li><strong>Product Management:</strong> <a href="https://www.linkedin.com/in/nimesh-narayan/">Nimesh Narayan</a>, <a href="https://www.linkedin.com/in/ian-krabacher-06816710/">Ian Krabacher</a>, <a href="https://www.linkedin.com/in/ananya-ani-poddar/">Ananya Poddar</a>, <a href="https://www.linkedin.com/in/meghanbailey02/">Meghan Bailey</a> and <a href="https://www.linkedin.com/in/anitakuc/">Anita Kuc</a> for defining the user requirements and product vision.</li><li><strong>Media Production Suite Team:</strong> <a href="https://www.linkedin.com/in/szymon-borodziuk/">Szymon Borodziuk</a>, <a href="https://www.linkedin.com/in/mike-czarnota/">Mike Czarnota</a>, <a href="https://www.linkedin.com/in/dsarkowicz/?locale=en">Dominika Sarkowicz</a>, <a href="mailto:bkoval@netflix.com">Bohdan Koval</a> and <a href="https://www.linkedin.com/in/sabov/">Sasha Sabov</a> for their work in engineering the end-user search experience.</li><li><strong>Asset Management Platform Team:</strong> For their collaborative efforts in operationalizing this design and bringing the system into production.</li></ul><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=3e0020cf1202" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/powering-multimodal-intelligence-for-video-search-3e0020cf1202">Powering Multimodal Intelligence for Video Search</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/powering-multimodal-intelligence-for-video-search-3e0020cf1202</link>
      <guid>https://netflixtechblog.com/powering-multimodal-intelligence-for-video-search-3e0020cf1202</guid>
      <pubDate>Sat, 04 Apr 2026 02:44:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Smarter Live Streaming at Scale: Rolling Out VBR for All Netflix Live Events]]></title>
      <description><![CDATA[<p>By Renata Teixeira, Zhi Li, Reenal Mahajan, and Wei Wei</p><p>On January 26, 2026, we flipped an important switch for Live at Netflix: <strong>all Live events are now encoded using VBR (Variable Bitrate) instead of CBR (Constant Bitrate)</strong>. It sounds like a small configuration change, but it required us to revisit some of the foundational assumptions behind how we deliver Live video at global scale.</p><p>VBR lets us tailor the bitrate to the actual complexity of the scene, instead of sending every second of video at roughly the same bitrate. When a scene is simple, VBR “shaves off” bits that wouldn’t improve what you see on screen; when a scene is complex, it spends more bits to preserve quality. (The more general idea is often referred to as capped variable bitrate, or capped VBR.) That makes our encodes more efficient and our network more scalable. But it also makes traffic much less predictable: large bitrate swings can overload servers and CDNs, and our old assumptions about “what bitrate equals what quality” no longer hold. As a result, we have to rethink both how we manage delivery and capacity, and which bitrates we offer for each version of the stream. In our live pipeline, we currently use AWS Elemental MediaLive, where this “capped” VBR is implemented using the QVBR (Quality‑Defined Variable Bitrate) setting.</p><h3>| Why Move Live from CBR to VBR?</h3><p>Our initial Live encoding pipeline used constant bitrate (CBR). For each encoded stream, we configured a resolution and a nominal bitrate — for example, a 1080p stream targeting 5 Mbps — and the actual bitrate stayed close to that target over time. This predictability made both capacity planning and day‑to‑day operations easier. If a server could safely deliver around 100 Gbps of Live traffic, and each stream averaged close to its nominal rate, we could admit on the order of twenty thousand concurrent sessions per server and be confident we were operating within limits. During an event, the total traffic sent by a server would change mainly when members joined or left; as long as concurrency was stable, traffic stayed relatively flat. The network saw a smooth, easy‑to‑reason‑about load profile, and large changes in throughput almost always reflected a real change in usage, not just a different scene on screen.</p><p>The problem is that content isn’t constant. A talking‑head segment in a studio or a simple animation is much easier to compress than a sequence of rapid camera moves and fast‑moving athletes in front of a highly detailed crowd. With CBR, both the easy and the hard segments get the same bitrate. In simple scenes, we spend more bits than we need; in complex scenes, we sometimes don’t spend enough.</p><p>VBR flips the objective. Rather than aiming for a fixed bitrate, the encoder aims for a target quality and is allowed to raise or lower the bitrate according to scene complexity. When the picture is easy to encode, VBR can drop the bitrate substantially below the old CBR level while keeping quality constant. When the action heats up, it can temporarily use more bits to avoid visible artifacts.</p><p>The figure below shows the per‑segment bitrate over time for the same episode of WWE RAW, encoded once with CBR and once with VBR at a nominal 8 Mbps. With CBR (blue), the bitrate wobbles a bit from segment to segment but stays close to the target; if you average it over a minute, it’s practically a straight line. We end up spending roughly the same number of bits on simple scenes, like the waiting room at the start of the stream (shaded region), as on complex scenes, like the confetti‑filled shot later in the show. With VBR (orange), the encoder can drop the bitrate for the waiting room to a small fraction of the nominal rate, while allowing the confetti‑filled shot to use a much higher bitrate to preserve quality.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*_MoqLhu01wArcD1n58SZyg.png"><figcaption>Per-segment bitrate over a WWE RAW episode for CBR and VBR encodes at a nominal of 8 Mbps.</figcaption></figure><figure><img alt="" src="https://cdn-images-1.medium.com/max/512/1*huWI8zAfmmwWq5IsxKIL3Q.png"><figcaption><strong><em>“Waiting room” scene:</em></strong><em> visually simple and easy to compress, so VBR can safely use a low bitrate.</em></figcaption></figure><figure><img alt="" src="https://cdn-images-1.medium.com/max/512/1*jBCcIjhCgSi9kPXLrsd1FA.png"><figcaption><strong><em>“Confetti-filled” shot:</em></strong><em> visually complex and noisy, so VBR spends many more bits to maintain quality.</em></figcaption></figure><p>At Netflix scale, the shift to VBR has several important effects. The most important is efficiency on the network: we reduce the average number of bytes we need to deliver a full event, which in turn reduces the traffic needed to fill all of the servers in Open Connect, our content delivery network (CDN), and the traffic needed to serve segments to members. The second is quality of experience. Because VBR sends fewer bits for similar video quality, we see fewer rebuffers and lower start‑up delay. Across multiple A/B tests on different Live events, we observed about 5% fewer rebuffers per hour, while transferring roughly 15% fewer bytes on average and around a 10% reduction in traffic at the peak minute.</p><h3>| When Efficiency Fights Stability</h3><p>The central challenge with VBR for Live is not that bitrate varies at all (it does under CBR as well), but that VBR can have much deeper, longer dips in bitrate, tightly coupled to what’s happening in the content.</p><p>Under CBR, a 5 Mbps stream is effectively that: on a per‑minute basis, traffic is remarkably flat. A server that is comfortably handling, say, ten thousand such sessions now is likely to still be comfortable in a minute, barring a wave of new joins. It’s safe for our steering logic to look at current traffic, see plenty of headroom, and route additional sessions to that server.</p><p>Under VBR, the same stream behaves very differently. During a slow, easy sequence, the encoder might only generate 2 Mbps for the 5 Mbps stream — or even less — to maintain its target quality, and it can stay at that lower level for an extended period. The server then appears to have plenty of unused capacity: per‑session bitrate is low and aggregate traffic is well below its limits. Our steering systems naturally interpret this as a signal that the server is under‑utilized and can accept more sessions.</p><p>The problem surfaces when the content changes. A fight starts, confetti begins to fall, or the camera cuts to a highly detailed, fast‑moving shot. To preserve quality, VBR may increase the bitrate to 6, 7, or 8 Mbps on the very next segments. If the server has admitted many additional sessions during the preceding low‑bitrate period, the aggregate traffic can suddenly exceed what the network link or NIC can sustain. Latency rises, packets are dropped, and devices start to experience stalls or quality downshifts. In extreme cases, this pattern of “bitrate dips followed by spikes” can destabilize parts of the system.</p><h3>| Making Servers Aware of Bitrate Variability</h3><p>These long VBR bitrate dips are great for efficiency, but — as we just saw — they can trick our delivery systems into thinking servers are under‑utilized and safe to load up. Under CBR, that behavior was predictable enough that current traffic was a good proxy for how “full” a server was; under VBR, it isn’t.</p><p>Our fix was to change how we decide whether a server can take more sessions. Instead of basing that decision only on current traffic, we reserve capacity based on each stream’s nominal bitrate, not just what it happens to be using at that moment. Even if a VBR stream is currently in a very cheap, low‑bitrate phase, we still treat it as something that can quickly return to its nominal rate.</p><p>This keeps our traffic‑steering behavior consistent between CBR and VBR and avoids the key failure mode where a server accepts too many sessions during a long low‑bitrate period and then becomes overloaded when bitrate rises again.</p><h3>| Tuning VBR Nominal Bitrates to Match CBR Quality</h3><p>The WWE example above already hints at why this isn’t automatic. We looked at a single 8 Mbps stream, encoded once with CBR and once with VBR. Both encodes have the same nominal bitrate, but the figure shows how differently they behave. The CBR encode stays clustered around 8 Mbps with frequent short spikes, while the VBR encode often drops far below that level and only spikes up when the content gets complex. That’s more efficient, but it also means that “same nominal bitrate” does not imply “same average number of bits” anymore — so simply reusing our CBR settings risks giving VBR less bitrate on average and losing some quality.</p><p>In practice, of course, we don’t just encode a single stream; we produce a set of streams at different resolutions and nominal bitrates — often called a bitrate ladder — so devices can adapt to their current network conditions by switching between them. When we first applied VBR using the existing CBR ladder, offline analysis with <a href="https://github.com/Netflix/vmaf">VMAF</a> (a <a href="https://netflixtechblog.com/toward-a-practical-perceptual-video-quality-metric-653f208b9652">perceptual video quality metric</a>) confirmed the concern from the WWE example: time‑averaged quality dropped slightly on a few streams, especially at the lowest bitrates. Early A/B tests showed the same pattern: overall VMAF about one point lower than CBR, with most of the gap at the bottom of the ladder.</p><p>To fix this, we compared CBR and VBR encodes rung by rung and looked at per‑stream VMAF. Wherever VBR fell more than about one VMAF point below CBR, we increased its nominal bitrate just enough to close the gap. Higher‑bitrate streams, where VBR quality was already very close to CBR, were left largely unchanged, including the 8 Mbps stream from the figure.</p><p>The result is a VBR ladder with slightly higher nominal bitrates on a few low‑end streams, but lower overall traffic, because VBR still drops the bitrate on simple scenes. This lets us match the quality of our CBR ladder while keeping the efficiency and stability gains that motivated the switch to VBR in the first place.</p><h3>| What’s Next for Live VBR</h3><p>With VBR in production for all Live events, our focus now is on using it more intelligently.</p><p>First, we are testing how to use the actual sizes of upcoming segments in our adaptive bitrate algorithms on devices, instead of relying only on nominal bitrates. This should help devices pick streams that better match how VBR will behave in the next few seconds, not just how a stream is labeled on paper.</p><p>Second, we are experimenting with making our capacity reservation less conservative. Today we reserve based on nominal bitrates to keep servers safe; by carefully applying a “discount” informed by real VBR behavior, we hope to free up additional headroom without sacrificing stability.</p><p>This work was the result of a broad, cross‑team effort. We’d like to thank Mariana Afonso, Dave Andrews, Mark Brady, Jake Freeland, Te-Yuan Huang, Ivan Ivanov, Yeshwenth Jayaraman, Patrick Kunka, Zheng Lu, Anirudh Mendiratta, Chris Pham, David Pfitzner, Jon Rivas, Garett Singer, Brenda So, Stan Surmay, Bowen Tan, Devashish Thakur, and Allan Zhou for their many contributions to making Live VBR a reality at Netflix.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=c8f833b238cc" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/smarter-live-streaming-at-scale-rolling-out-vbr-for-all-netflix-live-events-c8f833b238cc">Smarter Live Streaming at Scale: Rolling Out VBR for All Netflix Live Events</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/smarter-live-streaming-at-scale-rolling-out-vbr-for-all-netflix-live-events-c8f833b238cc</link>
      <guid>https://netflixtechblog.com/smarter-live-streaming-at-scale-rolling-out-vbr-for-all-netflix-live-events-c8f833b238cc</guid>
      <pubDate>Thu, 02 Apr 2026 23:46:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Scaling Global Storytelling: Modernizing Localization Analytics at Netflix]]></title>
      <description><![CDATA[<p><a href="https://www.linkedin.com/in/valentingeffrier/">Valentin Geffrier</a>, <a href="https://www.linkedin.com/in/tanguycornuau/">Tanguy Cornuau</a></p><p><em>Each year, we bring the Analytics Engineering community together for an Analytics Summit — a multi-day internal conference to share analytical deliverables across Netflix, discuss analytic practice, and build relationships within the community. This post is one of several topics presented at the Summit highlighting the breadth and impact of Analytics work across different areas of the business.</em></p><p>At Netflix, our goal is to entertain the world, which means we must speak the world’s languages. Given the company’s growth to serving 300 million+ members in more than 190+ countries and 50+ languages, the Localization team has had to scale rapidly in creating more dubs and subtitle assets than ever before. However, this growth created technical debt within our systems: a fragmented landscape of analytics workflows, duplicated pipelines, and siloed dashboards that we are now actively modernizing.</p><h4>The Challenge: “Who Made This Dub?”</h4><p>Historically, business logic for localization metrics was replicated across isolated domains. A question as simple as “<em>Who made this dub/subtitle?”</em> is actually complex — it requires mapping multiple data sources through intricate and constantly changing logic, which varies depending on the specific language asset type and creation workflow.</p><p>When this logic is copied into isolated pipelines for different use cases it creates two major risks: inconsistency in reporting and a massive maintenance burden whenever upstream logic changes. We realized we needed to move away from these vertical silos.</p><h4>Our Modernization Strategy</h4><p>To address this, we defined a vision centered on consolidation, standardization, and trust, executed through three strategic pillars:</p><p>1. The Audit and Consolidation Playbook</p><p>We initiated a comprehensive audit of over 40 dashboards and tools to assess usage and code quality. Our focus has shifted from patching frontend visualizations to consolidating backend pipelines. For example, we are currently merging three legacy dashboards related to dubbing partner KPIs (around operational performance, capacity, and finances), focusing first on a unified data and backend layer that can support a variety of future frontend iterations.</p><p>2. Reducing “Not-So-Tech” Debt</p><p>Technical debt isn’t just about code; it is also about the user experience. We define “Not-So-Tech Debt” as the friction stakeholders feel when tools are hard to interpret or can benefit from better storytelling. To fix this, we revamped our Language Asset Consumption tool — instead of reporting dub and subtitle metrics independently, we combine audio and text languages into one consumption language that helps differentiate Original Language versus Localized Consumption and measure member preferences between subtitles, dubs, or a combination of both for a given language. This unlocks more intuitive insights based on actual recurring stakeholder use cases.</p><p>3. Investing in Core Building Blocks</p><p>We are shifting to a <em>write once, read many</em> architecture. By centralizing business logic into unified tables — such as a “Language Asset Producer” table — we solve the “<em>Who made this dub?”</em> problem once. This centralized source now feeds into multiple downstream domains, including our Dub Quality and Translation Quality metrics, ensuring that any logic update propagates instantly across the ecosystem.</p><h4>The Future: Event-Level Analytics</h4><p>Looking ahead, we are moving beyond asset-level metrics to event-level analytics. We are building a generic data model to capture granular timed-text events, such as individual subtitle lines. This data helps us understand how subtitle characteristics (e.g. reading speed) affect member engagement and, in turn, refine the style guidelines we provide to our subtitle linguists to improve the member experience with localized content.</p><p>Ultimately, this modernization effort is about scaling our ability to measure and enhance the joy and entertainment we deliver to our diverse global audience, ensuring that every member, regardless of their language, has the best possible Netflix experience.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=816f47290641" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/scaling-global-storytelling-modernizing-localization-analytics-at-netflix-816f47290641">Scaling Global Storytelling: Modernizing Localization Analytics at Netflix</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/scaling-global-storytelling-modernizing-localization-analytics-at-netflix-816f47290641</link>
      <guid>https://netflixtechblog.com/scaling-global-storytelling-modernizing-localization-analytics-at-netflix-816f47290641</guid>
      <pubDate>Fri, 06 Mar 2026 16:01:00 +0100</pubDate>
    </item>
    <item>
      <title><![CDATA[Optimizing Recommendation Systems with JDK’s Vector API]]></title>
      <description><![CDATA[<p><em>By </em><a href="https://www.linkedin.com/in/harshad-sane-56711a11/"><em>Harshad San</em></a><em>e</em></p><p>Ranker is one of the largest and most complex services at Netflix. Among many things, it powers the personalized rows you see on the Netflix homepage, and runs at an enormous scale. When we looked at CPU profiles for this service, one feature kept standing out: <strong>video serendipity scoring</strong> — the logic that answers a simple question:</p><p><em>“How different is this new title from what you’ve been watching so far?”</em></p><p>This single feature was consuming about 7.5% of total CPU on each node running the service. What started as a simple idea — “just batch the video scoring feature” — turned into a deeper optimization journey. Along the way we introduced batching, re-architected memory layout and tried various libraries to handle the scoring kernels.</p><p>Read on to learn how we achieved the same serendipity scores, but at a meaningfully lower CPU per request, resulting in a reduced cluster footprint.</p><h3><strong>Problem: The Hotspot in Ranker</strong></h3><p>At a high level, serendipity scoring works like this: A candidate title and each item in a member’s viewing history are represented as embeddings in a vector space. For each candidate, we compute its similarity against the history embeddings, find the maximum similarity, and convert that into a “novelty” score. That score becomes an input feature to the downstream recommendation logic.</p><p>The original implementation was straightforward but expensive. For each candidate we fetch its embedding, loop over the history to compute cosine similarity one pair at a time and track the maximum similarity score. Although it is easy to reason about, at Ranker’s scale, this results in significant sequential work, repeated embedding lookups, scattered memory access, and poor cache locality. Profiling confirmed this.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/655/1*w8Z7CwTNc4dW84n8-S-CMw.png"><figcaption>Flamegraph showing inefficient scoring</figcaption></figure><p>A flamegraph made it clear: One of the top hotspots in the service was Java dot products inside the serendipity encoder. Algorithmically, the hotspot was a nested loop structure of M candidates × N history items where each pair generates its own cosine similarity i.e. O(M×N) separate dot product operations.</p><h3><strong>Solution</strong></h3><h4>The Original Implementation: Single video cosine loop</h4><p>In simplified form the code looked like this:</p><pre>for (Video candidate : candidates) {<br>  Vector c = embedding(candidate); // D-dimensional<br>  double maxSim = -1.0;<br><br>  for (Video h : history) {<br>    Vector v = embedding(h); // D-dimensional<br>    double sim = cosine(c, v); // dot(c, v) / (||c|| * ||v||)<br>    maxSim = Math.max(maxSim, sim);<br>  }<br><br>  double serendipity = 1.0 - maxSim;<br>  emitFeature(candidate, serendipity);<br>}</pre><p>The nested for loop with O(M×N) separate dot products brought upon its own overheads. One interesting detail we learned by instrumenting traffic shapes: most requests (about 98%) were single-video, but the remaining 2% were large batch requests. Because those batches were so large, the total volume of videos processed ended up being roughly 50:50 between single and batch jobs. This made batching worth pursuing even if it didn’t help the median request.</p><h4>Step 1 : Batching, from Nested Loops to Matrix Multiply</h4><p>The first idea was to stop thinking in terms of “many small dot products” and instead treat the work as a matrix operation. i.e. For batch candidates, implement a data layout to parallelize the math in a single operation i.e. matrix multiply. If D is the embedding dimension:</p><ol><li>Pack all candidate embeddings into a matrix A of shape M x D</li><li>Pack all history embeddings into a matrix B of shape N x D</li><li>Normalize all rows to unit length.</li><li>Compute: cosine similarities as <br> [ C = A x B^T ]; where C is an M x N matrix of cosine similarities.</li></ol><p>In pseudo‑code:</p><pre>// Build matrices<br>double[][] A = new double[M][D]; // candidates<br>double[][] B = new double[N][D]; // history<br><br>for (int i = 0; i &lt; M; i++) {<br>  A[i] = embedding(candidates[i]).toArray();<br>}<br>for (int j = 0; j &lt; N; j++) {<br>  B[j] = embedding(history[j]).toArray();<br>}<br><br>// Normalize rows to unit vectors<br>normalizeRows(A);<br>normalizeRows(B);<br><br>// Compute C = A * B^T<br>double[][] C = matmul(A, B);<br>C[i][j] = cosine(candidates[i], history[j])<br><br>// Derive serendipity<br>for (int i = 0; i &lt; M; i++) {<br>  double maxSim = max(C[i][0..N-1]);<br>  double serendipity = 1.0 - maxSim;<br>  emitFeature(candidates[i], serendipity);<br>}</pre><p>This turns <strong>M×N separate dot products into a single matrix multiply</strong>, which is exactly what CPUs and optimized kernels are built for. We integrated this into the existing framework by supporting both, encode()for single videos and batchEncode() for batches, while maintaining backward compatibility. At this point it seemed like we were “done”, but we weren't.</p><h4>Step 2: When Batching Isn’t Enough</h4><p>Once we had a batched implementation, we ran canaries and saw something surprising: about a 5% performance regression. The algorithm wasn’t the issue — turning M×N separate dot products into a matrix multiplication is mathematically sound. The problem was the overhead we introduced in the first implementation.</p><ol><li>Our initial version built double[][] matrices for candidates, history, and results on every batch. Those large, short-lived allocations created GC pressure, and the double[][] layout itself is non-contiguous in memory, which meant extra pointer chasing and worse cache behavior.</li><li>On top of that, the first-cut Java matrix multiply was a straightforward scalar implementation, so it couldn’t take advantage of SIMD. In other words, we paid the cost of batching without getting the compute efficiency we were aiming for.</li></ol><p>The lesson was immediate: algorithmic improvements don’t matter if the implementation details—memory layout, allocation strategy, and the compute kernel—work against you. That set up the next step for making the data layout cache-friendly and eliminating per-batch allocations before revisiting the matrix multiply kernel.</p><h4><strong>Step 3: Flat Buffers &amp; ThreadLocal Reuse</strong></h4><p>We reworked the data layout to be cache-friendly and allocation-light. Instead of double[m][n], we moved to flat double[] buffers in row-major order. That gave us contiguous memory and predictable access patterns. Then we introduced a ThreadLocal&lt;BufferHolder&gt; that owns reusable buffers for candidates, history, and any other scratch space. Buffers grow as needed but never shrink, which avoids per-request allocation while keeping each thread isolated (no contention). A simplified sketch:</p><pre>class BufferHolder {  <br>  double[] candidatesFlat = new double[0];  <br>  double[] historyFlat = new double[0];  <br><br>  double[] getCandidatesFlat(int required) {  <br>    if (candidatesFlat.length &lt; required) {  <br>      candidatesFlat = new double[required];  <br>    }  <br>    return candidatesFlat;  <br>  }  <br>  <br>  double[] getHistoryFlat(int required) {  <br>    if (historyFlat.length &lt; required) {  <br>      historyFlat = new double[required];  <br>    }  <br>    return historyFlat;  <br>  }  <br>}  <br><br>private static final ThreadLocal&lt;BufferHolder&gt; threadBuffers =  <br>    ThreadLocal.withInitial(BufferHolder::new);</pre><p>This change alone made the batched path far more predictable: fewer allocations, less GC pressure, and better cache locality.</p><p>Now the remaining question was the one we originally thought we were answering: what’s the best way to do the matrix multiply?</p><p><strong>Step 4: BLAS: Great in Tests, Not in Production</strong></p><p>The obvious next step was BLAS (Basic Linear Algebra Subprograms). In isolation, microbenchmarks looked promising. But once integrated into the real batch scoring path, the gains didn’t materialize. A few things were working against us:</p><ul><li>The default netlib-java path was using F2J (Fortran-to-Java) BLAS rather than a truly native implementation.</li><li>Even with native BLAS, we paid overhead for setup and JNI transitions.</li><li>Java’s row-major layout doesn’t match the column-major expectations of many BLAS routines, which can introduce conversion and temporary buffers.</li><li>Those extra allocations and copies mattered in the full pipeline, especially alongside TensorFlow embedding work.</li></ul><p>BLAS was still a useful experiment — it clarified where time was being spent, but it wasn’t the drop-in win we wanted. What we needed was something that stayed pure Java, fit our flat-buffer architecture, and could still exploit SIMD.</p><p><strong>Step 5: JDK Vector API to the rescue</strong></p><p><strong><em>A Short Note on the JDK Vector API: </em></strong>The JDK Vector API is an <em>incubating</em> feature that provides a portable way to express data-parallel operations in Java — think “SIMD without intrinsics”. You write in terms of vectors and lanes, and the JIT maps those operations to the best SIMD instructions available on the host CPU (SSE/AVX2/AVX-512), with a scalar fallback when needed. More crucially for us, it’s pure Java: no native dependencies, no JNI transitions, and a development model that looks like normal Java code rather than platform-specific assembly or intrinsics.</p><p>This was a particularly good match for our workload because we had already moved embeddings into flat, contiguous double[] buffers, and the hot loop was dominated by large numbers of dot products. The final step was to replace BLAS with a pure-Java SIMD implementation using the JDK Vector API. By this point we already had the right shape for high performance — batching, flat buffers, and ThreadLocal reuse. So the remaining work was to swap out the compute kernel without introducing JNI overhead or platform-specific code. We did that behind a small factory. At class load time, MatMulFactory selects the best available implementation:</p><ul><li>If jdk.incubator.vector is available, use a Vector API implementation.</li><li>Otherwise, fall back to a scalar implementation with a highly optimized loop-unrolled dot product (implemented by my colleague Patrick Strawderman, inspired by patterns used in <a href="https://github.com/apache/lucene/blob/6d4314d46fd69ca16edce0cd1c8507aa0e66ccd6/lucene/core/src/java/org/apache/lucene/util/VectorUtilDefaultProvider.java#L26">Lucene</a>)</li></ul><p>In the Vector API implementation, the inner loop computes a dot product by accumulating a * b into a vector accumulator using fma() (fused multiply-add). DoubleVector.SPECIES_PREFERRED lets the runtime pick an appropriate lane width for the machine. Here’s a simplified sketch of the inner loop:</p><pre>// Vector API path (simplified)  <br>for (int i = 0; i &lt; M; i++) {  <br>  for (int j = 0; j &lt; N; j++) {  <br>  <br>  DoubleVector acc = DoubleVector.zero(SPECIES);  <br>    int k = 0;  <br>    // SPECIES.length() (e.g. often 4 doubles on AVX2 and 8 doubles on AVX-512). <br>    for (; k + SPECIES.length() &lt;= D; k += SPECIES.length()) {  <br>      DoubleVector a = DoubleVector.fromArray(SPECIES, candidatesFlat, i*D + k);  <br>      DoubleVector b = DoubleVector.fromArray(SPECIES, historyFlat,   j*D + k);<br>      acc = a.fma(b, acc);  // fused multiply-add  <br>    }  <br>    double dot = acc.reduceLanes(VectorOperators.ADD);  <br>    // handle tail k..D-1  <br>    similaritiesFlat[i*N + j] = dot;  <br>  }  <br>}</pre><p>Figure below shows how the Vector API utilizes SIMD hardware to process multiple doubles per instruction (e.g., 4 lanes on AVX2 and 8 lanes on AVX‑512). What used to be many scalar multiply-adds becomes a smaller number of vector fma() operations plus a reduction—same algorithm, much better use of the CPU’s vector units.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*xU5DvGS8SVz3hfdrOfC1KQ.png"><figcaption>Vectorization with SIMD</figcaption></figure><h4>Fallbacks &amp; Safety: When the Vector API Isn’t Available</h4><p>Because the Vector API is still incubating, it requires a runtime flag: --add-modules=jdk.incubator.vector We didn’t want correctness or availability to depend on that flag. So we designed the fallback behavior explicitly: At startup, we detect Vector API support and use the SIMD batched matmul when available; otherwise we fall back to an optimized scalar path, with single-video requests continuing to use the per-item implementation.</p><p>That gives us a clean operational story: services can opt in to the Vector API for maximum performance, but the system remains safe and predictable without it.</p><h4><strong>Results in Production:</strong></h4><p>With the full design in place with batching, flat buffers, ThreadLocal reuse, and the Vector API, we ran canaries that run production traffic. We observed a ~7% drop in CPU utilization and ~12% drop in average latency. To normalize across any small throughput differences, we also tracked CPU/RPS (CPU consumed per request-per-second). That metric improved by roughly 10%, meaning we could handle the same traffic with about 10% less CPU, and we saw similar numbers hold after full production rollout.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/768/1*qV3G1oUywCDqssuz6oTaJw.png"><figcaption>CPU/RPS on Ranker</figcaption></figure><p>At the function operator level, we saw the CPU drop from the initial 7.5% to a merely ~1% with the optimization in place. At the assembly level, the shift was clear: from loop-unrolled scalar dot products to a vectorized matrix multiply on AVX-512 hardware.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/742/1*2oKKZ-jKZrHTZP_vU-ppxg.png"><figcaption>Assembly snippet from batchEncode</figcaption></figure><h3><strong>Closing Thoughts</strong></h3><p>This optimization ended up being less about finding the “fastest library” and more about getting the fundamentals right: choosing the right computation shape, keeping data layout cache-friendly, and avoiding overheads that can erase theoretical wins. Once those pieces were in place, the JDK Vector API was a great fit, as it let us express SIMD-style math in pure Java, without JNI, while still keeping a safe fallback path. Another bonus was the low developer overhead: compared to lower-level approaches, the Vector API let us replace a much larger, more complex implementation with a relatively small amount of readable Java code, which made it easier to review, maintain, and iterate on.</p><p>Have you tried the Vector API in a real service yet? I’d love to hear what workloads it helped (or didn’t), and what you learned about benchmarking and rollout in production.</p><p><em>Special thanks to </em><a href="https://www.linkedin.com/in/jason-koch-5692172/"><em>Jason Koch</em></a><em>, </em><a href="https://www.linkedin.com/in/patrickstrawderman/"><em>Patrick Strawderman</em></a><em>, </em><a href="https://www.linkedin.com/in/yuewangh/"><em>Daniel Huang</em></a><em>, </em><a href="https://www.linkedin.com/in/fan-yang-a15ba249/"><em>Fan Yang</em></a><em>, and the Performance Engineering team at Netflix</em></p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=30d2830401ec" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/optimizing-recommendation-systems-with-jdks-vector-api-30d2830401ec">Optimizing Recommendation Systems with JDK’s Vector API</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/optimizing-recommendation-systems-with-jdks-vector-api-30d2830401ec</link>
      <guid>https://netflixtechblog.com/optimizing-recommendation-systems-with-jdks-vector-api-30d2830401ec</guid>
      <pubDate>Tue, 03 Mar 2026 02:36:00 +0100</pubDate>
    </item>
    <item>
      <title><![CDATA[Mount Mayhem at Netflix: Scaling Containers on Modern CPUs]]></title>
      <description><![CDATA[<p>Authors: <a href="https://www.linkedin.com/in/harshad-sane-56711a11/">Harshad Sane</a>, <a href="https://www.linkedin.com/in/andrew-halaney/">Andrew Halaney</a></p><p>Imagine this — you click play on Netflix on a Friday night and behind the scenes hundreds of containers spring to action in a few seconds to answer your call. At Netflix, scaling containers efficiently is critical to delivering a seamless streaming experience to millions of members worldwide. To keep up with responsiveness at this scale, we modernized our container runtime, only to hit a surprising bottleneck: the CPU architecture itself.</p><p>Let us walk you through the story of how we diagnosed the problem and what we learned about scaling containers at the hardware level.</p><h3>The Problem</h3><p>When application demand requires that we scale up our servers, we get a new instance from AWS. To use this new capacity efficiently, pods are assigned to the node until its resources are considered fully allocated. A node can go from no applications running to being maxed out within moments of being ready to receive these applications.</p><p>As we migrated more and more from our old container platform to our new container platform, we started seeing some concerning trends. Some nodes were stalling for long periods of time, with a simple health check timing out after 30 seconds. An initial investigation showed that the mount table length was increasing dramatically in these situations, and reading it alone could take upwards of 30 seconds. Looking at systemd’s stack it was clear that it was busy processing these mount events as well and could lead to complete system lockup. Kubelet also timed out frequently talking to containerd in this period. Examining the mount table made it clear that these mounts were related to container creation.</p><p>The affected nodes were almost all r5.metal instances, and were starting applications whose container image contained many layers (50+).</p><h3>Challenge</h3><h4>Mount Lock Contention</h4><p>The flamegraph in Figure 1 clearly shows where containerd spent its time. Almost all of the time is spent trying to grab a kernel-level lock as part of the various mount-related activities when assembling the container’s root filesystem!</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*1w_zMp68xhQI0W1bTXG0Dw.png"><figcaption>Figure 1: Flamegraph depicting lock contention</figcaption></figure><p>Looking closer, containerd executes the following calls for each layer if using user namespaces:</p><ol><li>open_tree() to get a reference to the layer / directory</li><li>mount_setattr() to set the idmap to match the container’s user range, shifting the ownership so this container can access the files</li><li>move_mount() to create a bind mount on the host with this new idmap applied</li></ol><p>These bind mounts are owned by the container’s user range and are then used as the lowerdirs to create the overlayfs-based rootfs for the container. Once the overlayfs rootfs is mounted, the bind mounts are then unmounted since they are not necessary to keep around once the overlayfs is constructed.</p><p>If a node is starting many containers at once, every CPU ends up busy trying to execute these mounts and umounts. The kernel VFS has various global locks related to the mount table, and each of these mounts requires taking that lock as we can see in the top of the flamegraph. Any system trying to quickly set up many containers is prone to this, and this is a function of the number of layers in the container image.</p><p>For example, assume a node is starting 100 containers, each with 50 layers in its image. Each container will need 50 bind mounts to do the idmap for each layer. The container’s overlayfs mount will be created using those bind mounts as the lower directories, and then all 50 bind mounts can be cleaned up via umount. Containerd actually goes through this process twice, once to determine some user information in the image and once to create the actual rootfs. This means the total number of mount operations on the start up path for our 100 containers is 100 * 2 * (1 + 50 + 50) = 20200 mounts, all of which require grabbing various global mount related locks!</p><h3>Diagnosis</h3><h4>What’s Different In The New Runtime?</h4><p>As alluded to in the introduction, Netflix has been undergoing a modernization of its container runtime. In the past a virtual kubelet + docker solution was used, whereas now a kubelet + containerd solution is being used. Both the old runtime and the new runtime used user namespaces, so what’s the difference here?</p><ol><li>Old Runtime:<br>All containers shared a single host user range. UIDs in image layers were shifted at untar time, so file permissions matched when containers accessed files. This worked because all containers used the same host user.</li><li>New Runtime:<br>Each container gets a unique host user range, improving security — if a container escapes, it can only affect its own files. To avoid the costly process of untarring and shifting UIDs for every container, the new runtime uses the kernel’s idmap feature. This allows efficient UID mapping per container without copying or changing file ownership, which is why containerd performs many mounts.</li></ol><p>Figure 2 below is a simplified example of how this idmap feature looks like:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*TYCK-IO0aLp0jhS8KUM6QQ.png"><figcaption>Figure 2: idmap feature</figcaption></figure><h4>Why Does Instance Type Matter?</h4><p>As noted earlier, the issue was predominantly occurring on r5.metal instances. Once we identified the root issue we could easily reproduce by creating a container image with many layers and sending hundreds of workloads using the image to a test node.</p><p>To better understand why this bottleneck was more profound on some instances compared to others, we benchmarked container launches on different AWS instance types:</p><ul><li>r5.metal (5th gen Intel, dual-socket, multiple NUMA domains)</li><li>m7i.metal-24xl (7th gen Intel, single-socket, single NUMA domain)</li><li>m7a.24xlarge (7th gen AMD, single-socket, single NUMA domain)</li></ul><h4>Baseline Results</h4><p>Figure 3 shows the baseline results from scaling containers on each instance type</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/648/1*KcdXLeml8rDOdnk2aukv1w.png"></figure><ul><li>At low concurrency (≤ ~20 containers), all platforms performed similarly</li><li>As concurrency increased, r5.metal began to fail around 100 containers</li><li>7th generation AWS instances maintained lower launch times and higher success rates as concurrency grew</li><li>m7a instances showed the most consistent scaling behavior with the lowest failure rates even at high concurrency</li></ul><h3>Deep Dive</h3><p>Using perf record and custom microbenchmarks, we can see the hottest code path was in the Linux kernel’s Virtual Filesystem (VFS) path lookup code — specifically, a tight spin loop waiting on a sequence lock in path_init(). The CPU spent most of its time executing the pause instruction, indicating many threads were spinning, waiting for the global lock, as shown in the disassembly snippet below</p><pre>path_init():<br>…<br>mov mount_lock,%eax<br>test $0x1,%al<br>je 7c<br>pause<br>…</pre><p>Using Intel’s Topdown Microarchitecture Analysis (<a href="https://dyninst.github.io/scalable_tools_workshop/petascale2018/assets/slides/TMA%20addressing%20challenges%20in%20Icelake%20-%20Ahmad%20Yasin.pdf">TMA</a>), we observed:</p><ul><li>95.5% of pipeline slots were stalled on contested accesses (tma_contested_accesses).</li><li>57% of slots were due to false sharing (multiple cores accessing the same cache line).</li><li>Cache line bouncing and lock contention were the primary culprits.</li></ul><p>Given a high amount of time being spent in contested accesses, the natural thinking from a perspective of hardware variations led to investigation of NUMA and Hyperthreading impact coming from the architecture to this subset</p><h4>NUMA Effects</h4><p>Non-Uniform Memory Access (NUMA) is a system design where each processor has its own local memory for faster access but relies on an interconnect to access the memory attached to a remote processor. Introduced in the 1990s to improve scalability in multiprocessor systems, NUMA boosts performance but also introduces higher latency when a CPU needs to access memory attached to another processor. Figure 4 is a simple image describing local vs remote access patterns of a NUMA architecture</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/802/1*sDCMnrdke_7aAOI-wf5rGQ.png"><figcaption>Figure 4:<em> Source: </em><a href="https://pmem.io/images/posts/numa_overview.png"><em>https://pmem.io/images/posts/numa_overview.png</em></a></figcaption></figure><p>AWS instances come in a variety of shapes and sizes. To obtain the largest core count, we tested the 2-socket 5th generation metal instances (r5.metal), on which containers were orchestrated by the titus agent. Modern dual-socket architectures implement NUMA design, leading to faster local but higher remote access latencies. Although container orchestration can maintain locality, global locks can easily run into high latency effects due to remote synchronization. In order to test the impact of NUMA, we tested an AWS 48xl sized instance with 2 NUMA nodes or sockets versus an AWS 24xl sized instance, which represents a single NUMA node or socket. As seen from Figure 5, the extra hop introduces high latencies and hence failures very quickly.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*TXZhTqAQHDu3NffCO15JrA.png"><figcaption>Figure 5: Numa Impact</figcaption></figure><h4>Hyperthreading Effects</h4><ul><li>Hyperthreading (HT): Disabling HT on m7i.metal-24xl (Intel) improved container launch latencies by 20–30% as seen in Figure 6, since hyperthreads compete for shared execution resources, worsening the lock contention. When hyperthreading is enabled, each physical CPU core is split into two logical CPUs (hyperthreads) that share most of the core’s execution resources, such as caches, execution units, and memory bandwidth. While this can improve throughput for workloads that are not fully utilizing the core, it introduces significant challenges for workloads that rely heavily on global locks. By disabling hyperthreading, each thread runs on its own physical core, eliminating this competition for shared resources between hyperthreads. As a result, threads can acquire and release global locks more quickly, reducing overall contention and improving latency for operations that generally share underlying resources.</li></ul><figure><img alt="" src="https://cdn-images-1.medium.com/max/623/1*9I-QZB66dEnBtLLuu23Ydw.png"><figcaption>Figure 6: Hyperthreading impact</figcaption></figure><h3>Why Does Hardware Architecture Matter?</h3><h4>Centralized Cache Architectures</h4><p>Some modern server CPUs use a mesh-style interconnect to link cores and cache slices, with each intersection managing cache coherence for a subset of memory addresses. In these designs, all communication passes through a central queueing structure, which can only handle one request for a given address at a time. When a global lock (like the mount lock) is under heavy contention, all atomic operations targeting that lock are funneled through this single queue, causing requests to pile up and resulting in memory stalls and latency spikes.</p><p>In some well-known mesh-based architectures as shown in Figure 7 below, this central queue is called the “Table of Requests” (TOR), and it can become a surprising bottleneck when many threads are fighting for the same lock. If you’ve ever wondered why certain CPUs seem to “pause for breath” under heavy contention, this is often the culprit.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/975/1*kmbGU_M8AK_VHgh4QLJp2A.png"><figcaption>Figure 7: <em>Public document from one of the major CPU vendors Source:</em><a href="https://www.intel.com/content/dam/developer/articles/technical/ddio-analysis-performance-monitoring/Figure1.png">https://www.intel.com/content/dam/developer/articles/technical/ddio-analysis-performance-monitoring/Figure1.png</a></figcaption></figure><h4>Distributed Cache Architectures</h4><p>Some modern server CPUs use a distributed, chiplet-based architecture (Figure 8), where multiple core complexes, each with their own local last-level cache — are connected via a high-speed interconnect fabric. In these designs, cache coherence is managed within each core complex, and traffic between complexes is handled by a scalable control fabric. Unlike mesh-based architectures with centralized queueing structures, this distributed approach spreads contention across multiple domains, making severe stalls from global lock contention less likely. For those interested in the technical details, public documentation from major CPU vendors provides deeper insight into these distributed cache and chiplet designs.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*0hQAquZ36l8_7Ib9aAnH0g.png"><figcaption>Figure 8: <em>Public document from one of the major CPU vendors, Source: (</em><a href="https://www.servethehome.com/amd-epyc-genoa-gaps-intel-xeon-in-stunning-fashion/amd-epyc-9004-genoa-chiplet-architecture-8x-ccd/">AMD EPYC 9004 Genoa Chiplet Architecture 8x CCD — ServeTheHome</a>)</figcaption></figure><p>Here is a comparison of the same workload run on m7i (centralized cache architecture) vs m7a (distributed cache architecture). Note that, in order to make it closely comparable, Hyperthreading (HT) was disabled on m7i, given previous regression seen in Figure 6, and experiments were run using same core counts. The result clearly shows a fairly consistent difference in performance of approximately 20% as shown in Figure 9</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*85aavnaRC6TAcBlCZgN6xA.png"><figcaption>Figure 9: Architectural impact between m7i and m7a</figcaption></figure><h4>Microbenchmark Results</h4><p>To prove the above theory related to NUMA, HT and micro-architecture, we developed a small <a href="https://github.com/Netflix/global-lock-bench">microbenchmark</a> which basically invokes a given number of threads that then spins on a globally contended lock. Running the benchmark at increasing thread counts reveals the latency characteristics of the system under different scenarios. For example, Figure 10 below is the microbenchmark results with NUMA, HT and different microarchitectures.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*bs0MDU5xNs9VIguhHhl0zA.png"><figcaption>Figure 10: Global lock contention benchmark results</figcaption></figure><p>Results from this custom synthetic benchmark (pause_bench) confirmed:</p><ul><li>On r5.metal, eliminating NUMA by only using a single socket significantly drops latency at high thread counts</li><li>On m7i.metal-24xl, disabling hyperthreading further improves scaling</li><li>On m7a.24xlarge, performance scales the best, demonstrating that a distributed cache architecture handles cache-line contention in this case of global locks more gracefully.</li></ul><h3>Improving Software Architecture</h3><p>While understanding the impacts of the hardware architecture is important for assessing possible mitigations, the root cause here is contention over a global lock. Working with containerd upstream we came to two possible solutions:</p><ol><li>Use the newer kernel mount API’s fsconfig() lowerdir+ support to supply the idmap’ed lowerdirs as fd’s instead of filesystem paths. This avoids the move_mount() syscall mentioned prior which requires global locks to mount each layer to the mount table</li><li>Map the common parent directory of all the layers. This makes the number of mount operations go from O(n) to O(1) per container, where n is the number of layers in the image</li></ol><p>Since using the newer API requires using a new kernel, we opted to make the latter <a href="https://github.com/containerd/containerd/pull/12092">change</a> to benefit more of the community. With that in place, no longer do we see containerd’s flamegraph being dominated by mount-related operations. In fact, as seen in Figure 11 below we had to highlight them in purple below to see them at all!</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*ingCksvvc5HrLd-LvMX7Dw.png"><figcaption>Figure 11: Optimized solution</figcaption></figure><h3><strong>Conclusion</strong></h3><p>Our journey migrating to a modern kubelet + containerd runtime at Netflix revealed just how deeply intertwined software and hardware architecture can be when operating at scale. While kubelet/containerd’s usage of unique container users brought significant security gains, it also surfaced new bottlenecks rooted in kernel and CPU architecture — particularly when launching hundreds of many layered container images in parallel. Our investigation highlighted that not all hardware is created equal for this workload: centralized cache management amplified cache contention while distributed cache design smoothly scaled under load.</p><p>Ultimately, the best solution combined hardware awareness with software improvements. For an immediate mitigation we chose to route these workloads to CPU architectures that scaled better under these conditions. By changing the software design to minimize per-layer mount operations, we eliminated the global lock as a launch-time bottleneck — unlocking faster, more reliable scaling regardless of the underlying CPU architecture. This experience underscores the importance of holistic performance engineering: understanding and optimizing both the software stack and the hardware it runs on is key to delivering seamless user experiences at Netflix scale.</p><p>We trust these insights will assist others in navigating the evolving container ecosystem, transforming potential challenges into opportunities for building robust, high-performance platforms.</p><p><em>Special thanks to the Titus and Performance Engineering teams at Netflix.</em></p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=f3b09b68beac" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/mount-mayhem-at-netflix-scaling-containers-on-modern-cpus-f3b09b68beac">Mount Mayhem at Netflix: Scaling Containers on Modern CPUs</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/mount-mayhem-at-netflix-scaling-containers-on-modern-cpus-f3b09b68beac</link>
      <guid>https://netflixtechblog.com/mount-mayhem-at-netflix-scaling-containers-on-modern-cpus-f3b09b68beac</guid>
      <pubDate>Sat, 28 Feb 2026 23:55:00 +0100</pubDate>
    </item>
    <item>
      <title><![CDATA[MediaFM: The Multimodal AI Foundation for Media Understanding at Netflix]]></title>
      <description><![CDATA[<p><a href="https://www.linkedin.com/in/avneesh/">Avneesh Saluja</a>, <a href="https://www.linkedin.com/in/santiagocastroserra/">Santiago Castro</a>, <a href="https://www.linkedin.com/in/bowei-yan-0080a326/">Bowei Yan</a>, <a href="https://www.linkedin.com/in/ashish-rastogi-11362a/">Ashish Rastogi</a></p><h4>Introduction</h4><p>Netflix’s core mission is to connect millions of members around the world with stories they’ll love. This requires not just an incredible catalog, but also a deep, machine-level understanding of every piece of content in that catalog, from the biggest blockbusters to the most niche documentaries. As we onboard new types of content such as live events and podcasts, the need to scalably understand these nuances becomes even more critical to our productions and member-facing experiences.</p><p>Many of these media-related tasks require sophisticated long-form video understanding e.g., identifying subtle narrative dependencies and emotional arcs that span entire episodes or films. <a href="https://netflixtechblog.com/detecting-scene-changes-in-audiovisual-content-77a61d3eaad6">Previous work</a> has found that to truly grasp the content’s essence, our models must leverage the full multimodal signal. For example, the audio soundtrack is a crucial, non-visual modality that can help more precisely identify clip-level tones or when a new scene starts. Can we use our collection of shows and movies to learn how to a) fuse modalities like audio, video, and subtitle text together and b) develop robust representations that leverage the narrative structure that is present in long form entertainment? Consisting of tens of millions of individual <a href="https://en.wikipedia.org/wiki/Shot_(filmmaking)">shots</a> across multiple titles, our diverse yet entertainment-specific dataset provides the perfect foundation to train multimodal media understanding models that enable many capabilities across the company such as <a href="https://netflixtechblog.com/mediafm-the-multimodal-ai-foundation-for-media-understanding-at-netflix-e8c28df82e2d#befc">ads relevancy, clip popularity prediction, and clip tagging</a>.</p><p>For these reasons, we developed the <strong>Netflix Media Foundational Model (MediaFM)</strong>, our new, in-house, multimodal content embedding model. MediaFM is the first tri-modal (audio, video, text) model pretrained on portions of the Netflix catalog. Its core is a multimodal, Transformer-based encoder designed to generate rich, contextual embeddings¹ for shots from our catalog by learning the temporal relationships between them through integrating visual, audio, and textual information. The resulting shot-level embeddings are powerful representations designed to create a deeper, more nuanced, and machine-readable understanding of our content, providing the critical backbone for effective cold start of newly launching titles in recommendations, optimized promotional assets (like art and trailers), and internal content analysis tools.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*y5IOwk46Pu82512T00ywGA.png"><figcaption>Figure 1: MediaFM Architecture</figcaption></figure><h4>Input Representation &amp; Preprocessing</h4><p>The model’s fundamental unit of input is a shot, derived by segmenting a movie or episode (collectively referred to as “title”) using a <a href="https://arxiv.org/abs/2008.04838">shot boundary detection</a> algorithm. For each shot, we generate three distinct embeddings from its core modalities:</p><ul><li><strong>Video</strong>: an internal model called <a href="https://netflixtechblog.com/building-in-video-search-936766f0017c">SeqCLIP</a> (a CLIP-style model fine-tuned on video retrieval datasets) is used to embed frames sampled at uniform intervals from segmented shots</li><li><strong>Audio</strong>: the audio samples from the same shots are embedded using Meta FAIR’s <a href="https://arxiv.org/abs/2006.11477">wav2vec2</a></li><li><strong>Timed Text</strong>: OpenAI’s text-embedding-3-large <a href="https://openai.com/index/new-embedding-models-and-api-updates/">model</a> is used to encode the corresponding timed text (e.g., closed captions, audio descriptions, or subtitles) for each shot</li></ul><p>For each shot, the three embeddings² are concatenated and unit-normed to form a single 2304-dimensional fused embedding vector. The transformer encoder is trained on sequences of shots, so each example in our dataset is a temporally-ordered sequence of these fused embeddings from the same movie or episode (up to 512 shots per sequence). We also have access to title-level metadata which is used to provide global context for each sequence (via the [GLOBAL]token). The title-level embedding is computed by passing title-level metadata (such as synopses and tags) through the text-embedding-3-large model.</p><h4>Model Architecture and Training Objective</h4><p>The core of our model is a transformer encoder, architecturally similar to BERT. A sequence of preprocessed shot embeddings is passed through the following stages:</p><ol><li><strong>Input Projection</strong>: The fused shot embeddings are first projected down to the model’s hidden dimension via a linear layer.</li><li><strong>Sequence Construction &amp; Special Tokens</strong>: Before entering the Transformer, two special embeddings are prepended to the sequence:<br>• a learnable [CLS] embedding is added at the very beginning.<br>• the title-level embedding is projected to the model’s hidden dimension and inserted after the [CLS] token as the [GLOBAL] token, providing title-level context to every shot in the sequence and participating in the self-attention process.</li><li><strong>Contextualization</strong>: The sequence is enhanced with positional embeddings and fed through the Transformer stack to provide shot representations based on their surrounding context.</li><li><strong>Output Projection</strong>: The contextualized hidden states from the Transformer are passed through a final linear layer, projecting them from the hidden layers back up to the 2304-dimensional fused embedding space for prediction.</li></ol><p>We train the model using a <strong>Masked Shot Modeling (MSM)</strong> objective. In this self-supervised task, we randomly mask 20% of the input shot embeddings in each sequence by replacing them with a learnable [MASK] embedding. The model’s objective is to predict the original, unmasked fused embedding for these masked positions. The model is optimized by minimizing the <strong>cosine distance</strong> between its predicted embedding and the ground-truth embedding for each masked shot.</p><p>We optimized the hidden parameters with Muon and the remaining parameters with AdamW. It’s worth noting that the switch to Muon resulted in noticeable improvements.</p><h4>Evaluation</h4><p>To evaluate the learned embeddings, we learn task-specific linear layers on top of frozen representations (i.e., linear probes). Most of the tasks are clip-level, i.e., each example is a short clip ranging from a few seconds to a minute which are often presented to our members while recommending a title to them. When embedding these clips, we find that “embedding in context”, namely extracting the embeddings from within a larger sequence (e.g., the episode containing the clip), naturally does much better than embedding only the shots from a clip.</p><h4><a href="https://netflixtechblog.com/mediafm-the-multimodal-ai-foundation-for-media-understanding-at-netflix-e8c28df82e2d#befc">Tasks</a></h4><p>Our embeddings are foundational and we find that they bring value to applications across Netflix. Here are a few:</p><ul><li><strong>Ad Relevancy:</strong> A multilabel classification task to categorize Netflix clips for relevant ad placement, measured by <strong>Average Precision</strong>. In this task, these representations operate at the retrieval stage, where they help in identifying the candidate set and in turn are fed into the ad serving system for relevance optimization.</li><li><strong>Clip Popularity Ranking:</strong> A ranking task to predict the relative performance (in click-through rate, <a href="https://en.wikipedia.org/wiki/Click-through_rate">CTR</a>) of a media clip relative to other clips from that show or movie, measured by a ten-fold with <strong>Kendall’s tau correlation coefficient</strong>.</li><li><strong>Clip Tone:</strong> A multi-label classification of hook clips into 100 tone categories (e.g., creepy, scary, humorous) from our internal Metadata &amp; Ratings team, measured by <strong>micro Average Precision </strong>(averaged across tone categories).</li><li><strong>Clip Genre:</strong> A multi-label classification of clips into eleven core genres (Action, Anime, Comedy, Documentary, Drama, Fantasy, Horror, Kids, Romance, Sci-fi, Thriller) derived from the genre of the parent title, measured by <strong>macro Average Precision </strong>(averaged across genres).</li><li><strong>Clip Retrieval: </strong>a binary classification of clips from movies or episodes into “clip-worthy” (i.e., a good clip to showcase the title) or not, as determined by human annotators, and as measured by <strong>Average Precision</strong>. The positive to negative clip ratio is 1:3, and for each title we select 6–10 positive clips and the corresponding number of negatives.</li></ul><p>It’s worth noting that for the tasks above (as well as other tasks that use our model), the model outputs are utilized as information that the relevant teams use when driving to a decision rather than being used in a completely end-to-end fashion. Many of the improvements are also in various stages of deployment.</p><h4>Results</h4><p>Figure 2³ compares MediaFM to several strong baselines:</p><ul><li>The previously mentioned SeqCLIP, which also provides the video embedding input for MediaFM</li><li>Google’s <a href="https://docs.cloud.google.com/vertex-ai/generative-ai/docs/embeddings/get-multimodal-embeddings">VertexAI multimodal embeddings</a></li><li>TwelveLabs’ <a href="https://www.twelvelabs.io/blog/introducing-marengo-2-7">Marengo 2.7 embeddings</a></li></ul><figure><img alt="" src="https://cdn-images-1.medium.com/max/984/1*aydFwELbXQSMBRWauB3IKg.png"><figcaption>Figure 2: Performance of MediaFM vs. external and internal models.</figcaption></figure><p>On all tasks, MediaFM is better than the baselines. Improvements seem to be larger for tasks that require more detailed narrative understanding e.g., predicting the most relevant ads for an ad break given the surrounding context. We look further into this next.</p><h4>Ablations</h4><p>MediaFM’s primary improvements over previous Netflix work stem from two key areas: combining multiple modalities and learning to contextualize shot representations. To determine the contribution of each factor across different tasks, we compared MediaFM to a baseline. This baseline concatenates the three input embeddings, essentially providing the same complete, shot-level input as MediaFM but without the contextualization step. This comparison allows us to isolate which tasks benefit most from the contextualization aspect.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*sLEMAVa5NAra9h5WvG4D2g.png"></figure><p>Additional modalities help somewhat for tone but the main improvement comes from contextualization.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*mWYcJIXT_lPBkcTHpvQwGA.png"></figure><p>Oddly, multiple uncontextualized modalities <strong>hurts </strong>the clip popularity ranking model, but adding contextualization significantly improves performance.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*y0MEiCxlpGmvIIgxjmarVg.png"></figure><p>For clip retrieval we see a natural progression of around 15% for each improvement.</p><h4>Next Steps</h4><p>MediaFM presents a way to learn how to fuse and/or contextualize shot-level information by leveraging Netflix’s catalog in a self-supervised manner. With this perspective, we are actively investigating how pretrained multimodal (audio, video/image, text) LLMs like <a href="https://qwen.ai/blog?id=65f766fc2dcba7905c1cb69cc4cab90e94126bf4&amp;from=research.latest-advancements-list">Qwen3-Omni</a>, where the modality fusion has already been learned, can provide an even stronger starting point for subsequent model generations.</p><p>Next in this series of blog posts, we will present our method to embed title-level metadata and adapt it to our needs. Stay tuned!</p><h4>Footnotes</h4><ol><li>We chose embeddings over generative text outputs to prioritize modular design. This provides a tighter, cleaner abstraction layer: we generate the representation once, and it is consumed across our entire suite of services. This avoids the architectural fragility of fine-tuning, allowing us to enhance our existing embedding-based workflows with new modalities more flexibly.</li><li>All of our data has audio and video; we zero-pad for missing timed text data, which is relatively likely to occur (e.g., in shots without dialogue).</li><li>The title-level tasks couldn’t be evaluated with the VertexAI MM and Marengo embedding models as the videos exceed the length limit set by the APIs.</li></ol><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=e8c28df82e2d" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/mediafm-the-multimodal-ai-foundation-for-media-understanding-at-netflix-e8c28df82e2d">MediaFM: The Multimodal AI Foundation for Media Understanding at Netflix</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/mediafm-the-multimodal-ai-foundation-for-media-understanding-at-netflix-e8c28df82e2d</link>
      <guid>https://netflixtechblog.com/mediafm-the-multimodal-ai-foundation-for-media-understanding-at-netflix-e8c28df82e2d</guid>
      <pubDate>Mon, 23 Feb 2026 19:24:00 +0100</pubDate>
    </item>
    <item>
      <title><![CDATA[Scaling LLM Post-Training at Netflix]]></title>
      <description><![CDATA[<p><a href="https://www.linkedin.com/in/baolin-li-659426115/">Baolin Li</a>, <a href="https://www.linkedin.com/in/lingyi-liu-4b866016/">Lingyi Liu</a>, <a href="https://www.linkedin.com/in/binh-tang-3b76557b/">Binh Tang</a>, <a href="https://www.linkedin.com/in/shaojingli/">Shaojing Li</a></p><h3>Introduction</h3><p>Pre-training gives Large Language Models (LLMs) broad linguistic ability and general world knowledge, but post-training is the phase that actually aligns them to concrete intents, domain constraints, and the reliability requirements of production environments. At Netflix, we are exploring how LLMs can enable new member experiences across recommendation, personalization, and search, which requires adapting generic foundation models so they can better reflect our catalog and the nuances of member interaction histories. At Netflix scale, post-training quickly becomes an engineering problem as much as a modeling one: building and operating complex data pipelines, coordinating distributed state across multi-node GPU clusters, and orchestrating workflows that interleave training and inference. This blog describes the architecture and engineering philosophy of our internal <strong>Post-Training Framework</strong>, built by the AI Platform team to hide infrastructure complexity so researchers and model developers can focus on model innovation — not distributed systems plumbing.</p><h3>A Model Developer’s Post-Training Journey</h3><p>Post-training often starts deceptively simply: curate proprietary domain data, load an open-weight model from Hugging Face, and iterate batches through it. At the experimentation scale, that’s a few lines of code. But when fine-tuning production-grade LLMs at scale, the gap between “running a script” and “robust post-training” becomes an abyss of engineering edge cases.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/873/1*at60qZXd0j6SphkjzdmazQ.png"><figcaption>Figure 1. Simple steps to post-train an open-weight model.</figcaption></figure><h4>Getting the data right</h4><p>On paper, post-training is straightforward: choose a tokenizer, preprocess the dataset, and build a dataloader. In practice, data preparation is where things break. High-quality post-training — instruction following, multi-turn dialogue, Chain-of-Thought — depends on precisely controlling which tokens contribute to the loss. Hugging Face chat templates serialize conversations, but don’t specify what to train on versus ignore. The pipeline must apply explicit loss masking so only assistant tokens are optimized; otherwise the model learns from prompts and other non-target text, degrading quality.</p><p>Variable sequence length is another pitfall. Padding within a batch can waste compute, and uneven shapes across FSDP workers can cause GPU synchronization overhead. A more GPU-efficient approach is to pack multiple samples into fixed-length sequences and use a “document mask” to prevent cross-attention across samples, reducing padding and keeping shapes consistent.</p><h4>Setting up the model</h4><p>Loading an open-source checkpoint sounds simple until the model no longer fits on one GPU. At that point you need a sharding strategy (e.g., FSDP, TP) and must load partial weights directly onto the device mesh to avoid ever materializing the full model on a single device.</p><p>After loading, you still need to make the model trainable: choose full fine-tuning vs. LoRA, and apply optimizations like activation checkpointing, compilation, and correct precision settings (often subtle for RL, where rollout and policy precision must align). Large vocabularies (&gt;128k) add a further memory trap: logits are<em> [batch, seq_len, vocab] </em>and can spike peak memory. Common mitigations include dropping ignored tokens before projection and computing logits/loss in chunks along the sequence dimension.</p><h4>Starting the training</h4><p>Even with data and models ready, production training is not a simple “for loop”. The system must support everything from SFT’s forward/backward pass to on-policy RL workflows that interleave rollout generation, reward/reference inference, and policy updates.</p><p>At Netflix scale, training runs as a distributed job. We use Ray to orchestrate workflows via actors, decoupling modeling logic from hardware. Robust runs also require experiment tracking (model quality metrics like loss and efficiency metrics like MFU) and fault tolerance via standardized checkpoints to resume cleanly after failures.</p><p>These challenges motivate a post-training framework that lets developers focus on modeling rather than distributed systems and operational details.</p><h3>The Netflix Post-Training Framework</h3><p>We built Netflix’s LLM post-training framework so Netflix model developers can turn ideas like those in Figure 1 into scalable, robust training jobs. It addresses the engineering hurdles described above, and also constraints that are specific to the Netflix ecosystem. Existing tools (e.g., Thinking Machines’ <a href="https://thinkingmachines.ai/tinker/">Tinker</a>) work well for standard chat and instruction-tuning, but their structure can limit deeper experimentation. In contrast, our internal use cases often require architectural variation (for example, customizing output projection heads for task-specific objectives), expanded or nonstandard vocabularies driven by semantic IDs or special tokens, and even transformer models pre-trained from scratch on domain-specific, non-natural-language sequences. Supporting this range requires a framework that prioritizes flexibility and extensibility over a fixed fine-tuning paradigm.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*z2IFH-4iIQ5qxARTur-wQA.png"><figcaption>Figure 2. The post-training library within Netflix stack</figcaption></figure><p>Figure 2 shows the end-to-end stack from infrastructure to trained models. At the base is Mako, Netflix’s internal ML compute platform, which provisions GPUs on AWS. On top of Mako, we run robust open-source components — PyTorch, Ray, and vLLM — largely out of the box. Our post-training framework sits above these foundations as a library: it provides reusable utilities and standardized training recipes for common workflows such as Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), Reinforcement Learning (RL), and Knowledge Distillation. Users typically express jobs as configuration files that select a recipe and plug in task-specific components.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*RzZzr3wADPwism5CQP8mYw.png"><figcaption>Figure 3. Main components developed for the post-training framework</figcaption></figure><p>Figure 3 summarizes the modular components we built to reduce complexity across four dimensions. As with most ML systems, training success hinges on three pillars — <strong>Data</strong>, <strong>Model</strong>, and <strong>Compute</strong> — and the rise of RL fine-tuning adds a fourth pillar: <strong>Workflow</strong>, to support multi-stage execution patterns that don’t fit a simple training loop. Below, we detail the specific abstractions and features the framework provides for each of these dimensions:</p><ul><li><strong>Data:</strong> Dataset abstractions for SFT, reward modeling, and RL; high-throughput streaming from cloud and disk for datasets that exceed local storage; and asynchronous, on-the-fly sequence packing to overlap CPU-heavy packing with GPU execution and reduce idle time.</li><li><strong>Model:</strong> Support for modern architectures (e.g., Qwen3, Gemma3) and Mixture-of-Experts variants (e.g., Qwen3 MoE, GPT-OSS); LoRA integrated into model definitions; and high-level sharding APIs so developers can distribute large models across device meshes without writing low-level distributed code.</li><li><strong>Compute:</strong> A unified job submission interface that scales from a single node to hundreds of GPUs; MFU (Model FLOPS Utilization) monitoring that remains accurate under custom architectures and LoRA; and comprehensive checkpointing (states of trained parameters, optimizer, dataloader, data mixer, etc.) to enable exact resumption after interruptions.</li><li><strong>Workflow:</strong> Support for training paradigms beyond SFT, including complex online RL. In particular, we extend Single Program, Multiple Data (SPMD) style SFT workloads to run online RL with a hybrid single-controller + SPMD execution model, which we’ll describe next.</li></ul><p>Today, this framework supports research use cases ranging from post-training large-scale foundation models to fine-tuning specialized expert models. By standardizing these workflows, we’ve lowered the barrier for teams to experiment with advanced techniques and iterate more quickly.</p><h3>Learnings from Building the Post-Training Framework</h3><p>Building a system of this scope wasn’t a linear implementation exercise. It meant tracking a fast-moving open-source ecosystem, chasing down failure modes that only appear under distributed load, and repeatedly revisiting architectural decisions as the post-training frontier shifted. Below are three engineering learnings and best practices that shaped the framework.</p><h4>Scaling from SFT to RL</h4><p>We initially designed the library around Supervised Fine-Tuning (SFT): relatively static data flow, a single training loop, and a Single Program, Multiple Data (SPMD) execution model. That assumption stopped holding in 2025. With DeepSeek-R1 and the broader adoption of efficient on-policy RL methods like GRPO, SFT became table stakes rather than the finish line. Staying close to the frontier required infrastructure that could move from “offline training loop” to “multi-stage, on-policy orchestration.”</p><p>SFT’s learning signal is dense and immediate: for each token position we compute logits over the full vocabulary and backpropagate a differentiable loss. Infrastructure-wise, this looks a lot like pre-training and maps cleanly to SPMD — every GPU worker runs the same step function over a different shard of data, synchronizing through Pytorch distributed primitives.</p><p>On-policy RL changes the shape of the system. The learning signal is typically sparse and delayed (e.g., a scalar reward at the end of an episode), and the training step depends on data generated by the current policy. Individual sub-stages — policy updates, rollout generation, reference model inference, reward model scoring — can each be implemented as SPMD workloads, but the end-to-end algorithm needs explicit coordination: you’re constantly handing off artifacts (prompts, sampled trajectories, rewards, advantages) across stages and synchronizing their lifecycle.</p><p>In our original SFT architecture, the driver node was intentionally “thin”: it launched N identical Ray actors, each encapsulating the full training loop, and scaling meant launching more identical workers. That model breaks down for RL. RL required us to decompose the system into distinct roles — Policy, Rollout Workers, Reward Model, Reference Model, etc. — and evolve the driver into an active controller that encodes the control plane: when to generate rollouts, how to batch and score them, when to trigger optimization, and how to manage cluster resources across phases.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*5lhd3rDmexD0KGoHy78CZQ.png"><figcaption>Figure 4. Architectural differences of SFT and RL framework</figcaption></figure><p>Figure 4 highlights this shift. To add RL support without reinventing distributed orchestration from scratch, we integrated the core infrastructure from the open-source <a href="https://github.com/verl-project/verl"><strong>Verl</strong></a> library to manage Ray actor lifecycle and GPU resource allocation. Leveraging Verl’s backend let us focus on the “modeling surface area” — our Data/Model/Compute abstractions and internal optimizations — while keeping orchestration concerns decoupled. The result is a hybrid design: a unified user interface where developers can move between SFT and RL workflows without adopting an entirely different mental model or API set.</p><h4>Hugging Face-Centric Experience</h4><p>The Hugging Face Hub has effectively become the default distribution channel for open-weight LLMs, tokenizers, and configs. We designed the framework to stay close to that ecosystem rather than creating an isolated internal standard. Even when we use optimized internal model representations for speed, we load and save checkpoints in standard Hugging Face formats. This avoids “walled garden” friction and lets teams pull in new architectures, weights, and tokenizers quickly.</p><p>This philosophy also shaped our tokenizer story. Early on, we bound directly to low-level tokenization libraries (e.g., SentencePiece, tiktoken) to maximize control. In practice, that created a costly failure mode: silent training–serving skew. Our inference stack (vLLM) defaults to Hugging Face AutoTokenizer, and tiny differences in normalization, special token handling, or chat templating can yield different token boundaries — exactly the kind of mismatch that shows up later as inexplicable quality regressions. We fixed this by making Hugging Face AutoTokenizer the single source of truth. We then built a thin compatibility layer (BaseHFModelTokenizer) to handle post-training needs — setting padding tokens, injecting generation markers to support loss masking, and managing special tokens / semantic IDs — while ensuring the byte-level tokenization path matches production.</p><p>We do take a different approach for model implementations. Rather than training directly on transformers model classes, we maintain our own optimized, unified model definitions that can still load/save Hugging Face checkpoints. This layer is what enables framework-level optimizations — e.g., FlexAttention, memory-efficient chunked cross-entropy, consistent MFU accounting, and uniform LoRA extensibility — without re-implementing them separately for every model family. A unified module naming convention also makes it feasible to programmatically locate and swap components (Attention, MLP, output heads) across architectures, and provides a consistent surface for Tensor Parallelism and FSDP wrapping policies.</p><p>The trade-off is clear: supporting a new model family requires building a bridge between the Hugging Face reference implementation and our internal definition. To reduce that overhead, we use AI coding agents to automate much of the conversion work, with a strict <strong>logit verifier</strong> as the gate: given random inputs, our internal model must match the Hugging Face logits within tolerance. Because the acceptance criterion is mechanically checkable, agents can iterate autonomously until the implementation is correct, dramatically shortening the time-to-support for new architectures.</p><p>Today, this design means we can only train architectures we explicitly support — an intentional constraint shared by other high-performance systems like <a href="https://huggingface.co/docs/transformers/main/transformers_as_backend">vLLM, SGLang</a>, and <a href="https://github.com/pytorch/torchtitan/pull/2048">torchtitan</a>. To broaden coverage, we plan to add a fallback Hugging Face backend, similar to the compatibility patterns these projects use: users will be able to run training directly on native transformers models for rapid exploration of novel architectures, with the understanding that some framework optimizations and features may not apply in that mode.</p><h4>Providing Differential Value</h4><p>A post-training framework is only worth owning if it delivers clear value beyond assembling OSS components. We build on open source for velocity, but we invest heavily where off-the-shelf tools tend to be weakest: performance tuned to our workload characteristics, and integration with Netflix-specific model and business requirements. Here are some concrete examples:</p><p>First, we optimize training efficiency for our real use cases. A representative example is extreme variance in sequence length. In FSDP-style training, long-tail sequences create stragglers: faster workers end up waiting at synchronization points for the slowest batch, lowering utilization. Standard bin-packing approaches help, but doing them offline at our data scale can add substantial preprocessing latency and make it harder to keep datasets fresh. Instead, we built on-the-fly sequence packing that streams samples from storage and dynamically packs them in memory. Packing runs asynchronously, overlapping CPU work with GPU compute. Figure 5 shows the impact: for our most skewed dataset, on-the-fly packing improved the effective token throughput by up to 4.7x.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*Z2mdtXVFJsr764NihkguIA.png"><figcaption>Figure 5. Training throughput on two of our internal datasets on A100 and H200 GPUs</figcaption></figure><p>We also encountered subtler performance cliffs around vocabulary expansion. Our workloads frequently add custom tokens and semantic IDs. We found that certain vocabulary sizes could cause the language model head to fall back from a highly optimized cuBLAS kernel to a much slower CUTLASS path, tripling that layer’s execution time. The framework now automatically pads vocabulary sizes to multiples of 64 so the compiler selects the fast kernel, preserving throughput without requiring developers to know these low-level constraints.</p><p>Second, owning the framework lets us support “non-standard” transformer use cases that generic LLM tooling rarely targets. For example, some internal models are trained on member interaction event sequences rather than natural language, and may require bespoke RL loops that integrate with highly-customized inference engines and optimize business-defined metrics. These workflows demand custom environments, reward computation, and orchestration patterns — while still needing the same underlying guarantees around performance, tracking, and fault tolerance. The framework is built to accommodate these specialized requirements without fragmenting into one-off pipelines, enabling rapid iteration.</p><h3>Wrap up</h3><p>Building the Netflix Post-Training Framework has been a continual exercise in balancing standardization with specialization. By staying anchored to the open-source ecosystem, we’ve avoided drifting into a proprietary stack that diverges from where the community is moving. At the same time, by owning the core abstractions around Data, Model, Compute, and Workflow, we’ve preserved the freedom to optimize for Netflix-scale training and Netflix-specific requirements.</p><p>In the process, we’ve moved post-training from a loose collection of scripts into a managed, scalable system. Whether the goal is maximizing SFT throughput, orchestrating multi-stage on-policy RL, or training transformers over member interaction sequences, the framework provides a consistent set of primitives to do so reliably and efficiently. As the field shifts toward more agentic, reasoning-heavy, and multimodal architectures, this foundation will help us translate new ideas into scalable GenAI prototypes — so experimentation is constrained by our imagination, not by operational complexity.</p><h3>Acknowledgements</h3><p>This work builds on the momentum of the broader open-source ML community. We’re especially grateful to the teams and contributors behind Torchtune, Torchtitan, and Verl, whose reference implementations and design patterns informed many of our training framework choices — particularly around scalable training recipes, distributed execution, and RL-oriented orchestration. We also thank our partner teams in Netflix AI for Member Systems for close collaboration, feedback, and shared problem-solving throughout the development and rollout of the Post-Training Framework, and the Training Platform team for providing the robust infrastructure and operational foundation that makes large-scale post-training possible.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=0046f8790194" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/scaling-llm-post-training-at-netflix-0046f8790194">Scaling LLM Post-Training at Netflix</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/scaling-llm-post-training-at-netflix-0046f8790194</link>
      <guid>https://netflixtechblog.com/scaling-llm-post-training-at-netflix-0046f8790194</guid>
      <pubDate>Fri, 13 Feb 2026 09:05:00 +0100</pubDate>
    </item>
    <item>
      <title><![CDATA[Automating RDS Postgres to Aurora Postgres Migration]]></title>
      <description><![CDATA[<p><a href="https://www.linkedin.com/in/ramsrivatsa/">Ram Srivasta Kannan,</a> <a href="https://www.linkedin.com/in/wale-akintayo-30782a82/">Wale Akintayo</a>, <a href="https://www.linkedin.com/in/jay-bharadwaj-4b310ab8/">Jay Bharadwaj</a>, <a href="https://www.linkedin.com/in/john-crimmins-39730b3a/">John Crimmins</a>, <a href="https://www.linkedin.com/in/shengwei4721/">Shengwei Wang,</a> <a href="https://www.linkedin.com/in/zhitao-cathy-zhu/">Zhitaou Zhu</a></p><h3>Introduction</h3><p>In 2024, the Online Data Stores team at Netflix conducted a comprehensive review of the relational database technologies used across the company. This evaluation examined functionality, performance, and total cost of ownership across our database ecosystem. Based on this analysis, we decided to standardize on <strong>Amazon Aurora PostgreSQL as the primary relational database</strong> offering for Netflix teams.</p><p>Several key factors influenced this decision:</p><ul><li><strong>PostgreSQL already </strong>underpinned<strong> </strong>the majority of our relational workloads, which made it a natural foundation for standardization. Internal evaluations revealed that Aurora PostgreSQL had supported over 95% of the applications and workloads running on other relational databases across our internal services.</li><li><strong>Industry momentum had continued to shift toward PostgreSQL, </strong>driven by its open ecosystem, strong community support, and broad adoption across modern data platforms.</li><li><strong>Aurora’s cloud-native, distributed architecture</strong> provided clear advantages in scalability, high availability, and elasticity compared to traditional single-node PostgreSQL deployments.</li><li>Aurora PostgreSQL offered a <strong>rich feature set</strong>, along with a <strong>strong, forward-looking roadmap</strong> aligned with the needs of large-scale, globally distributed applications.</li></ul><h3>A Clear Migration Path Forward</h3><p>As part of this strategic shift, one of our key initiatives for 2024/2025 was migrating existing users to Aurora PostgreSQL. This effort began with RDS PostgreSQL migrations and will expand to include migrations from other relational systems in subsequent phases.</p><p>As a data platform organization, our goal is to make this evolution predictable, well-supported, and minimally disruptive. This allows teams to adopt Aurora PostgreSQL at a pace that aligns with their product and operational roadmaps, while we move toward a unified and scalable relational data platform across the organization.</p><h3>Database Migration: More Than a Simple Transfer</h3><p>Migrating a database involves far more than copying rows from one system to another. It is a coordinated process of transitioning both data and database functionality while preserving correctness, availability, and performance. At scale, a well-designed migration must minimize disruption to applications and ensure a clean, deterministic handoff from the old system to the new one.</p><p>Most database migrations follow a common set of high-level steps:</p><ol><li><strong>Data Replication</strong>: Data is first copied from the source database to the destination, typically using replication, so that ongoing changes are continuously captured and applied.</li><li><strong>Quiescence</strong>: Write traffic to the source database is halted, allowing the destination to fully catch up and eliminate any remaining divergence.</li><li><strong>Validation</strong>: The system verifies that the source and destination databases are fully synchronized and contain identical data.</li><li><strong>Cutover</strong>: Client applications are reconfigured to point to the destination database, which becomes the new primary source of truth.</li></ol><h3>Challenges</h3><h4>Operational Challenges</h4><p>Migrating to a new relational database at Netflix scale presents substantial operational challenges. With a fleet approaching 400 PostgreSQL clusters, manually migrating each one is simply not scalable for the data platform team. Such an approach would require a significant amount of time, introduce the risk of human error, and necessitate considerable hands-on engineering effort. Compounding the problem, coordinating downtime across the many interconnected services that depend on each database is extremely cumbersome at this scale.</p><p>To address these challenges, we designed a self-service migration workflow that enables service owners to run their own RDS PostgreSQL to Aurora PostgreSQL migrations. The workflow automatically handles orchestration, safety checks, and correctness guarantees end-to-end, resulting in lower operational overhead and a predictable, reliable migration experience.</p><h3>Technical challenges</h3><ul><li><strong>Zero data loss</strong> — We must guarantee that all data from the source cluster is fully and safely migrated to the destination within a very tight window, with no possibility of data loss.</li><li><strong>Minimal downtime — </strong>Some downtime is unavoidable during migration, as applications must briefly pause write traffic while cutting over to Aurora PostgreSQL. For higher-tier services that power critical parts of the Netflix ecosystem, this window must be kept extremely short to prevent user-facing impact and maintain service reliability.</li><li><strong>No control over client applications</strong> — As the platform team, we manage the databases, but application teams handle the read and write operations. We cannot assume that they have the ability to pause writes on demand, nor do we want to expose such controls to them, as mistakes could lead to data inconsistencies post migration. Therefore, building a self-service migration pipeline requires creative control-plane solutions to halt traffic, ensuring that no writes occur during the validation and cutover phases.</li><li><strong>No direct access to RDS credentials</strong> — The migration automation must perform replication, quiescence, and validation without requesting database credentials from users or relying on manual authentication. Source databases are often tightly secured, allowing access only from client applications, but more importantly, requiring credential access — even if it were possible — would significantly increase operational overhead and risk. At the same time, the migration platform may operate in environments without direct access to the source database, making traditional verification or parity checks impossible.</li><li><strong>No Degradation in Performance</strong> — The migration process must not impact the performance or stability of production databases once they are running in the Aurora PostgreSQL ecosystem.</li><li><strong>Full Ecosystem Parity</strong> — Beyond migrating the core database, associated components such as parameter groups, read replicas, and replication slots must also be migrated to ensure functional equivalence.</li></ul><p><strong>Minimal User Effort</strong> — Since we rely on teams who are not database experts to perform migrations, the process must be simple, intuitive, and fully self-guided.</p><h3>AWS recommended migration techniques</h3><h4>Using a snapshot</h4><p>One of the simplest AWS-recommended approaches for migrating from RDS PostgreSQL to Aurora PostgreSQL is based on snapshots. In this model, write traffic to the source PostgreSQL database is first stopped. A manual snapshot of the RDS PostgreSQL instance is then taken and migrated to Aurora, where AWS converts it into an Aurora-compatible format.<br> <br>Once the conversion completes, a new Aurora PostgreSQL cluster is created from the snapshot. After the cluster is brought online and validated, application traffic is redirected to the Aurora endpoint, completing the migration.</p><p><a href="https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide/AuroraPostgreSQL.Migrating.RDSPostgreSQL.Import.Console.html">Reference</a></p><h4>Using an Aurora read replica</h4><p>In the read-replica–based approach, an Aurora PostgreSQL read replica is created from an existing RDS PostgreSQL instance. AWS establishes continuous, asynchronous replication from the RDS source to the Aurora replica, allowing ongoing changes to be streamed in near real time.</p><p>Because replication runs continuously, the Aurora replica remains closely synchronized with the source database. This enables teams to provision and validate the Aurora environment — including configuration, connectivity, and performance characteristics — while production traffic continues to flow to the source.</p><p>When the replication lag is sufficiently low, write traffic is briefly paused to allow the replica to fully catch up. The Aurora read replica is then promoted to a standalone Aurora PostgreSQL cluster, and application traffic is redirected to the new Aurora endpoint. This approach significantly reduces downtime compared to snapshot-based migrations and is well-suited for production systems that require minimal disruption.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/647/1*_xydKq_VWpaxyqcwCwQyuA.png"><figcaption>Migration Strategy Trade-Offs</figcaption></figure><p>These differences represent the key considerations when choosing a migration strategy from RDS PostgreSQL to Aurora PostgreSQL. For our automation, we opted for the Aurora Read Replica approach, trading increased implementation complexity for a significantly shorter downtime window for client applications.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/682/1*SALMQ9wXHsJY9kBTmZ1D5w.png"><figcaption>Netflix RDS PostgreSQL Deployment Architecture</figcaption></figure><p>In Netflix’s RDS setup, a <a href="https://netflixtechblog.medium.com/data-gateway-a-platform-for-growing-and-protecting-the-data-tier-f1ed8db8f5c6">Data Access Layer</a> (DAL) sits between applications and backend databases, acting as middleware that centralizes database connectivity, security, and traffic routing on behalf of client applications.</p><p>On the client side, applications connect through a forward proxy that manages mutual TLS (mTLS) authentication and establishes a secure tunnel to the Data Gateway service. The Data Gateway, acting as a reverse proxy for database servers, terminates client connections, enforces centralized authentication and authorization, and forwards traffic to the appropriate RDS PostgreSQL instance.</p><p>This layered design ensures that applications never handle raw database credentials, provides a consistent and secure access pattern across all datastore types, and delivers isolated, transparent connectivity to managed PostgreSQL clusters. While the primary goal of this architecture is to enforce strong security controls and standardize how applications access external AWS data stores, it also allows backend databases to be switched transparently via configuration, enabling controlled, low-downtime migrations.</p><h3>Migration Process</h3><p>The Platform team’s goal is to deliver a fully automated, self-service workflow that helps with the migration of customer RDS PostgreSQL instances to Aurora PostgreSQL clusters. This migration tool orchestrates the entire process — from preparing the source environment, initializing the Aurora read replica, and maintaining continuous synchronization, all the way through to cutover — without requiring any database credentials or manual intervention from the customer.</p><p>Designed for minimal downtime and seamless user experience, the workflow ensures full ecosystem parity between RDS and Aurora, preserving performance characteristics and operational behavior while enabling customers to benefit from Aurora’s improved scalability, resilience, and cost efficiency.</p><h3>Data Replication Phase</h3><h4>Enable Automated Backups</h4><p>Automated backups must be enabled on the source database because the Aurora <a href="https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/USER_PostgreSQL.Replication.ReadReplicas.Configuration.html?utm_source=chatgpt.com">read replica</a> is initialized from a consistent snapshot of the source and then kept in sync through continuous replication. Automated backups provide the stable snapshot required to bootstrap the replica, along with the continuous streaming of write-ahead log (WAL) records needed to keep the read replica closely synchronized with the source.</p><h4>Port RDS parameters to an Aurora parameter group</h4><p>We create a dedicated Aurora parameter group for each cluster and migrate all RDS-compatible parameters from the source RDS instance. This ensures that the Aurora cluster inherits the same configuration settings — such as memory configuration, connection limits, query planner behavior, and other PostgreSQL engine parameters that have equivalents in Aurora. Parameters that are unsupported or behave differently in Aurora are either omitted or adjusted according to Aurora best practices.</p><h4>Create an Aurora read replica cluster and instance</h4><p>Creating an Aurora read replica cluster is a critical step in migrating from RDS PostgreSQL to Aurora PostgreSQL. At this stage, the Aurora cluster is created and attached to the RDS PostgreSQL primary as a replica, establishing continuous replication from the source RDS PostgreSQL instance. These Aurora read replicas stay nearly in sync with ongoing changes by streaming write-ahead logs (WAL) from the source, enabling minimal downtime during cutover. The cluster is fully operational for validation and performance testing, but it is not yet writable — RDS remains the authoritative primary.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*a3JN0YW1wuYTuez8oWw5SA.png"></figure><h4>Quiescence Phase</h4><p>The goal of the quiescence phase is to transition client applications from the source RDS PostgreSQL instance to the Aurora PostgreSQL cluster as the new primary database, while preserving data consistency during cutover.</p><p>The first step in this process is to stop all write traffic to the source RDS PostgreSQL instance to guarantee consistency. To achieve this, we instruct users to halt application-level traffic, which helps prevent issues such as retry storms, queue backlogs, or unnecessary resource consumption when connectivity changes during cutover. This coordination also gives teams time to prepare operationally, for example, by suppressing alerts, notifying downstream consumers, or communicating planned maintenance to their customers.</p><p>However, relying solely on application-side controls is unreliable. Operational gaps, misconfigurations, or lingering connections can still modify the source database state, potentially resulting in changes that are not replicated to the destination and leading to data inconsistency or loss. To enforce a clean and deterministic cutover, we also block traffic at the infrastructure layer. This is done by detaching the RDS instance’s security groups to prevent new inbound connections, followed by a reboot of the instance. With security groups removed, no new SQL sessions can be established, and the reboot forcibly terminates any existing connections.</p><p>This approach intentionally avoids requiring database credentials or logging into the PostgreSQL server to manually terminate connections. While it may be slower than application- or database-level intervention, it provides a reliably automated and repeatable mechanism to fully quiesce the source RDS PostgreSQL instance before Aurora promotion, eliminating the risk of divergent writes or an inconsistent WAL state.</p><h4>Validation Phase</h4><p>To determine whether the Aurora read replica has fully caught up with the source RDS PostgreSQL instance, we track replication progress using Aurora’s OldestReplicationSlotLag metric. This metric represents how far the Aurora replica is behind the source in applying write-ahead log (WAL) records.</p><p>Once client traffic is halted during quiescence, the source RDS PostgreSQL instance stops producing meaningful WAL entries. At that point, the replication lag should converge to zero, indicating that all WAL records corresponding to real writes have been fully replayed on Aurora.</p><p>However, in practice, our experiments show that the metric never settles at a steady zero. Instead, it briefly drops to <strong>0</strong>, then quickly returns to <strong>64 MB</strong>, repeating this pattern every few minutes as shown in the figure below.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/674/1*0pgUo0X6RiwoeATejg6NfQ.png"><figcaption>OldestReplicationSlotLag</figcaption></figure><p>This behavior stems from how OldestReplicationSlotLag is calculated. Internally, the lag is derived using the following query:</p><pre>SELECT<br>  slot_name,<br>  pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn) AS slot_lag_bytes<br>FROM pg_replication_slots;</pre><p>Conceptually, this translates to:</p><pre>OldestReplicationSlotLag = current_WAL_position_on_RDS <br>                           – restart_lsn </pre><p>See AWS references <a href="https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/USER_PostgreSQL.Replication.ReadReplicas.Monitor.html">here</a> and <a href="https://repost.aws/knowledge-center/rds-postgresql-use-logical-replication">here</a>.</p><p>The <a href="https://www.morling.dev/blog/postgres-replication-slots-confirmed-flush-lsn-vs-restart-lsn/#:~:text=And%20this%20is%20exactly%20the,has%20a%20few%20important%20implications:"><em>restart_lsn</em></a> represents the oldest write-ahead log (WAL) record that PostgreSQL must retain to ensure a replication consumer can safely resume replication.</p><p>When PostgreSQL performs a WAL segment switch, Aurora typically catches up almost immediately. At that moment, the restart_lsn briefly matches the source’s current WAL position, causing the reported lag to drop to 0. During idle periods, PostgreSQL performs an empty WAL segment rotation approximately every five minutes, driven by the archive_timeout = 300s setting in the database parameter group.</p><p>Immediately afterward, PostgreSQL begins writing to the new WAL segment. Since this new segment has not yet been fully flushed or consumed by Aurora, the WAL position in source RDS PostgreSQL advances ahead of the restart_lsn of Aurora PostgreSQL by exactly one segment. As a result, OldestReplicationSlotLag jumps to 64 MB, which corresponds to the configured WAL segment size at database initialization, and remains there until the next segment switch occurs.</p><p>Because idle PostgreSQL performs an empty WAL rotation approximately every five minutes, this zero-then-64 MB oscillation is expected. Importantly, the moment when the lag drops to 0 indicates that all meaningful WAL records have been fully replicated, and the Aurora read replica is fully caught up with the source.</p><h4>Cutover Phase</h4><p>Once the Aurora read replica has fully caught up with the source RDS PostgreSQL instance — as confirmed through replication lag analysis — the final step is to promote the replica and redirect application traffic. Promoting the Aurora read replica converts it into an independent, writable Aurora PostgreSQL cluster with its own writer and reader endpoints. At this point, the source RDS PostgreSQL instance is no longer the authoritative primary and is made inaccessible.</p><p>Because Netflix’s RDS ecosystem is fronted by a Data Access Layer (DAL), consisting of client-side forward proxies and a centralized Data Gateway, switching databases does not require application code changes or database credential access. Instead, traffic redirection is handled entirely through configuration updates in the reverse-proxy layer. Specifically, we update the runtime configuration of the Envoy-based Data Gateway to route traffic to the newly promoted Aurora cluster. Once this configuration change propagates, all client-initiated database connections are transparently routed through the DAL to the Aurora writer endpoint, completing the migration without requiring any application changes.</p><p>This proxy-level cutover, combined with Aurora promotion, enables a seamless transition for service owners, minimizes downtime, and preserves data consistency throughout the migration process.</p><h3>Customer Experience: Migrating a Business-Critical Partner Platform</h3><p>One of the critical teams to adopt the RDS PostgreSQL to Aurora PostgreSQL migration workflow was the Enablement Applications team. This team owns a set of databases that model Netflix’s entire ecosystem of partner integrations, including device manufacturers, discovery platforms, and distribution partners. These databases power a suite of enterprise applications that partners worldwide rely on to build, test, certify, and launch Netflix experiences on their devices and services.</p><p>Because these databases sit at the center of Netflix’s partner enablement and certification workflows, they are consumed by a diverse set of client applications across both internal and external organizations. <strong>Internally</strong>, reliability teams use this data to identify streaming failures for specific devices and configurations, supporting quality improvements across the device ecosystem. At the same time, these databases directly serve <strong>external</strong> partners operating across many regions. Device manufacturers rely on them to configure, test, and certify new hardware, while payment partners use them to set up and launch bundled offerings with Netflix.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*DJHayui83MkdSMYYd4TIQw.png"><figcaption>Simplified Enablement Applications Overview</figcaption></figure><p><strong>Device Lifecycle Management</strong></p><p>Netflix works with a wide range of device partners to ensure Netflix streams seamlessly across a diverse ecosystem of consumer devices. A core responsibility of Device Lifecycle Management is to provide tools and workflows that allow partners to develop, test, and certify Netflix integrations on their devices.</p><p>As part of the device lifecycle, partners run Netflix-provided test suites against their NRDP implementation. We store <strong>signals that represent the current stage for each device in the certification process</strong>. This certification data forms the backbone of Netflix’s device enablement program, ensuring that only validated devices can launch Netflix experiences.</p><p><strong>Partner Billed Integrations</strong></p><p>In addition to device enablement, the same partner metadata is also consumed by Netflix’s Partner Billed Integrations organization. This group enables external partners to offer Netflix as part of bundled subscription and billing experiences.</p><p>Any disruption in these databases affects partner integration workflows. If the database is unavailable, partners may be unable to configure or launch service bundles with Netflix. Maintaining high availability and data correctness is essential to preserving smooth integration operations.</p><p>The global nature of these workflows makes it difficult to schedule downtime windows. Any disruption would impact partner productivity and risk eroding trust in Netflix’s integration and certification processes.</p><h3>Preparation</h3><p>Given the criticality of the Enablement Applications databases, thorough preparation was essential before initiating the migration. The team invested significant effort upfront to understand traffic patterns, identify all consumers, and establish clear communication channels.</p><p><strong>Understand Client Fan-Out and Traffic Patterns<br></strong>The first step was to gain a complete view of how the databases were being used in production. Using observability tools like CloudWatch metrics, the team analyzed PostgreSQL connection counts, read and write patterns, and overall load characteristics. This helped establish a baseline for normal behavior and ensured there were no unexpected traffic spikes or hidden dependencies that could complicate the migration.</p><p>Just as importantly, this baseline gave the Enablement Applications team a rough idea of the post-migration behavior on Aurora. For example, they expected to see a similar number of active database connections and comparable traffic patterns after cutover, making it easier to validate that the migration had preserved operational characteristics.</p><p><strong>Identify and Enumerate All Database Consumers<br></strong>Unlike most databases, where the set of consumers is well known to the owning team, these databases were accessed by a wide range of internal services and external-facing systems that were not fully enumerated upfront. To address this, we leveraged a tool called flowlogs, an eBPF-based network attribution tooling was used to capture TCP flow data to identify the services and applications establishing connections to the database(<a href="https://netflixtechblog.com/how-netflix-accurately-attributes-ebpf-flow-logs-afe6d644a3bc">link</a>).<br> <br>This approach allowed the team to enumerate active consumers, including those that were not previously documented, ensuring no clients were missed during migration planning.</p><p><strong>Establish Dedicated Communication Channels<br></strong>Once all consumers were identified, a dedicated communication channel was created to provide continuous updates throughout the migration process. This channel was used to share timelines, readiness checks, status updates, and cutover notifications, ensuring that all stakeholders remained aligned and could respond quickly if issues arose.</p><h3>Migration Process</h3><p>After completing application-side preparation, the Enablement Applications team initiated the data replication phase of the migration workflow. The automation successfully provisioned the Aurora read replica cluster and ported the RDS PostgreSQL parameter group to a corresponding Aurora parameter group, bringing the destination environment up with equivalent configuration.</p><h4><strong>Unexpected Replication Slot Behavior</strong></h4><p>However, shortly after replication began, we observed that the OldestReplicationSlotLag metric was unexpectedly high. This was counterintuitive, as Aurora read replicas are designed to remain closely synchronized with the source database by continuously streaming write-ahead logs (WAL).</p><p>Further investigation revealed the presence of an inactive logical replication slot on the source RDS PostgreSQL instance. An inactive replication slot can cause elevated OldestReplicationSlotLag because PostgreSQL must retain all WAL records required by the slot’s last known position (restart_lsn), even if no client is actively consuming data from it. Replication slots are intentionally designed to prevent data loss by ensuring that a consumer can resume replication from where it left off. As a result, PostgreSQL will not recycle or delete WAL segments needed by a replication slot until the slot advances. When a slot becomes inactive — such as when a client migration task is stopped or abandoned — the slot’s position no longer moves forward. Meanwhile, the database continues to generate WAL, forcing PostgreSQL to retain increasingly older WAL files. This growing gap between the current WAL position and the slot’s restart_lsn manifests as a high OldestReplicationSlotLag.</p><p>Identifying and addressing these inactive replication slots was a critical prerequisite to proceeding safely with the migration and ensuring accurate replication state during cutover.</p><p><strong>Successful Migration After Remediation<br> </strong>After identifying the inactive logical replication slot, the team safely cleaned it up on the source RDS PostgreSQL instance and resumed the migration workflow. With the stale slot removed, replication progressed as expected, and the Aurora read replica quickly converged with the source. The migration then proceeded smoothly through the quiescence phase, with no unexpected behavior or replication anomalies observed.</p><p>Following promotion, application traffic transitioned seamlessly to the newly writable Aurora PostgreSQL cluster. Through the Data Access Layer, new client connections were automatically routed to Aurora, and observability metrics confirmed healthy behavior — connection counts, read/write patterns, and overall load closely matched pre-migration baselines. From the application and partner perspective, the cutover was transparent, validating both the correctness of the migration workflow and the effectiveness of the preparation steps.</p><h3>Open questions</h3><h4>How do we select target Aurora PostgreSQL instance types based on the existing production RDS PostgreSQL instance?</h4><p>When selecting the target Aurora PostgreSQL instance type for a production migration, our guidance is intentionally conservative. We prioritize stability and performance first, and optimize for cost only after observing real workload behavior on Aurora.</p><p>In practice, the recommended approach is to adopt Graviton2-based instances (particularly the <em>r6g</em> family) whenever possible, maintain the same instance family and size where feasible, and — at minimum — preserve the memory footprint of the existing RDS instance.</p><p>Unlike RDS PostgreSQL, Aurora does not support the <em>m</em>-series, making a direct family match impossible for those instances. In such cases, simply keeping the same “size” (e.g., 2xlarge → 2xlarge) is not meaningful because the memory profiles differ across families. Instead, we map instances by memory equivalence. For example, an Aurora <em>r6g.xlarge</em> provides a memory footprint comparable to an RDS <em>m5.2xlarge</em>, making it a practical replacement. This memory-aligned strategy offers a safer and more predictable baseline for production migrations.</p><h4><strong>Downtime During RDS → Aurora Cutover?</strong></h4><p>To achieve minimal downtime during an RDS PostgreSQL → Aurora PostgreSQL migration, we front-load as much work as possible into the preparation phase. By the time we reach cutover, the Aurora read replica is already provisioned and continuously replicating WAL from the source RDS instance. Before initiating downtime, we ensure that the replication lag between Aurora and RDS has stabilized within an acceptable threshold. If the lag is large or fluctuating significantly, forcing a cutover will only inflate downtime.</p><p>Downtime begins the moment we remove the security groups from the source RDS instance, blocking all inbound traffic. We then reboot the instance to forcibly terminate existing connections, which typically takes up to a minute. From this point forward, no writes can be performed.</p><p>After traffic is halted, the next objective is to verify that Aurora has fully replayed all meaningful WAL records from RDS. We track this using <strong>OldestReplicationSlotLag</strong>. We first wait for the metric to drop to <strong>0</strong>, indicating that Aurora has consumed all WAL with real writes. Under normal idle behavior, PostgreSQL triggers an empty WAL switch every five minutes. After observing one data point at 0, we wait for an additional idle WAL rotation and confirm that the lag oscillates within the expected <strong>0 → 64 MB</strong> pattern — signifying that the only remaining WAL segments are empty ones produced during idle time. At this point, we know the Aurora replica is fully caught up and can be safely promoted.</p><p>While these validation steps run, we perform the configuration updates on the Envoy reverse proxy in parallel. Once promotion completes and Envoy is restarted with the new runtime configuration, all client-initiated connections begin routing to the Aurora cluster. In practice, the total write-downtime observed across services averages <strong>around 10 minutes</strong>, dominated largely by the RDS reboot and the idle WAL switch interval.</p><p><strong>Optimization: Reducing Idle-Time Wait</strong></p><p>For services requiring stricter downtime budgets, waiting the full five minutes for an idle WAL switch can be prohibitively expensive. In such cases, we can force a WAL rotation immediately after traffic is cut off by issuing:</p><p>SELECT pg_switch_wal();</p><p>Once the switch occurs, OldestReplicationSlotLag will drop to 0 again as Aurora consumes the new (empty) WAL segment. This approach eliminates the need to wait for the default archive_timeout interval, which can significantly reduce overall downtime.</p><h4>How do we migrate CDC consumers?</h4><p>As part of the data platform organization in Netflix, we provide a managed Change Data Capture (CDC) service across a variety of datastores. For PostgreSQL, logical replication slots is the way of implementing change data capture. At Netflix, we build a managed abstraction on top of these replication slots called <strong>datamesh</strong> to manage customers who are leveraging them (<a href="https://netflixtechblog.com/data-mesh-a-data-movement-and-processing-platform-netflix-1288bcab2873">link</a>).</p><p>Each logical replication slot tracks a consumer’s position in the write-ahead log (WAL), ensuring that WAL records are retained until the consumer has successfully processed them. This guarantees ordered and reliable delivery of row-level changes to downstream systems. At the same time, it tightly couples the lifecycle of replication slots to database operations, making their management a critical consideration during database migrations.</p><p>A key challenge in migrating from RDS PostgreSQL to Aurora PostgreSQL is transitioning these CDC consumers safely — without data loss, stalled replication, or extended downtime — while ensuring that replication slots are correctly managed throughout the cutover process.</p><p>Each row-level change in PostgreSQL is emitted as a CDC event with an operation type of INSERT, UPDATE, DELETE, or REFRESH. REFRESH events are generated during backfills by querying the database directly and emitting the current state of rows in chunks. Downstream consumers are designed to be idempotent and eventually consistent, allowing them to safely process retries, replays, and backfills.</p><p><strong>Handling Replication Slots During Migration</strong></p><p>Before initiating database cutover, we temporarily pause CDC consumption by stopping the infrastructure responsible for consuming from PostgreSQL replication slots and writing into datamesh source. This also drops the replication slot from the database and cleans up our internal state around replication slot offsets. This essentially resets the state of the connector to one of a brand new one.</p><p>This step is critical for two reasons. First, it prevents replication slots from blocking WAL recycling during migration. Second, it ensures that no CDC consumers are left pointing at the source database once traffic is quiesced and cutover begins. While CDC consumers are paused, downstream systems temporarily stop receiving new change events, but remain stable. Once CDC consumers are paused, we proceed with stopping other client traffic and executing the RDS-to-Aurora cutover.</p><p><strong>Reinitializing CDC After Cutover</strong></p><p>After the Aurora PostgreSQL cluster has been promoted and traffic has been redirected, CDC consumers are reconfigured to point to the Aurora endpoint and restarted. Because their previous state was intentionally cleared, consumers initialize as if they are starting fresh.</p><p>On startup, new logical replication slots are created on Aurora, and a full backfill is performed by querying the database and emitting REFRESH events for all existing rows. These events let the consumer know that a manual refresh was done from Aurora and to treat this as an upsert operation. This establishes a clean and consistent baseline from which ongoing CDC can resume. Consumers are expected to handle these refresh events correctly as part of normal operation.</p><p>By explicitly managing PostgreSQL replication slots as part of the migration workflow, we are able to migrate CDC consumers safely and predictably, without leaving behind stalled slots, retained WAL, or consumers pointing to the wrong database. This approach allows CDC pipelines to be cleanly re-established on Aurora while preserving correctness and operational simplicity.</p><h4>How do we roll back in the middle of the process?</h4><p><strong>Pre-quiescence<br></strong>Rolling back before the pre-quienscence phase is quite easy. Your primary RDS database is still the source. Rolling back before the quiescence phase is straightforward. At this stage, the primary RDS PostgreSQL instance continues to serve as the sole source of truth, and no client traffic has been redirected.</p><p>If a rollback is required, the migration can be safely aborted by deleting the newly created Aurora PostgreSQL cluster along with its associated parameter groups. No changes are needed on the application side, and normal operations on RDS PostgreSQL can continue without impact.</p><p><strong>During-quiescence<br></strong>Rolling back during the quiescence phase is more involved. At this point, client traffic to the source RDS PostgreSQL instance has already been stopped by detaching its security groups. To roll back safely, access must first be restored by reattaching the original security groups to the RDS instance, allowing client connections to resume. In addition, any logical replication slots removed during the migration must be recreated so that CDC consumers can continue processing changes from the source database.</p><p>Once connectivity and replication slots are restored, the RDS PostgreSQL instance can safely resume its role as the primary source of truth.</p><p><strong>Post-quiescence <br></strong>Rolling back after cutover, once the Aurora PostgreSQL cluster is serving production traffic, is significantly more complex. At this stage, Aurora has become the primary source of truth, and client applications may already have written new data to it.</p><p>In this scenario, rollback requires setting up replication in the opposite direction, with Aurora as the source and RDS PostgreSQL as the destination. This can be achieved using a service such as AWS Database Migration Service (DMS). AWS provides detailed guidance for setting up this reverse replication flow, which can be followed to migrate data back to RDS if necessary.</p><h3>Conclusion</h3><p>Standardizing and reducing the surface area of data technologies is crucial for any large-scale platform. For the Netflix platform team, this strategy allows us to concentrate engineering effort, deliver deeper value on a smaller set of well-understood systems, and significantly cut the operational overhead of running multiple database technologies that serve similar purposes. Within the relational database ecosystem, Aurora PostgreSQL has become the paved-path datastore — offering strong scalability, resilience, and consistent operational patterns across the fleet.</p><p>Migrations of this scale demand solutions that are reliable, low-touch, and minimally disruptive for service owners. Our automated RDS PostgreSQL → Aurora PostgreSQL workflow represents a major step forward, providing predictable cutovers, strong correctness guarantees, and a migration experience that works uniformly across diverse workloads.</p><p>As we continue this journey, the Relational Data Platform team is building higher-level abstractions and capabilities on top of Aurora, enabling service owners to focus less on the complexities of database internals and more on delivering product value. More to come — stay tuned.</p><h3>Acknowledgements</h3><p>Special thanks to our other stunning colleagues/customers who contributed to the success of the RDS PostgreSQL to Aurora PostgreSQL migration. <a href="mailto:spasupuleti@netflix.com">Sumanth Pasupuleti</a>, <a href="mailto:coleantoniop@netflix.com">Cole Perez</a>, <a href="mailto:akhaku@netflix.com">Ammar Khaku</a></p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=261ca045447f" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/automating-rds-postgres-to-aurora-postgres-migration-261ca045447f">Automating RDS Postgres to Aurora Postgres Migration</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/automating-rds-postgres-to-aurora-postgres-migration-261ca045447f</link>
      <guid>https://netflixtechblog.com/automating-rds-postgres-to-aurora-postgres-migration-261ca045447f</guid>
      <pubDate>Thu, 12 Feb 2026 15:07:00 +0100</pubDate>
    </item>
    <item>
      <title><![CDATA[The AI Evolution of Graph Search at Netflix]]></title>
      <description><![CDATA[<h3>The AI Evolution of Graph Search at Netflix: From Structured Queries to Natural Language</h3><p>By <a href="https://www.linkedin.com/in/ahutter/">Alex Hutter</a> and <a href="https://www.linkedin.com/in/bartosz-balukiewicz/">Bartosz Balukiewicz</a></p><p>Our previous blog posts (<a href="https://netflixtechblog.com/how-netflix-content-engineering-makes-a-federated-graph-searchable-5c0c1c7d7eaf">part 1</a>, <a href="https://netflixtechblog.com/how-netflix-content-engineering-makes-a-federated-graph-searchable-part-2-49348511c06c">part 2</a>, <a href="https://netflixtechblog.com/reverse-searching-netflixs-federated-graph-222ac5d23576">part 3</a>) detailed how Netflix’s Graph Search platform addresses the challenges of searching across federated data sets within Netflix’s enterprise ecosystem. Although highly scalable and easy to configure, it still relies on a structured query language for input. Natural language based search has been possible for some time, but the level of effort required was high. The emergence of readily-available AI, specifically Large Language Models (LLMs), has created new opportunities to integrate AI search features, with a smaller investment and improved accuracy.</p><p>While Text-to-Query and Text-to-SQL are established problems, the complexity of distributed Graph Search data in the GraphQL ecosystem necessitates innovative solutions. This is the first in a three-part series where we will detail our journey: how we implemented these solutions, evaluated their performance, and ultimately evolved them into a self-managed platform.</p><h3>The Need for Intuitive Search: Addressing Business and Product Demands</h3><p>Natural language search is the ability to use everyday language to retrieve information as opposed to complex, structured query languages like the Graph Search Filter Domain Specific Language (DSL). When users interact with 100’s of various UIs within the suite of Content and Business Products applications, a frequent task is filtering a data table like the one below:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*qujlm9wJ_fZTdOrG-EiBcA.png"><figcaption>Example Content and Business Products application view</figcaption></figure><p>Ideally, a user simply wants to satisfy a query like <strong>“I want to see all movies from the 90s about robots from the US.”</strong> Because the underlying platform operates on the Graph Search Filter DSL, the application acts as an intermediary. Users input their requirements through UI elements — toggling facets or using query builders — and the system programmatically converts these interactions into a valid DSL query to filter the data.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/612/1*vCXVNhRXGbjLjseNY2vmlQ.png"><figcaption>The Complexity of filtering and DSL generation</figcaption></figure><p>This process presents a few issues.</p><p>Today, many applications have bespoke components for collecting user input — the experience varies across them and they have inconsistent support for the DSL. Users need to “learn” how to use each application to achieve their goals.</p><p>Additionally, some domains have hundreds of fields in an index that could be faceted or filtered by. A <em>subject matter expert </em>(SME) may know exactly what they want to accomplish, but be bottlenecked by the inefficient pace of filling out a large scale UI form and translating their questions in order to encode it in a representation Graph Search needs.</p><p>Most importantly, users think and operate using natural language, not technical constructs like query builders, components, or DSLs. By requiring them to switch contexts, we introduce friction that slows them down or even prevents their progress.</p><p>With readily-available AI components, our users can now interact with our systems through natural language. The challenge now is to make sure our offering, searching Netflix’s complex enterprise state with natural language, is an intuitive and trustworthy experience.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/615/1*EklLSPYRp9yrPP3wLW5qfg.png"><figcaption>Natural language queries translated into Graph Search Filter DSL</figcaption></figure><p>We’ve made a decision to pursue generating Graph Search Filter statements from natural language to meet this need. Our intention is to augment and not replace existing applications with <a href="https://en.wikipedia.org/wiki/Retrieval-augmented_generation">retrieval augmented generation</a> (RAG), providing tooling and capabilities so that applications in our ecosystem have newly accessible means of processing and presenting their data in their distinct domain flavours. It should be noted that all the work here has direct application to building a RAG system on top of Graph Search in the future.</p><h3>Under the Hood: Our Approach to Text-to-Query</h3><p>The core function of the text-to-query process is converting a user’s (often ambiguous) natural language question into a structured query. We primarily achieve this through the use of an LLM.</p><p>Before we dive deeper, let’s quickly revisit the structure of Graph Search Filter DSL. Each Graph Search index is <a href="https://netflixtechblog.com/how-netflix-content-engineering-makes-a-federated-graph-searchable-5c0c1c7d7eaf#:~:text=of%20configuration%20required.-,Configuration,-For%20collecting%20the">defined by a GraphQL query</a>, made up of a collection of fields. Each field has a type e.g. boolean, string, and some have their permitted values governed by controlled vocabularies — a standardized and governed list of values (like an enumeration, or a foreign key). The names of those fields can be used to construct expressions using comparison (e.g. &gt; or ==) or inclusion/exclusion operators (e.g. IN). In turn those expressions can be combined using logical operators (e.g. AND) to construct complex statements.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/615/1*UbP1eIXDDqt5Q8nlCdqIKQ.png"><figcaption>Graph Search Filter DSL</figcaption></figure><p>With that understanding, we can now more rigorously define the conversion process. We need the LLM to generate a Graph Search Filter DSL statement that is syntactically, semantically, and pragmatically correct.</p><p><strong>Syntactic correctness</strong> is easy — does it parse? To be syntactically correct, the generated statement must be well formed<strong> </strong>i.e. follow the grammar of the Graph Search Filter DSL.</p><p><strong>Semantic correctness </strong>adds some additional complexity as it requires more knowledge of the index itself. To be semantically correct:</p><ul><li>it must respect the field types i.e. only use comparisons that make sense given the underlying type;</li><li>it must only use fields that are actually present in the index, i.e. does not <em>hallucinate;</em></li><li>when the values of a field are constrained to a controlled vocabulary, any comparison must only use values from that controlled vocabulary.</li></ul><p><strong>Pragmatic correctness</strong> is much more difficult. It asks the question: does the generated filter actually capture the intent of the user’s query?</p><p>The following sections will detail how we pre-process the user’s question to create appropriate context for the instructions that we will provide to the LLM — both of <a href="https://developers.google.com/machine-learning/resources/intro-llms">which are fundamental to LLM interaction</a> — as well as post-processing we perform on the generated statement to validate it, and help users understand and trust the results they receive.</p><p>At a high level that process looks like this:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/284/1*s34kKg5FI3TDd6UeXaQWOA.png"><figcaption>Graph Search FIlter DSL generation process</figcaption></figure><h3>Context Engineering</h3><p>Preparation for the filter generation task is predominantly engineering the appropriate context. The LLM will need access to the fields of an index and their metadata in order to construct semantically correct filters. As the indices are defined by GraphQL queries, we can use the type information from the GraphQL schema to derive much of the required information. For some fields, there is additional information we can provide beyond what’s available in the schema as well, in particular permissible values that pull from controlled vocabularies.</p><p>Each field in the index is associated with metadata as seen below, and that metadata is provided as part of the context.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/655/1*Z34Ek0ib34dd600ZH9iT0Q.png"><figcaption>Graph Search index representation</figcaption></figure><ul><li>The <strong>field</strong> is derived from the document path as characterized by the GraphQL query.</li><li>The <strong>description</strong> is the comment from the GraphQL schema for the field.</li><li>The <strong>type</strong> is derived from the GraphQL schema for the field e.g. Boolean, String, enum. We also support an additional controlled vocabulary type we will discuss more of shortly.</li><li>The <strong>valid values</strong> are derived from enum values for the enum type or from a controlled vocabulary as we will now discuss.</li></ul><p>A <em>controlled vocabulary</em> is a specific field type that consists of a finite set of allowed values, which are defined by a SMEs or domain owners. Index fields can be associated with a particular controlled vocabulary, e.g. countries with members such as Spain and Thailand, and any usage of that field within a generated statement must refer to values from that vocabulary.</p><p>Naively providing all the metadata as context to the LLM worked for simple cases but did not scale. Some indices have hundreds of fields and some controlled vocabularies have thousands of valid values. Providing all of those, especially the controlled vocabulary values and their accompanying metadata, expands the context; this proportionally increases latency and decreases the correctness of generated filter statements. Not providing the values wasn’t an option as we needed to ground the LLMs generated statements- without them, the LLM would frequently hallucinate values that did not exist.</p><p>Curating the context to an appropriate subset was a problem we addressed using the well known RAG pattern.</p><h4>Field RAG</h4><p>As mentioned previously, some indices have hundreds of fields, however, most user’s questions typically refer only to a handful of them. If there was no cost in including them all, we would, but as mentioned prior, there is a cost in terms of the latency of query generation as well as the correctness of the generated query (e.g. needle-in-the-hackstack problem) and non-deterministic results.</p><p>To determine which subset of fields to include in the context, we “match” them against the intent of the user’s question.</p><ul><li>Embeddings are created for index fields and their metadata (name, description, type) and are indexed in a vector store</li><li>At filter generation time, the user’s question is chunked with an overlapping strategy. For each chunk, we perform a vector search to identify the top K most relevant values and the fields to which they belong.</li><li><strong>Deduplication:</strong> The top K fields from each chunk are both consolidated and deduplicated before being provided as context to the system instructions.</li></ul><figure><img alt="" src="https://cdn-images-1.medium.com/max/599/1*_fkaUvg0ljRYJnws1xu_Rw.png"><figcaption>Field RAG process (chunking, merge, deduplicate)</figcaption></figure><h4>Controlled Vocabularies RAG</h4><p>Index fields of the controlled vocabulary type are associated with a particular controlled vocabulary, again, countries are one example. Given a user’s question, we can infer whether or not it refers to values of a particular controlled vocabulary. In turn, by knowing which controlled vocabulary values are present, we can identify additional, related index fields that should be included in the context that may not have been identified by the field RAG step.</p><p>Each controlled vocabulary value has:</p><ul><li>a unique<strong> identifier</strong> within its type;</li><li>a human readable <strong>display name;</strong></li><li>a <strong>description</strong> of the value;</li><li>also-known-as values or <strong>AKA</strong> display names, e.g. “romcom” for “Romantic Comedy”.</li></ul><p>To determine which subset of values to include in the context for controlled vocabulary fields (and also possibly infer additional fields), we “match” them against the user’s question.</p><ul><li>Embeddings are created for controlled vocabulary values and their metadata, and these are indexed in a vector store. The controlled vocabularies are available via GraphQL and are regularly fetched and reindexed so this system stays up to date with any changes in the domain.</li><li>At filter generation time, the user’s question is chunked. For each chunk, we perform a vector search to identify the top K most relevant values (but only for the controlled vocabularies that are associated with fields in the index)</li><li>The top K values from each chunk are deduplicated by their controlled vocabulary type. The associated field definition is then injected into the context along with the matched values.</li></ul><figure><img alt="" src="https://cdn-images-1.medium.com/max/461/1*N70dw5GEqNDVDYPmpRmqUQ.png"><figcaption>Controlled Vocabularies RAG</figcaption></figure><p><strong>Combining both approaches, the RAG of fields and controlled vocabularies, we end up with the solution that each input question resolves in available and matched fields and values:</strong></p><figure><img alt="" src="https://cdn-images-1.medium.com/max/706/1*hIALZVLkkOjZ5ePgsdVWIA.png"><figcaption>Field and CV RAG</figcaption></figure><p>The quality of results generated by the RAG tool can be significantly enhanced by tuning its various parameters, or “levers.” These include strategies for reranking, chunking, and the selection of different embedding generation models. The careful and systematic evaluation of these factors will be the focus of the subsequent parts of this series.</p><h3>The Instructions</h3><p>Once the context is constructed, it is provided to the LLM with a set of instructions and the user’s question. The instructions can be summarised as follows: <strong>“<em>Given a natural language question, generate a syntactically, semantically, and pragmatically correct filter statement given the availability of the following index fields and their metadata</em>.”</strong></p><ul><li>In order to generate a <em>syntactically</em> correct filter statement, the instructions include the syntax rules of the DSL.</li><li>In order to generate a <em>semantically</em> correct filter statement, the instructions tell the LLM to ground the generated statement in the provided context.</li><li>In order to generate a <em>pragmatically</em> correct filter statement, so far we focus on better context engineering to ensure that only the most relevant fields and values are provided. We haven’t identified any instructions that make the LLM just “do better” at this aspect of the task.</li></ul><figure><img alt="" src="https://cdn-images-1.medium.com/max/893/1*gCfHD-yVJhpcj6zPYRhzqQ.png"><figcaption>Graph Search Filter DSL generation</figcaption></figure><p>After the filter statement is generated by the LLM, we deterministically validate it prior to returning the values to the user.</p><h3>Validation</h3><h4>Syntactic Correctness</h4><p>Syntactic correctness ensures the LLM output is a parsable filter statement. We utilize an Abstract Syntax Tree (AST) parser built for our custom DSL. If the generated string fails to parse into a valid AST, we know immediately that the query is malformed and there is a fundamental issue with the generation.</p><p>The other approach to solve this problem could be using the <a href="https://platform.openai.com/docs/guides/structured-outputs">structured outputs</a> modes provided by some LLMs. However, our initial evaluation yielded mixed results, as the custom DSL is not natively supported and requires further work.</p><h4>Semantic Correctness</h4><p>Despite careful context engineering using the RAG pattern, the LLM sometimes hallucinates both fields and available values in the generated filter statement. The most straightforward way of preventing this phenomenon is validating the generated filters against available index metadata. This approach does not impact the overall latency of the system, as we are already working with an AST of the filter statement, and the metadata is freely available from the context engineering stage.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/553/1*0LUUAK1G7CtwFDBZsAjgDA.png"><figcaption>DSL verification &amp; hallucinations</figcaption></figure><p>If a hallucination is detected it can be returned as an error to a user, indicating the need to refine the query, or can be provided back to the LLM in the form of a feedback loop for self correction.</p><p>This increases the filter generation time, so should be used cautiously with a limited number of retries.</p><h3>Building Confidence</h3><p>You probably noticed we are not validating the generated filter for pragmatic correctness. That task is the hardest challenge: The filter parses (<em>syntactic</em>) and uses real fields (<em>semantic</em>), but is it what the user meant? When a user searches for <strong>“Dark”</strong>, do they mean <strong>the specific German sci-fi series <em>Dark</em>, </strong>or are they browsing for the mood category<strong> “dark TV shows”</strong>?</p><p>The gap between what a user intended and the generated filter statement is often caused by ambiguity. Ambiguity stems from the <a href="https://en.wikipedia.org/wiki/Semantic_compression">compression of natural language</a>. A user says <strong>“German time-travel mystery with the missing boy and the cave”</strong> but the index contains <strong>discrete metadata fields</strong> like <strong>releaseYear</strong>, <strong>genreTags</strong>, and <strong>synopsisKeywords</strong>.</p><p>How do we ensure users aren’t inadvertently led to wrong answers or to answers for questions they didn’t ask?</p><h4>Showing Our Work</h4><p>One way we are handling ambiguity is by <em>showing our work</em>. We visualise the generated filters in the UI in a user-friendly way allowing them to very clearly see if the answer we’re returning is what they were looking for so they can trust the results..</p><p>We cannot show a raw DSL string (e.g., <em>origin.country == ‘Germany’ AND genre.tags CONTAINS ‘Time Travel’ AND synopsisKeywords LIKE ‘*cave*’</em>) to a non-technical user. Instead, we reflect its underlying AST into UI components.</p><p>After the LLM generates a filter statement, we parse it into an AST, and then map that AST to the existing “Chips” and “Facets” in our UI (see below). If the LLM generates a filter for <em>origin.country == ‘Germany’</em>, the user sees the “Country” dropdown pre-selected to “Germany.” This gives users immediate visual feedback and the ability to easily fine-tune the query using standard UI controls when the results need improvement or further experimentation.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*_Gx0THjlWW9_jwb8KwCrYw.png"><figcaption>Generated filters visualisation</figcaption></figure><h4>Explicit Entity Selection</h4><p>Another strategy we’ve developed to remove ambiguity happens at query time. We give users the ability to constrain their input to refer to known entities using “@mentions”. Similar to Slack, typing @ lets them search for entities directly from our specialized UI Graph Search component, giving them easy access to multiple controlled vocabularies (plus other identifying metadata like launch year) to feel confident they’re choosing the entity they intend.</p><p>If a user types, “When was <em>@dark</em> produced”, we explicitly know they are referring to the <em>Series</em> controlled vocabulary, allowing us to bypass the RAG inference step and hard-code that context, significantly increasing pragmatic correctness (and building user trust in the process).</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*u3YrDuaYJr4LlVB1286Xzg.png"><figcaption>Example @mentions usage in the UI</figcaption></figure><h3>End-to-end architecture</h3><p>As mentioned previously, the solution architecture is divided into <em>pre-processing</em>, filter statement generation, and then <em>post-processing</em> stages. The pre-processing handles context building and involves a RAG pattern for similarity search, while the post-processing validation stage checks the correctness of the LLM-generated filter statements and provides visibility into the results for end users. This design strategically balances LLM involvement with more deterministic strategies.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*lw47a7h67i4EUmvhlRKcYQ.png"><figcaption>End-to-end architecture</figcaption></figure><p>The end-to-end process is as follows:</p><ol><li>A user’s natural language question (with optional `@mentions` statements) are provided as input, along with the Graph Search index context</li><li>The context is scoped by using the RAG pattern on both fields and possible values</li><li>The pre-processed context and the question are fed into the LLM with an instruction asking for<em> a syntactically and semantically correct filter statement</em></li><li>The generated filer statement DSL is verified and checked for hallucinations</li><li>The final response contains the related AST in order to build “Chips” and “Facets”</li></ol><h3>Summary</h3><p>By combining our existing Graph Search infrastructure with the power and flexibility of LLMs, we’ve bridged the gap between complex filter statements and user intent. We moved from requiring users to speak our language (DSL) to our systems understanding theirs.</p><p>The initial challenge for our users was successfully addressed. However, our next steps involve transforming this system into a comprehensive and expandable platform, rigorously evaluating its performance in a live production environment, and expanding its capabilities to support GraphQL-first user interfaces. These topics, and others, will be the focus of the subsequent installments in this series. Be sure to follow along!</p><p>You may have noticed that we have a lot more to do on this project, including named entity recognition and extraction, intent detection so we can route questions to the appropriate indices, and query rewriting among others. If this kind of work interests you, reach out! We’re hiring in our Warsaw office, check for open roles <a href="https://explore.jobs.netflix.net/careers?location=Warsaw%2C%20Masovian%20Voivodeship%2C%20Poland&amp;pid=790302168096&amp;domain=netflix.com&amp;sort_by=relevance&amp;triggerGoButton=false">here</a>.</p><h3>Credits</h3><p>Special thanks to <a href="https://www.linkedin.com/in/quesadaalejandro/">Alejandro Quesada</a>, <a href="https://www.linkedin.com/in/yevgeniya-li-9877ba160/">Yevgeniya Li</a>, <a href="https://www.linkedin.com/in/dkyrii/">Dmytro Kyrii</a>, <a href="https://www.linkedin.com/in/razvan-gabriel-gatea/">Razvan-Gabriel Gatea</a>, <a href="https://www.linkedin.com/in/milodorif/">Orif Milod</a>, <a href="https://www.linkedin.com/in/michal-krol-45973411a/">Michal Krol</a>, <a href="https://www.linkedin.com/in/jeffbalis/">Jeff Balis</a>, <a href="https://www.linkedin.com/in/czhao/">Charles Zhao</a>, <a href="https://www.linkedin.com/in/shilpamotukuri/">Shilpa Motukuri</a>, <a href="https://www.linkedin.com/in/shervineamidi/">Shervine Amidi</a>, <a href="https://www.linkedin.com/in/aborysov/">Alex Borysov</a>, <a href="https://www.linkedin.com/in/mike-azar-7064883b/">Mike Azar</a>, <a href="https://www.linkedin.com/in/bernardo-g-4414b41/">Bernardo Gomez Palacio</a>, <a href="https://www.linkedin.com/in/haoyuan-h-98b587134/">Haoyun He</a>, <a href="https://www.linkedin.com/in/edyr96/">Eduardo Ramirez</a>, <a href="https://www.linkedin.com/in/yujiaxie2019/">Cynthia Xie</a>.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=d416ec5b1151" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/the-ai-evolution-of-graph-search-at-netflix-d416ec5b1151">The AI Evolution of Graph Search at Netflix</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/the-ai-evolution-of-graph-search-at-netflix-d416ec5b1151</link>
      <guid>https://netflixtechblog.com/the-ai-evolution-of-graph-search-at-netflix-d416ec5b1151</guid>
      <pubDate>Mon, 26 Jan 2026 20:01:00 +0100</pubDate>
    </item>
    <item>
      <title><![CDATA[How Temporal Powers Reliable Cloud Operations at Netflix]]></title>
      <description><![CDATA[<p>By <a href="https://www.linkedin.com/in/jacobmeyers35/">Jacob Meyers</a> and <a href="https://www.linkedin.com/in/robzienert/">Rob Zienert</a></p><p><a href="https://temporal.io/">Temporal</a> is a <a href="https://docs.temporal.io/evaluate/understanding-temporal#durable-execution">Durable Execution</a> platform which allows you to write code “as if failures don’t exist”. It’s become increasingly critical to Netflix since its initial adoption in 2021, with users ranging from the operators of our <a href="https://about.netflix.com/en/news/how-netflix-works-with-isps-around-the-globe-to-deliver-a-great-viewing-experience">Open Connect</a> global CDN to our <a href="https://medium.com/netflix-techblog/behind-the-streams-live-at-netflix-part-1-d23f917c2f40">Live</a> reliability teams now depending on Temporal to operate their business-critical services. In this post, I’ll give a high-level overview of what Temporal offers users, the problems we were experiencing operating Spinnaker that motivated its initial adoption at Netflix, and how Temporal helped us reduce the number of transient deployment failures at Netflix from <strong>4% to 0.0001%</strong>.</p><h3>A Crash Course on (some of) Spinnaker</h3><p><a href="https://netflixtechblog.com/global-continuous-delivery-with-spinnaker-2a6896c23ba7">Spinnaker</a> is a multi-cloud continuous delivery platform that powers the vast majority of Netflix’s software deployments. It’s composed of several (mostly nautical themed) microservices. Let’s double-click on two in particular to understand the problems we were facing that led us to adopting Temporal.</p><p>In case you’re completely new to Spinnaker, Spinnaker’s fundamental tool for deployments is the <em>Pipeline</em>. A Pipeline is composed of a sequence of steps called <em>Stages</em>, which themselves can be decomposed into one or more <em>Tasks</em>, or other Stages. An example deployment pipeline for a production service may consist of these stages: Find Image -&gt; Run Smoke Tests -&gt; Run Canary -&gt; Deploy to us-east-2 -&gt; Wait -&gt; Deploy to us-east-1.</p><figure><img alt="An example Spinnaker Pipeline" src="https://cdn-images-1.medium.com/max/1024/1*7sGhc8LhyqQlW9Uiq76TWQ.png"><figcaption>An example Spinnaker Pipeline for a Netflix service</figcaption></figure><p>Pipeline configuration is extremely flexible. You can have Stages run completely serially, one after another, or you can have a mix of concurrent and serial Stages. Stages can also be executed conditionally based on the result of previous stages. This brings us to our first Spinnaker service: <em>Orca</em>. Orca is the <a href="https://raw.githubusercontent.com/spinnaker/orca/refs/heads/master/logo.jpg">orca-stration</a> engine of Spinnaker. It’s responsible for managing the execution of the Stages and Tasks that a Pipeline unrolls into and coordinating with other Spinnaker services to actually execute them.</p><p>One of those collaborating services is called <em>Clouddriver</em>. In the example Pipeline above, some of the Stages will require interfacing with cloud infrastructure. For example, the canary deployment involves creating ephemeral hosts to run an experiment, and a full deployment of a new version of the service may involve spinning up new servers and then tearing down the old ones. We call these sorts of operations that mutate cloud infrastructure <em>Cloud Operations</em>. Clouddriver’s job is to decompose and execute Cloud Operations sent to it by Orca as part of a deployment. Cloud Operations sent from Orca to Clouddriver are relatively high level (for example: createServerGroup), so Clouddriver understands how to translate these into lower-level cloud provider API calls.</p><p>Pain points in the interaction between Orca and Clouddriver and the implementation details of Cloud Operation execution in Clouddriver are what led us to look for new solutions and ultimately migrate to Temporal, so we’ll next look at the anatomy of a Cloud Operation. Cloud Operations in the OSS version of Spinnaker still work as described below, so motivated readers can follow along in <a href="https://github.com/spinnaker/clouddriver">source code</a>, however our migration to Temporal is entirely closed-source following a fork from OSS in 2020 to allow Netflix to make larger pivots to the product such as this one.</p><h4><strong>The Original Cloud Operation Flow</strong></h4><p>A Cloud Operation’s execution goes something like this:</p><ol><li>Orca, in orchestrating a Pipeline execution, decides a particular Cloud Operation needs to be performed. It sends a POST request to Clouddriver’s /ops endpoint with an untyped bag-of-fields.</li><li>Clouddriver attempts to resolve the operation Orca sent into a set of AtomicOperation s— internal operations that only Clouddriver understands.</li><li>If the payload was valid and Clouddriver successfully resolved the operation, it will immediately return a Task ID to Orca.</li><li>Orca will immediately begin polling Clouddriver’s GET /task/&lt;id&gt; endpoint to keep track of the status of the Cloud Operation.</li><li>Asynchronously, Clouddriver begins executing AtomicOperations using <em>its own</em> internal orchestration engine. Ultimately, the AtomicOperations resolve into cloud provider API calls. As the Cloud Operation progresses, Clouddriver updates an internal state store to surface progress to Orca.</li><li>Eventually, if all went well, Clouddriver will mark the Cloud Operation complete, which eventually surfaces to Orca in its polling. Orca considers the Cloud Operation finished, and the deployment can progress.</li></ol><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*Y57y00EsM2YGRph9IRNmLQ.png"><figcaption>A sequence diagram of a Cloud Operation execution</figcaption></figure><p>This works well enough on the happy path, but veer off the happy path and dragons begin to emerge:</p><ol><li>Clouddriver has its own internal orchestration system independent of Orca to allow Orca to query the progress of Cloud Operation. This is largely undifferentiated lifting relative to Clouddriver’s goal of actuating cloud infrastructure changes, and ultimately adds complexity and surface area for bugs to the application. Additionally, Orca is tightly coupled to Clouddriver’s orchestration system — it must understand how to poll Clouddriver, interpret the status, and handle errors returned by Clouddriver.</li><li>Distributed systems are messy — networks and external services are unreliable. While executing a Cloud Operation, Clouddriver could experience transient network issues, or the cloud provider it’s attempting to call into may be having an outage, or any number of issues in between. Despite all of this, Clouddriver must be as reliable as reasonably possible as a core platform service. To deal with this shape of issue, Clouddriver internally evolved complex retry logic, further adding cognitive complexity to the system.</li><li>Remember how a Cloud Operation gets decomposed by Clouddriver into AtomicOperations? Sometimes, if there’s a failure in the middle of a Cloud Operation, we need to be able to roll back what was done in AtomicOperations prior to the failure. This led to a homegrown Saga framework being implemented inside Clouddriver. While this did result in a big step forward in reliability of Cloud Operations facing transient failures because the Saga framework <em>also</em> allowed replaying partially-failed Cloud Operations, it added yet more undifferentiated lifting inside the service.</li><li>The task state kept by Clouddriver was <em>instance-local</em>. In other words, if the Clouddriver instance carrying out a Cloud Operation crashed, that Cloud Operation state was lost, and Orca would eventually time out polling for the task status. The Saga implementation mentioned above mitigated this for certain operations, but was not widely adopted across all cloud providers supported by Spinnaker.</li></ol><p>We introduced a <em>lot</em> of incidental complexity into Clouddriver in an effort to keep Cloud Operation execution reliable, and despite all this deployments still failed around 4% of the time due to transient Cloud Operation failures.</p><p>Now, I can already hear you saying: “So what? Can’t people re-try their deployments if they fail?” While true, some pipelines take <em>days</em> to complete for complex deployments, and a failed Cloud Operation mid-way through requires re-running the <em>whole</em> thing. This was detrimental to engineering productivity at Netflix in a non-trivial way. Rather than continue trying to build a faster horse, we began to look elsewhere for our reliable orchestration requirements, which is where Temporal comes in.</p><h3>Temporal: Basic Concepts</h3><p>Temporal is an open source product that offers a durable execution platform for your applications. Durable execution means that the platform will ensure your programs run to completion despite adverse conditions. With Temporal, you organize your business logic into <em>Workflows</em>, which are a deterministic series of steps. The steps inside of Workflows are called <em>Activities</em>, which is where you encapsulate all your non-deterministic logic that needs to happen in the course of executing your Workflows. As your Workflows execute in processes called <em>Workers</em>, the Temporal server durably stores their execution state so that in the event of failures your Workflows can be retried or even migrated to a different Worker. This makes Workflows incredibly resilient to the sorts of transient failures Clouddriver was susceptible to. Here’s a simple example Workflow in Java that runs an Activity to send an email once every 30 days:</p><pre>@WorkflowInterface<br>public interface SleepForDaysWorkflow {<br>    @WorkflowMethod<br>    void run();<br>}<br><br>public class SleepForDaysWorkflowImpl implements SleepForDaysWorkflow {<br><br>    private final SendEmailActivities emailActivities =<br>            Workflow.newActivityStub(<br>                    SendEmailActivities.class,<br>                    ActivityOptions.newBuilder()<br>                            .setStartToCloseTimeout(Duration.ofSeconds(10))<br>                            .build());<br><br>    @Override<br>    public void run() {<br>        while (true) {<br>            // Activities already carry retries/timeouts via options.<br>            emailActivities.sendEmail();<br><br>            // Pause the workflow for 30 days before sending the next email.<br>            Workflow.sleep(Duration.ofDays(30));<br>        }<br>    }<br>}<br><br>@ActivityInterface<br>public interface SendEmailActivities {<br>    void sendEmail();<br>}</pre><p>There’s some interesting things to note about this Workflow:</p><ol><li>Workflows and Activities are just code, so you can test them using the same techniques and processes as the rest of your codebase.</li><li>Activities are automatically retried by Temporal with configurable exponential backoff.</li><li>Temporal manages all the execution state of the Workflow, including timers (like the one used by Workflow.sleep). If the Worker executing this workflow were to have its power cable unplugged, Temporal would ensure another Worker continues to execute it (even during the 30 day sleep).</li><li>Workflow sleeps are not compute-intensive, and they don’t tie up the process.</li></ol><p>You might already begin to see how Temporal solves a lot of the problems we had with Clouddriver. Ultimately, we decided to pull the trigger on migrating Cloud Operation execution to Temporal.</p><h3>Cloud Operations with Temporal</h3><p>Today, we execute Cloud Operations as Temporal workflows. Here’s what that looks like.</p><ol><li>Orca, using a Temporal client, sends a request to Temporal to execute an UntypedCloudOperationRunner Workflow. The contract of the Workflow looks something like this:</li></ol><pre>@WorkflowInterface<br>interface UntypedCloudOperationRunner {<br>  /**<br>   * Runs a cloud operation given an untyped payload.<br>   *<br>   * WorkflowResult is a thin wrapper around OutputType providing a standard contract for<br>   * clients to determine if the CloudOperation was successful and fetching any errors.<br>   */<br>  @WorkflowMethod<br>  fun &lt;OutputType : CloudOperationOutput&gt; run(stageContext: Map&lt;String, Any?&gt;, operationType: String): WorkflowResult&lt;OutputType&gt;<br>}</pre><p>2. The Clouddriver Temporal worker is constantly polling Temporal for work. A worker will eventually see a task for an UntypedCloudOperationRunner Workflow and start executing it.</p><p>3. Similar to before with resolution into AtomicOperations, Clouddriver does some pre-processing of the bag-of-fields in stageContext and resolves it to a strongly typed implementation of the CloudOperation Workflow interface based on the operationType input and the stageContext:</p><pre>interface CloudOperation&lt;I : CloudOperationInput, O : CloudOperationOutput&gt; {<br>  @WorkflowMethod<br>  fun operate(input: I, credentials: AccountCredentials&lt;out Any&gt;): O<br>}</pre><p>4. Clouddriver starts a <a href="https://docs.temporal.io/child-workflows">Child Workflow</a> execution of the CloudOperation implementation it resolved. The child workflow will execute Activities which handle the actual cloud provider API calls to mutate infrastructure.</p><p>5. Orca uses its Temporal Client to await completion of the UntypedCloudOperationRunner Workflow. Once it’s complete, Temporal notifies the client and sends the result and Orca can continue progressing the deployment.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*leM3bH8iyb65_cmtl3vm4A.png"><figcaption>Sequence diagram of a Cloud Operation execution with Temporal</figcaption></figure><h3>Results and Lessons Learned from the Migration</h3><p>A shiny new architecture is great, but equally important is the non-glamorous work of refactoring legacy systems to fit the new architecture. How did we integrate Temporal into critical dependencies of all Netflix engineers transparently?</p><p>The answer, of course, is a combination of abstraction and dynamic configuration. We built a CloudOperationRunner interface in Orca to encapsulate whether the Cloud Operation was being executed via the legacy path or Temporal. At runtime, <a href="https://netflixtechblog.com/announcing-archaius-dynamic-properties-in-the-cloud-bc8c51faf675">Fast Properties</a> (Netflix’s dynamic configuration system) determined which path a stage that needed to execute a Cloud Operation would take. We could set these properties quite granularly — by Stage type, cloud provider account, Spinnaker application, Cloud Operation type (createServerGroup), and cloud provider (either AWS or <a href="https://netflix.github.io/titus/">Titus</a> in our case). The Spinnaker services themselves were the first to be deployed using Temporal, and within two quarters, all applications at Netflix were onboarded.</p><h4>Impact</h4><p>What did we have to show for it all? With Temporal as the orchestration engine for Cloud Operations, the percentage of deployments that failed due to transient Cloud Operation failures dropped from 4% to 0.0001%. For those keeping track at home, that’s a four and a half order of magnitude reduction. Virtually eliminating this failure mode for deployments was a huge win for developer productivity, especially for teams with long and complex deployment pipelines.</p><p>Beyond the improvement in deployment success metrics, we saw a number of other benefits:</p><ol><li>Orca no longer needs to directly communicate with Clouddriver to start Cloud Operations or poll their status with Temporal as the intermediary. The services are less coupled, which is a win for maintainability.</li><li>Speaking of maintainability, with Temporal doing the heavy lifting of orchestration and retries inside of Clouddriver, we got to remove a lot of the homegrown logic we’d built up over the years for the same purpose.</li><li>Since Temporal manages execution state, Clouddriver instances became stateless and Cloud Operation execution can bounce between instances with impunity. We can treat Clouddriver instances more like cattle and enable things like <a href="https://netflixtechblog.com/the-netflix-simian-army-16e57fbab116">Chaos Monkey</a> for the service which we were previously prevented from doing.</li><li>Migrating Cloud Operation steps into Activities was a forcing function to re-write the logic to be idempotent. Since Temporal retries activities by default, it’s generally recommended they be idempotent. This alone fixed a number of issues that existed previously when operations were retried in Clouddriver.</li><li>We set the retry timeout for Activities in Clouddriver to be two hours by default. This gives us a long leash to fix-forward or rollback Clouddriver if we introduce a regression before customer deployments fail — to them, it might just look like a deployment is taking longer than usual.</li><li>Cloud Operations are much easier to introspect than before. Temporal ships with a great UI to help visualize Workflow and Activity executions, which is a huge boon for debugging live Workflows executing in production. The Temporal SDKs and server also emit a lot of useful metrics.</li></ol><figure><img alt="A Cloud Operation Workflow as seen from the Temporal UI. This operation executes 3 Activities: DescribeAutoScalingGroup, GetHookConfigurations, and ResizeServerGroup" src="https://cdn-images-1.medium.com/max/1024/1*zmCyjwzTXji921mulJjmTw.png"><figcaption>Execution of a resizeServerGroup Cloud Operation as seen from the Temporal UI. This operation executes 3 Activities: DescribeAutoScalingGroup, GetHookConfigurations, and ResizeServerGroup</figcaption></figure><h4>Lessons Learned</h4><p>With the benefit of hindsight, there are also some lessons we can share from this migration:</p><p>1. <strong>Avoid unnecessary Child Workflows</strong>: Structuring Cloud Operations as an UntypedCloudOperationRunner Workflow that starts Child Workflows to actually execute the Cloud Operation’s logic was unnecessary and the indirection made troubleshooting more difficult. There are <a href="https://community.temporal.io/t/purpose-of-child-workflows/652">situations</a> where Child Workflows are appropriate, but in this case we were using them as a tool for code organization, which is generally unnecessary. We could’ve achieved the same effect with class composition in the top-level parent Workflow.</p><p>2. <strong>Use single argument objects</strong>: At first, we structured Workflow and Activity functions with variable arguments, much as you’d write normal functions. This can be problematic for Temporal because of Temporal’s <a href="https://community.temporal.io/t/workflow-determinism/4027">determinism constraints</a>. Adding or removing an argument from a function signature is <strong>not</strong> a backward-compatible change, and doing so can break long-running workflows — and it’s not immediately obvious in code review your change is problematic. The preferred pattern is to use a single serializable class to host all your arguments for Workflows and Activities — these can be more freely changed without breaking determinism.</p><p>3. <strong>Separate business failures from workflow failures</strong>: We like the pattern of the WorkflowResult type that UntypedCloudOperationRunner returns in the interface above. It allows us to communicate business process failures without failing the Workflow itself and have more overall nuance in error handling. This is a pattern we’ve carried over to other Workflows we’ve implemented since.</p><h3>Temporal at Netflix Today</h3><p>Temporal adoption has skyrocketed at Netflix since its initial introduction for Spinnaker. Today, we have hundreds of use cases, and we’ve seen adoption double in the last year with no signs of slowing down.</p><p>One major difference between initial adoption and today is that Netflix migrated from an on-prem Temporal deployment to using <a href="https://temporal.io/cloud">Temporal Cloud</a>, which is Temporal’s SaaS offering of the Temporal server. This has let us scale Temporal adoption while running a lean team. We’ve also built up a robust internal platform around Temporal Cloud to integrate with Netflix’s internal ecosystem and make onboarding for our developers as easy as possible. Stay tuned for a future post digging into more specifics of our Netflix Temporal platform.</p><h3>Acknowledgement</h3><p>We all stand on the shoulders of giants in software. I want to call out that I’m retelling the work of my two stunning colleagues <a href="https://www.linkedin.com/in/chris-smalley/">Chris Smalley</a> and <a href="https://www.linkedin.com/in/robzienert/">Rob Zienert</a> in this post, who were the two aforementioned engineers who introduced Temporal and carried out the migration.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=73c69ccb5953" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/how-temporal-powers-reliable-cloud-operations-at-netflix-73c69ccb5953">How Temporal Powers Reliable Cloud Operations at Netflix</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/how-temporal-powers-reliable-cloud-operations-at-netflix-73c69ccb5953</link>
      <guid>https://netflixtechblog.com/how-temporal-powers-reliable-cloud-operations-at-netflix-73c69ccb5953</guid>
      <pubDate>Tue, 16 Dec 2025 00:51:00 +0100</pubDate>
    </item>
    <item>
      <title><![CDATA[Netflix Live Origin]]></title>
      <description><![CDATA[<p><a href="https://www.linkedin.com/in/xiaomei-liu-b475711/">Xiaomei Liu</a>, <a href="https://www.linkedin.com/in/joseph-lynch-9976a431/">Joseph Lynch</a>, <a href="https://www.linkedin.com/in/chrisnewton2/">Chris Newton</a></p><h3>Introduction</h3><p><a href="https://netflixtechblog.com/building-a-reliable-cloud-live-streaming-pipeline-for-netflix-8627c608c967">Behind the Streams: Building a Reliable Cloud Live Streaming Pipeline for Netflix</a> introduced the architecture of the streaming pipeline. This blog post looks at the custom Origin Server we built for Live — the Netflix Live Origin. It sits at the demarcation point between the cloud live streaming pipelines on its upstream side and the distribution system, Open Connect, Netflix’s in-house Content Delivery Network (CDN), on its downstream side, and acts as a broker managing what content makes it out to Open Connect and ultimately to the client devices.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*44sJszKXEHZvSHnEQgYiiw.png"><figcaption><strong>Live Streaming Distribution and Origin Architecture</strong></figcaption></figure><p>Netflix Live Origin is a multi-tenant microservice operating on EC2 instances within the AWS cloud. We lean on standard HTTP protocol features to communicate with the Live Origin. The Packager pushes segments to it using PUT requests, which place a file into storage at the particular location named in the URL. The storage location corresponds to the URL that is used when the Open Connect side issues the corresponding GET request.</p><p>Live Origin architecture is influenced by key technical decisions of the live streaming architecture. First, resilience is achieved through redundant regional live streaming pipelines, with failover orchestrated at the server-side to reduce client complexity. The implementation of <a href="https://netflixtechblog.com/building-a-reliable-cloud-live-streaming-pipeline-for-netflix-8627c608c967">epoch locking at the cloud encoder</a> enables the origin to select a segment from either encoding pipeline. Second, Netflix adopted a manifest design with <a href="https://netflixtechblog.com/behind-the-streams-live-at-netflix-part-1-d23f917c2f40">segment templates and constant segment duration</a> to avoid frequent manifest refresh. The constant duration templates enable Origin to predict the segment publishing schedule.</p><h3>Multi-pipeline and multi-region aware origin</h3><p>Live streams inevitably contain defects due to the non-deterministic nature of live contribution feeds and strict real-time segment publishing timelines. Common defects include:</p><ul><li><strong>Short segments:</strong> Missing video frames and audio samples.</li><li><strong>Missing segments:</strong> Entire segments are absent.</li><li><strong>Segment timing discontinuity:</strong> Issues with the Track Fragment Decode Time.</li></ul><p>Communicating segment discontinuity from the server to the client via a segment template-based manifest is impractical, and these defective segments can disrupt client streaming.</p><p>The redundant cloud streaming pipelines operate independently, encompassing distinct cloud regions, contribution feeds, encoder, and packager deployments. This independence substantially mitigates the probability of simultaneous defective segments across the dual pipelines. Owing to its strategic placement within the distribution path, the live origin naturally emerges as a component capable of intelligent candidate selection.</p><p>The Netflix Live Origin features multi-pipeline and multi-region awareness. When a segment is requested, the live origin checks candidates from each pipeline in a deterministic order, selecting the first valid one. Segment defects are detected via lightweight media inspection at the packager. This defect information is provided as metadata when the segment is published to the live origin. In the rare case of concurrent defects at the dual pipeline, the segment defects can be communicated downstream for intelligent client-side error concealment.</p><h3>Open Connect streaming optimization</h3><p>When the Live project started, Open Connect had become highly optimised for VOD content delivery — <a href="https://freenginx.org/en/">nginx</a> had been chosen many years ago as the Web Server since it is highly capable in this role, and a number of enhancements had been added to it and to the underlying operating system (BSD). Unlike traditional CDNs, Open Connect is more of a distributed origin server — VOD assets are pre-positioned onto carefully selected server machines (OCAs, or Open Connect Appliances) rather than being filled on demand.</p><p>Alongside the VOD delivery, an on-demand fill system has been used for non-VOD assets — this includes artwork and the downloadable portions of the clients, etc. These are also served out of the same <a href="https://freenginx.org/en/">nginx</a> workers, albeit under a distinct server block, using a distinct set of hostnames.</p><p>Live didn’t fit neatly into this ‘small object delivery’ model, so we extended the proxy-caching functionality of <a href="https://freenginx.org/en/">nginx</a> to address Live-specific needs. We will touch on some of these here related to optimized interactions with the Origin Server. Look for a future blog post that will go into more details on the Open Connect side.</p><p>The segment templates provided to clients are also provided to the OCAs as part of the Live Event Configuration data. Using the Availability Start Time and Initial Segment number, the OCA is able to determine the legitimate range of segments for each event at any point in time — requests for objects outside this range can be rejected, preventing unnecessary requests going up through the fill hierarchy to the origin. If a request makes it through to the origin, and the segment isn’t available yet, the origin server will return a 404 Status Code (indicating File Not Found) with the expiration policy of that error so that it can be cached within Open Connect until just before that segment is expected to be published.</p><p>If the Live Origin knows when segments are being pushed to it, and knows what the live edge is — when a request is received for the immediately next object, rather than handing back another 404 error (which would go all the way back through Open Connect to the client), the Live Origin can ‘hold open’ the request, and service it once the segment has been published to it. By doing this, the degree of chatter within the network handling requests that arrive early has been significantly reduced. As part of this, millisecond grain caching was added to <a href="https://freenginx.org/en/">nginx</a> to enhance the standard HTTP Cache Control, which only works at second granularity, a long time when segments are generated every 2 seconds.</p><h4>Streaming metadata enhancement</h4><p>The HTTP standard allows for the addition of request and response headers that can be used to provide additional information as files move between clients and servers. The HTTP headers provide notifications of events within the stream in a highly scalable way that is independently conveyed to client devices, regardless of their playback position within the stream.</p><p>These notifications are provided to the origin by the live streaming pipeline and are inserted by the origin in the form of headers, appearing on the segments generated at that point in time (and persist to future segments — they are cumulative). Whenever a segment is received at an OCA, this notification information is extracted from the response headers and used to update an in-memory data structure, keyed by event ID; and whenever a segment is served from the OCA, the latest such notification data is attached to the response. This means that, given any flow of segments into an OCA, it will always have the most recent notification data, even if all clients requesting it are behind the live edge. In fact, the notification information can be conveyed on any response, not just those supplying new segments.</p><h4>Cache invalidation and origin mask</h4><p>An invalidation system has been available since the early days of the project. It can be used to “flush” all content associated with an event by altering the key used when looking up objects in cache — this is done by incorporating a version number into the cache key that can then be bumped on demand. This is used during pre-event testing so that the network can be returned to a pristine state for the test with minimal fuss.</p><p>Each segment published by the Live Origin conveys the encoding pipeline it was generated by, as well as the region it was requested from. Any issues that are found after segments make their way into the network can be remedied by an enhanced invalidation system that takes such variants into account. It is possible to invalidate (that is, cause to be considered expired) segments in a range of segment numbers, but only if they were sourced from encoder A, or from Encoder A, but only if retrieved from region X.</p><p>In combination with Open Connect’s enhanced cache invalidation, the Netflix Live Origin allows <em>selective encoding pipeline masking</em> to exclude a range of segments from a particular pipeline when serving segments to Open Connect. The enhanced cache invalidation and origin masking enable live streaming operations to hide known problematic segments (e.g., segments causing client playback errors) from streaming clients once the bad segments are detected, protecting millions of streaming clients during the DVR playback window.</p><h3>Origin storage architecture</h3><p>Our original storage architecture for the Live Origin was simple: just use <a href="https://aws.amazon.com/s3/">AWS S3</a> like we do for SVOD. This served us well initially for our low-traffic events, but as we scaled up we discovered that Live streaming has unique latency and workload requirements that differ significantly from on-demand where we have significant time ahead-of-time to pre-position content. While S3 met its stated uptime guarantees, our strict 2-second retry budget inherent to Live events (where every write is critical) led us to explore optimizations specifically tailored for real-time delivery at scale. AWS S3 is an amazing object store, but our Live streaming requirements were closer to those of a global low-latency highly-available database. So, we went back to the drawing board and started from the requirements. The Origin required:</p><ol><li>[HA Writes] Extremely high <em>write</em> availability, ideally as close to full write availability within a single AWS region, with low second replication delay to other regions. Any failed write operation within 500ms is considered a bug that must be triaged and prevented from re-occurring.</li><li>[Throughput] High write throughput, with hundreds of MiB replicating across regions</li><li>[Large Partitions] Efficiently support O(MiB) writes that accumulate to O(10k) keys per partition with O(GiB) total size per event.</li><li>[Strong Consistency] Within the same region, we needed read-your-write semantics to hit our &lt;1s read delay requirements (must be able to read published segments)</li><li>[Origin Storm] During worst-case load involving Open Connect edge cases, we may need to handle O(<strong>GiB</strong>) of read throughput <em>without affecting writes</em>.</li></ol><p>Fortunately, Netflix had previously invested in building a <a href="https://netflixtechblog.com/introducing-netflixs-key-value-data-abstraction-layer-1ea8a0a11b30">KeyValue Storage Abstraction</a> that cleverly leveraged <a href="https://youtu.be/sQ-_jFgOBng?t=1061">Apache Cassandra</a> to provide chunked storage of MiB or even GiB values. This abstraction was initially built to support cloud saves of Game state. The Live use case would push the boundaries of this solution, however, in terms of availability for writes (#1), cumulative partition size (#3), and read throughput during Origin Storm (#5).</p><h4>High Availability for Writes of Large Payloads</h4><p>The <a href="https://youtu.be/paTtLhZFsGE?t=1077">KeyValue Payload Chunking and Compression Algorithm</a> breaks O(MiB) work down so each part can be idempotently retried and hedged to maintain strict latency service level objectives, as well as spreading the data across the full cluster. When we combine this algorithm with Apache Cassandra’s local-quorum consistency model, which allows write availability even with an entire Availability Zone outage, plus a write-optimized <a href="https://en.wikipedia.org/wiki/Log-structured_merge-tree">Log-Structured Merge Tree</a> (LSM) storage engine, we could meet the first four requirements. After iterating on the performance and availability of this solution, we were not only able to achieve the write availability required, but did so with a P99 <em>tail</em> latency that was similar to the status quo’s P50 <em>average </em>latency while also handling cross-region replication behind the scenes for the Origin. This new solution was significantly more expensive (as expected, databases backed by SSD cost more), but minimizing cost was <em>not</em> a key objective and low latency with high availability was:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*bUPc4gC-mSDcybayhBJQ8g.png"><figcaption><strong>Storage System Write Performance</strong></figcaption></figure><h4>High Availability Reads at Gbps Throughputs</h4><p>Now that we solved the write reliability problem, we had to handle the Origin Storm failure case, where potentially dozens of Open Connect top-tier caches could be requesting multiple O(MiB) video segments at once. Our back-of-the-envelope calculations showed worst-case read throughput in the O(100Gbps) range, which would normally be extremely expensive for a strongly-consistent storage engine like Apache Cassandra. With careful tuning of chunk access, we were able to respond to reads at network line rate (100Gbps) from Apache Cassandra, but we observed unacceptable performance and availability degradation on concurrent writes. To resolve this issue, we introduced write-through caching of chunks using our distributed caching system <a href="https://github.com/Netflix/EVCache">EVCache</a>, which is based on Memcached. This allows almost all reads to be served from a highly scalable cache, allowing us to easily hit 200Gbps and beyond without affecting the write path, achieving read-write separation.</p><h4>Final Storage Architecture</h4><p>In the final storage architecture, the Live Origin writes and reads to KeyValue, which manages a write-through cache to EVCache (memcached) and implements a safe chunking protocol that spreads large values and partitions them out across the storage cluster (Apache Cassandra). This allows almost all read load to be handled from cache, with only misses hitting the storage. This combination of cache and highly available storage has met the demanding needs of our Live Origin for over a year now.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*yA9H-BFemM_-99FXBlVOMg.png"><figcaption><strong>Storage System High Level Architecture</strong></figcaption></figure><p>Delivering this consistent low latency for large writes with cross-region replication and consistent write-through caching to a distributed cache required solving numerous hard problems with novel techniques, which we plan to share in detail during a future post.</p><h3>Scalability and scalable architecture</h3><p>Netflix’s live streaming platform must handle a high volume of diverse stream renditions for each live event. This complexity stems from supporting various video encoding formats (each with multiple encoder ladders), numerous audio options (across languages, formats, and bitrates), and different content versions (e.g., with or without advertisements). The combination of these elements, alongside concurrent event support, leads to a significant number of unique stream renditions per live event. This, in turn, necessitates a high Requests Per Second (RPS) capacity from the multi-tenant live origin service to ensure publishing-side scalability.</p><p>In addition, Netflix’s global reach presents distinct challenges to the live origin on the retrieval side. During the Tyson vs. Paul fight event in 2024, a historic peak of 65 million concurrent streams was observed. Consequently, a scalable architecture for live origin is essential for the success of large-scale live streaming.</p><h4>Scaling architecture</h4><p>We chose to build a highly scalable origin instead of relying on the traditional origin shields approach for better end-to-end cache consistency control and simpler system architecture. The live origin in this architecture directly connects with top-tier Open Connect nodes, which are geographically distributed across several sites. To minimize the load on the origin, only designated nodes per stream rendition at each site are permitted to directly fill from the origin.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*jW7eBCQtjlna0VKaWrKz_A.png"><figcaption><strong>Netflix Live Origin Scalability Architecture</strong></figcaption></figure><p>While the origin service can autoscale horizontally using EC2 instances, there are other system resources that are not autoscalable, such as storage platform capacity and AWS to Open Connect backbone bandwidth capacity. Since in live streaming, not all requests to the live origin are of the same importance, the origin is designed to prioritize more critical requests over less critical requests when system resources are limited. The table below outlines the request categories, their identification, and protection methods.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*dYKFJkq22KI8sDBW_Njmog.png"></figure><h4>Publishing isolation</h4><p>Publishing traffic, unlike potentially surging CDN retrieval traffic, is predictable, making path isolation a highly effective solution. As shown in the scalability architecture diagram, the origin utilizes separate EC2 publishing and CDN stacks to protect the latency and failure-sensitive origin writes. In addition, the storage abstraction layer features distinct clusters for key-value (KV) read and KV write operations. Finally, the storage layer itself separates read (EVCache) and write (Cassandra) paths. This comprehensive path isolation facilitates independent cloud scaling of publishing and retrieval, and also prevents CDN-facing traffic surges from impacting the performance and reliability of origin publishing.</p><h4>Priority rate limiting</h4><p>Given Netflix’s scale, managing incoming requests during a traffic storm is challenging, especially considering non-autoscalable system resources. The Netflix Live Origin implemented priority-based rate limiting when the underlying system is under stress. This approach ensures that requests with greater user impact are prioritized to succeed, while requests with lower user impact are allowed to fail during times of stress in order to protect the streaming infrastructure and are permitted to retry later to succeed.</p><p>Leveraging Netflix’s microservice platform priority rate limiting feature, the origin prioritizes live edge traffic over DVR traffic during periods of high load on the storage platform. The live edge vs. DVR traffic detection is based on the predictable segment template. The template is further cached in memory on the origin node to enable priority rate limiting without access to the datastore, which is valuable especially during periods of high datastore stress.</p><p>To mitigate traffic surges, TTL cache control is used alongside priority rate limiting. When the low-priority traffic is impacted, the origin instructs Open Connect to slow down and cache identical requests for 5 seconds by setting a max-age = 5s and returns an HTTP 503 error code. This strategy effectively dampens traffic surges by preventing repeated requests to the origin within that 5-second window.</p><p>The following diagrams illustrate origin priority rate limiting with simulated traffic. The nliveorigin_mp41 traffic is the low-priority traffic and is mixed with other high-priority traffic. In the first row: the 1st diagram shows the request RPS, the 2nd diagram shows the percentage of request failure. In the second row, the 1st diagram shows datastore resource utilization, and the 2nd diagram shows the origin retrieval P99 latency. The results clearly show that only the low-priority traffic (nliveorigin_mp41) is impacted at datastore high utilization, and the origin request latency is under control.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*-_YP6H3sEDaw1lS8prH4mQ.png"><figcaption><strong>Origin Priority Rate Limiting</strong></figcaption></figure><h4>404 storm and cache optimization</h4><p>Publishing isolation and priority rate limiting successfully protect the live origin from DVR traffic storms. However, the traffic storm generated by requests for non-existent segments presents further challenges and opportunities for optimization.</p><p>The live origin structures metadata hierarchically as event &gt; stream rendition &gt; segment, and the segment publishing template is maintained at the stream rendition level. This hierarchical organization allows the origin to preemptively reject requests with an HTTP 404(not found)/410(Gone) error, leveraging highly cacheable event and stream rendition level metadata, avoiding unnecessary queries to the segment level metadata:</p><ul><li>If the event is unknown, reject the request with 404</li><li>If the event is known, but the segment request timing does not match the expected publishing timing, reject the request with 404 and cache control TTL matching the expected publishing time</li><li>If the event is known, the requested segment is never generated or misses the retry deadline, reject the request with a 410 error, preventing the client from repeatedly requesting</li></ul><p>At the storage layer, metadata is stored separately from media data in the control plane datastore. Unlike the media datastore, the control plane datastore does not use a distributed cache to avoid cache inconsistency. Event and rendition level metadata benefits from a high cache hit ratio when in-memory caching is utilized at the live origin instance. During traffic storms involving non-existent segments, the cache hit ratio for control plane access easily exceeds 90%.</p><p>The use of in-memory caching for metadata effectively handles 404 storms at the live origin without causing datastore stress. This metadata caching complements the storage system’s distributed media cache, providing a complete solution for traffic surge protection.</p><h3>Summary</h3><p>The Netflix Live Origin, built upon an optimized storage platform, is specifically designed for live streaming. It incorporates advanced media and segment publishing scheduling awareness and leverages enhanced intelligence to improve streaming quality, optimize scalability, and improve Open Connect live streaming operations.</p><h3>Acknowledgement</h3><p>Many teams and stunning colleagues contributed to the Netflix live origin. Special thanks to <a href="https://www.linkedin.com/in/flavioribeiro/?originalSubdomain=br">Flavio Ribeiro</a> for advocacy and sponsorship of the live origin project; to <a href="https://www.linkedin.com/in/rummadis/">Raj Ummadisetty</a>, <a href="https://www.linkedin.com/in/prudhviraj9/">Prudhviraj Karumanchi</a> for the storage platform; to <a href="https://www.linkedin.com/in/rosanna-lee-197920/">Rosanna Lee</a>, <a href="https://www.linkedin.com/in/hunterford/">Hunter Ford</a>, and <a href="https://www.linkedin.com/in/thiagopnts/">Thiago Pontes</a> for storage lifecycle management; to <a href="https://www.linkedin.com/in/ameya-vasani-8904304/">Ameya Vasani</a> for e2e test framework; <a href="https://www.linkedin.com/in/thomas-symborski-b4216728/">Thomas Symborski</a> for orchestrator integration; to <a href="https://www.linkedin.com/in/jschek/">James Schek</a> for Open Connect integration; to <a href="https://www.linkedin.com/in/kzwang/">Kevin Wang</a> for platform priority rate limit; to <a href="https://www.linkedin.com/in/di-li-09663968/">Di Li</a>, <a href="mailto:nhubbard@netflix.com">Nathan Hubbard</a> for origin scalability testing.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=41f1b0ad5371" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/netflix-live-origin-41f1b0ad5371">Netflix Live Origin</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/netflix-live-origin-41f1b0ad5371</link>
      <guid>https://netflixtechblog.com/netflix-live-origin-41f1b0ad5371</guid>
      <pubDate>Mon, 15 Dec 2025 18:38:00 +0100</pubDate>
    </item>
    <item>
      <title><![CDATA[AV1 — Now Powering 30% of Netflix Streaming]]></title>
      <description><![CDATA[<h3><strong>AV1 — Now Powering 30% of Netflix Streaming</strong></h3><p><a href="https://www.linkedin.com/in/liwei-guo/">Liwei Guo</a>, <a href="https://www.linkedin.com/in/henryzhili/">Zhi Li</a>, <a href="https://www.linkedin.com/in/sheldon-radford/">Sheldon Radford</a>, <a href="https://www.linkedin.com/in/jeffrwatts/">Jeff Watts</a></p><p>Streaming video has become an integral part of our daily lives. At Netflix, our top priority is delivering the best possible entertainment experience to our members, regardless of their devices or network conditions. One of the key technologies enabling this is <a href="https://aomedia.org/specifications/av1/">AV1</a>, a modern, open video codec that is rapidly transforming both how we stream content and how users experience it. Today, AV1 powers approximately 30% of all Netflix viewing, marking a major milestone in our efforts to bring more efficient and higher-quality streaming to our members.</p><p>In this post, we’ll revisit Netflix’s AV1 journey to date, highlight emerging use cases, and share adoption trends across the device ecosystem. Having witnessed AV1’s significant impact，and with <a href="https://aomedia.org/press%20releases/AOMedia-Announces-Year-End-Launch-of-Next-Generation-Video-Codec-AV2-on-10th-Anniversary/">AV2 on the horizon</a>, we’re more excited than ever about how open codecs will continue to revolutionize streaming for everyone.</p><h3>AV1: A Modern, Open Codec</h3><p>Since entering the streaming business in 2007, Netflix has primarily relied on H.264/AVC as its streaming format. However, we quickly recognized that a modern, open codec would benefit not only Netflix, but the entire multimedia industry. In 2015, together with a group of like-minded industry leaders, Netflix co-founded the <a href="https://aomedia.org/">Alliance for Open Media (AOMedia)</a> to develop and promote next generation, open source media technologies. The AV1 codec became the first major project of this collaboration, with ambitious goals: to deliver significant improvements in compression efficiency over state-of-the-art codecs, and to introduce rich features that enable new use cases. After three years of collaborative development, AV1 was officially released in 2018.</p><h3>Netflix’s AV1 Journey: From Android to TVs and Beyond</h3><h4><strong>Piloting on Android Mobile</strong></h4><p>When we first set out to bring AV1 streaming to Netflix members, Android was the ideal starting point. Android’s flexibility allowed us to quickly integrate a software AV1 decoder using the efficient <a href="https://code.videolan.org/videolan/dav1d">dav1d</a> library, which was already optimized for ARM chipsets in mobile devices.</p><p>AV1’s superior compression efficiency was especially valuable for mobile users, many of whom are mindful of their data usage and network conditions. By adopting AV1, we were able to deliver noticeably better video quality at lower bitrates. For members relying on cellular data, this meant crisper images with fewer compression artifacts, even when bandwidth was limited. <a href="https://netflixtechblog.com/netflix-now-streaming-av1-on-android-d5264a515202">Launching AV1 support on Android</a> in 2020 marked a significant step forward for Netflix on mobile, making high-quality streaming more accessible and enjoyable for members everywhere.</p><h4><strong>Front-and-Center for Netflix VOD Streaming</strong></h4><p>The success of our AV1 launch on Android proved its value for Netflix streaming, motivating us to expand support to smart TVs and other large-screen devices, where most of our members watch their favorite shows.</p><p>Smart TVs depend on hardware decoders for efficient high-quality playback. We worked closely with device manufacturers and SoC vendors to certify these devices, ensuring they are both conformant and performant. This collaborative effort enabled our AV1 streaming to TV devices in <a href="https://netflixtechblog.com/bringing-av1-streaming-to-netflix-members-tvs-b7fc88e42320">late 2021</a>. Shortly thereafter, we expanded AV1 streaming to web browsers (in 2022) and continued to broaden device support. In 2023, this included Apple devices with the introduction of AV1 hardware support in the new M3 and A17 Pro chips.</p><p>As more devices began shipping with AV1 hardware support, a rapidly growing share of our members could enjoy the benefits of this advanced codec. Combined with our investment in adding AV1 streams across the entire catalog, AV1 viewing share has been consistently increasing in recent years. Today, AV1 accounts for approximately 30% of all Netflix streaming, making it our second most-used codec — and it’s on track to become number one very soon. The payoff has been substantial.</p><ul><li><strong>Elevating Streaming Experience Across the Board</strong>: Large-screen TVs and other devices demand higher bitrates to deliver stunning 4K, high frame rate (HFR) experiences. AV1’s superior compression efficiency has allowed us to provide these experiences using less data, making high-quality streaming more accessible and reliable. On average, AV1 streaming sessions achieve VMAF scores¹ that are 4.3 points higher than AVC and 0.9 points higher than HEVC sessions. At the same time, AV1 sessions use one-third less bandwidth than both AVC and HEVC, resulting in 45% fewer buffering interruptions. Moreover, Netflix’s diverse content catalog benefits universally from AV1, with improvements across all content types.</li><li><strong>Driving Network Efficiency Worldwide</strong>: Netflix streams are delivered through our own content delivery network (<a href="https://openconnect.netflix.com/en/?utm_referrer=https://www.google.com/">Open Connect</a>), in partnership with local ISPs around the globe. With more than 300 million members, Netflix streaming constitutes a non-trivial portion of global internet traffic. Because AV1 is a more efficient codec, its streams are smaller in size (while providing even better visual quality). By shifting a substantial share of our streaming to AV1, we reduce overall internet bandwidth consumption, and lessen system and network load for both Netflix and our partners.</li></ul><h4>Unlocking Advanced Experiences</h4><p>In addition to its superior compression efficiency, AV1 was designed to support a rich set of features. Once we established a robust framework for the continuous expansion of AV1 streaming, we quickly shifted our focus towards exploring AV1’s unique features to unlock even more advanced and immersive experiences for our members.</p><p><strong>High-Dynamic-Range(HDR)<br></strong>HDR brings enhanced detail, vivid colors, and greater clarity to images. As a premium streaming service, Netflix has been a pioneer in adopting HDR, offering HDR streaming since 2016. In March 2025, we launched <a href="https://netflixtechblog.com/hdr10-now-streaming-on-netflix-c9ab1f4bd72b">AV1 HDR streaming</a>. We chose HDR10+ as the HDR format for its use of dynamic metadata, which enabled us to adapt the tone mapping per device in a scene-dependent manner.</p><p>As anticipated, the combination of AV1 and HDR10+ allows us to deliver images with greater detail, more vibrant colors, and an overall heightened sense of immersion for our members. At the moment, 85% of our HDR catalog (from the perspective of view-hours) has AV1-HDR10+ coverage, and this number is expected to reach 100% in the next couple of months.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/759/1*Ubhj9prgqb0zuTHt6oOx0g.png"><figcaption><strong><em>Photographs of devices displaying the same (cropped) frame with HDR10 metadata (left) and HDR10+ metadata (right). Notice the preservation of the flashlight detail in the HDR10+ capture, and the over-exposure of the region under the flashlight in the HDR10 one.</em></strong></figcaption></figure><p><strong>Cinematic Film Grain<br></strong>Film grain is a hallmark of the cinematic experience, widely used in the movie industry to enhance a film’s depth, texture, and realism. However, because film grain is inherently random, faithfully representing it in digital video requires a significant amount of data. This presents a unique challenge for streaming: restricting the bitrate can result in grain that appears unnatural or distorted, while increasing the bitrate to accurately preserve cinematic grain almost inevitably leads to elevated rebuffering. The AV1 specification incorporates a unique solution called Film Grain Synthesis (FGS). Instead of encoding grain as part of every frame, the grain is stripped out before encoding and then resynthesized at the decoder using parameters sent in the bitstream, delivering a realistic cinematic film grain experience without the usual data costs.</p><p>This approach represents a significant shift from traditional compression and streaming techniques. Our team invested substantial effort in fine-tuning the media processing pipeline, ensuring FGS delivers robust performance at scale. In July 2025, we successfully <a href="https://netflixtechblog.com/av1-scale-film-grain-synthesis-the-awakening-ee09cfdff40b">productized AV1 FGS</a>, and the results were astonishing: AV1 with FGS could deliver videos with cinematic film grain at a bitrate well within the capabilities of typical household internet connections. For non-FGS AV1 encodings, even at much higher bitrate, they may not be able to achieve comparable quality.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1016/1*9fB5xuoFpbpN8ZQIzTHDtg.png"><figcaption><strong><em>The same (cropped) frame from source (left), regular AV1 stream encoded at 8274kbps (middle) and AV1 FGS stream encoded at 2804 kbps (right). The AV1 FGS stream reduces the bitrate by 66% while delivering clearly better quality.</em></strong></figcaption></figure><h4><strong>Beyond VOD Streaming</strong></h4><p>So far, our AV1 journey has been mainly on VOD, but we see significant opportunities for AV1 beyond traditional VOD streaming. On a mission to entertain the world, Netflix has constantly explored and established other ways to bring joy to our members, and we believe AV1 could contribute to the success of these new products.</p><p><strong>Live Streaming<br></strong>Debuting in 2023, live streaming has experienced <a href="https://help.netflix.com/en/node/129840">rapid growth</a> at Netflix, becoming a key part of our streaming offerings in just two short years. We are actively evaluating the use of AV1 in live streaming, as we believe it could help further scale Netflix’s live programming:</p><ul><li><strong>Hyper-scale concurrent viewership: </strong>Live streaming at Netflix means delivering content to <a href="https://www.netflix.com/tudum/articles/jake-paul-vs-mike-tyson-live-release-date-news">tens of millions</a> of viewers simultaneously. AV1’s superior compression efficiency could significantly reduce the required bandwidth, enabling us to deliver high-quality live experiences to large audiences without compromising video quality.</li><li><strong>Customizable graphics overlay</strong>: for live sport events such as football, tennis and boxing, graphics overlays have become an integral part of the member experience — from embedding game statistics to delivering sponsorships. AV1 offers an opportunity to make the graphics highly customizable: layered coding is supported in AV1’s main profile, allowing encoding the main content in the base layer, and graphics in the enhancement layer, and easily swapping out one version of the enhancement layer with another. We envision that the use of AV1’s layered coding can greatly simplify the live streaming workflow and reduce delivery costs.</li></ul><p><strong>Cloud Gaming<br></strong>Cloud gaming is a new Netflix offering that is currently in the <a href="https://help.netflix.com/en/node/132197">beta phase</a> and is available to members in select countries. The game engines run on cloud servers, while the rendered graphics are streamed directly to members’ devices. By removing barriers and transforming every Netflix-enabled device into a game console, Cloud gaming aims to deliver a seamless, “play anywhere” experience for our members. For a glimpse of this in action, <a href="https://www.linkedin.com/feed/update/urn:li:activity:7382077927875825664/">watch as Co-CEO Greg Peters and CTO Elizabeth Stone play a round of Boggle Party — powered entirely by Netflix’s cloud gaming platform</a>!</p><p>Unlike traditional video streaming, cloud gaming requires that every player action is reflected instantly on the screen to ensure a responsive and immersive experience. This makes delivering high-quality video frames with extremely low latency, despite fluctuating network conditions, one of the biggest challenges in cloud gaming.</p><p>Our team is actively working on productizing AV1 for cloud gaming. Given AV1’s high compression efficiency, we can reduce frame sizes, helping video frames get through even when network conditions become challenging. This positions AV1 as a promising technology for enabling a high-quality, low-latency gaming experience across a wide range of devices.</p><h3>A Device Ecosystem United for AV1</h3><p>Netflix is a streaming company, and we have worked diligently to create highly efficient and standards-conformant AV1 streams for our catalog. However, an equally, if not more, important factor in AV1’s success is the widespread support from device manufacturers. Throughout our AV1 journey, we have been impressed by the unprecedented pace at which the device ecosystem has embraced AV1.</p><p>Just six months after the AV1 specification was finalized, the open-source AV1 decoder library sponsored by AOM, dav1d, was released. Small, performant, and highly resource-efficient, dav1d bridged the gap for early adopters like Netflix while hardware solutions were still in development. Continuous improvements to its performance and compatibility have made dav1d the preferred choice for a wide range of platforms and practical applications. Today, it serves as <a href="https://aomedia.org/av1-adoption-showcase/google-story/">Android’s default software decoder</a>. Additionally, it plays a key role in web browsers — for Netflix, it powers approximately 40% of our browser playback. This broad adoption has significantly expanded access to high-quality AV1 streaming, even in the absence of dedicated hardware decoders.</p><p>Netflix maintains a close working relationship with device manufacturers and SoC vendors, and we have witnessed first-hand their enthusiasm for adopting AV1. To ensure optimal streaming performance, Netflix has a rigorous certification process to verify proper support for our streaming formats on devices. AV1 was added to this certification process in 2019, and since then, we have seen a steady increase in the number of devices with full AV1 decoding capabilities. Over the past five years (2021–2025), 88% of large-screen devices, including TVs, set-top boxes, and streaming sticks, submitted for Netflix certification have supported AV1, with the vast majority offering full 4K@60fps capability. Notably, since 2023, almost all devices we have received for certification are AV1-capable.</p><p>We have also been impressed by the robustness of AV1 implementations across these devices. As mentioned earlier, FGS is an innovative tool that departs from traditional codec architectures and was not included in our initial full-scale AV1 streaming rollout. When we launched FGS this July, we worked closely with our partners to ensure broad device compatibility. We are pleased with the successful progress made, and AV1 with FGS is now supported across a significant and growing number of in-field devices.</p><h3>Looking Ahead: AV1 Today, AV2 Tomorrow</h3><p>As we reflect on our AV1 journey, it’s clear that the codec has already transformed the streaming experience for hundreds of millions of Netflix members worldwide. Thanks to industry-wide collaboration and rapid device adoption, AV1 is delivering higher quality, greater efficiency, and new cinematic features to more screens than ever before.</p><p>Looking ahead, we are excited about the forthcoming release of AV2, announced by the Alliance for Open Media for the end of 2025. <a href="https://www.youtube.com/watch?v=RUMwMe_2Dqo">AV2 is poised to set a new benchmark for compression efficiency and streaming capabilities, building on the solid foundation laid by AV1</a>. At Netflix, we remain committed to adopting the best open technologies to delight our members around the globe. While AV2 represents the future of streaming, AV1 is very much the present — serving as the backbone of our platform and powering exceptional entertainment experiences across a vast and ever-expanding ecosystem of devices.</p><h3>Acknowledgement</h3><p>The success of AV1 at Netflix is the result of the dedication, expertise, and collaboration of many teams across the company — including Encoding, Clients, Device Certification, Partner Engineering, Data Science &amp; Engineering, Infra, Platform, etc.</p><p>We would also like to thank <a href="https://www.linkedin.com/in/artemdanylenko/">Artem Danylenko</a>, <a href="https://www.linkedin.com/in/aditya-mavlankar-7139791/">Aditya Mavlankar</a>, <a href="https://www.linkedin.com/in/anne-aaron/">Anne Aaron</a>, <a href="https://www.linkedin.com/in/cyril-concolato-567a522/">Cyril Concolato</a>, <a href="https://www.linkedin.com/in/allanzp/">Allan Zhou</a> and <a href="https://www.linkedin.com/in/anush-moorthy-b8451142/">Anush Moorthy</a> for their valuable comments and feedback on earlier drafts of this post.</p><h3>Footnotes</h3><ol><li>These numbers represent a snapshot of data from November 13, 2025. Actual values may vary slightly from day to day and across different regions, depending on the mix of content, devices, and internet connectivity.</li></ol><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*xuLP8glDDcj-DBYO8djmNA.png"></figure><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=02f592242d80" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/av1-now-powering-30-of-netflix-streaming-02f592242d80">AV1 — Now Powering 30% of Netflix Streaming</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/av1-now-powering-30-of-netflix-streaming-02f592242d80</link>
      <guid>https://netflixtechblog.com/av1-now-powering-30-of-netflix-streaming-02f592242d80</guid>
      <pubDate>Thu, 04 Dec 2025 21:09:00 +0100</pubDate>
    </item>
    <item>
      <title><![CDATA[Supercharging the ML and AI Development Experience at Netflix]]></title>
      <description><![CDATA[<h3>Supercharging the ML and AI Development Experience at Netflix with Metaflow</h3><p><a href="https://www.linkedin.com/in/shashanksrikanth/"><em>Shashank Srikanth</em></a>, <a href="https://www.linkedin.com/in/romain-cledat-4a211a5/"><em>Romain Cledat</em></a></p><p><a href="https://docs.metaflow.org/">Metaflow</a> — a framework we started and <a href="https://netflixtechblog.com/open-sourcing-metaflow-a-human-centric-framework-for-data-science-fa72e04a5d9">open-sourced</a> in 2019 — now powers <a href="https://netflixtechblog.com/supporting-diverse-ml-systems-at-netflix-2d2e6b6d205d">a wide range of ML and AI systems across Netflix</a> and at <a href="https://github.com/Netflix/metaflow/blob/master/ADOPTERS.md">many other companies</a>. It is well loved by users for helping them take their ML/AI workflows from <a href="https://docs.metaflow.org/introduction/what-is-metaflow#how-does-metaflow-support-prototyping-and-production-use-cases">prototype to production</a>, allowing them to focus on building cutting-edge systems that bring joy and entertainment to audiences worldwide.</p><p>Metaflow allows users to:</p><ol><li><strong>Iterate and ship quickly </strong>by minimizing friction</li><li><strong>Operate systems reliably</strong> in production with minimal overhead, at Netflix scale.</li></ol><p>Metaflow works with many battle-hardened tooling to address the second point — among them <a href="https://netflixtechblog.com/100x-faster-how-we-supercharged-netflix-maestros-workflow-engine-028e9637f041">Maestro</a>, our newly open-sourced workflow orchestrator that powers nearly every ML and AI system at Netflix and serves as a backbone for Metaflow itself.</p><p>In this post, we focus on the first point and introduce a new Metaflow functionality, <strong>Spin</strong>, that helps users <strong>accelerate their iterative development process</strong>. By the end, you’ll have a solid understanding of Spin’s capabilities and learn how to try it out yourself with <strong>Metaflow 2.19</strong>.</p><h3>Iterative development in ML and AI workflows</h3><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*0I8DAvXCQEN1RpTTC0ZyaQ.jpeg"><figcaption>Developing a Metaflow flow with cards in VSCode</figcaption></figure><p>To understand our approach to improving the ML and AI development experience, it helps to consider how these workflows differ from traditional software engineering.</p><p>ML and AI development revolves not just around code but also around data and models, which are large, mutable, and computationally expensive to process. Iteration cycles can involve long-running data transformations, model training, and stochastic processes that yield slightly different results from run to run. These characteristics make fast, stateful iteration a critical part of productive development.</p><p>This is where notebooks — such as Jupyter, <a href="https://observablehq.com/documentation/notebooks/">Observable</a>, or <a href="https://marimo.io/">Marimo</a> — shine. Their ability to preserve state in memory allows developers to load a dataset once and iteratively explore, transform, and visualize it without reloading or recomputing from scratch. This persistent, interactive environment turns what would otherwise be a slow, rigid loop into a fluid, exploratory workflow — perfectly suited to the needs of ML and AI practitioners.</p><p>Because ML and AI development is computationally intensive, stochastic, and data- and model-centric, tools that optimize iteration speed must treat state management as a first-class design concern. Any system aiming to improve the development experience in this domain must therefore enable quick, incremental experimentation without losing continuity between iterations.</p><h3>New: rapid, iterative development with spin</h3><p>At first glance, Metaflow code looks like a workflow — similar to <a href="https://airflow.apache.org/">Airflow</a> — but there’s another way to look at it: each Metaflow @step serves as <a href="https://docs.metaflow.org/metaflow/basics#what-should-be-a-step">a checkpoint boundary</a>. At the end of every step, Metaflow automatically persists all instance variables as <em>artifacts</em>, allowing the execution to <a href="https://docs.metaflow.org/metaflow/debugging#how-to-use-the-resume-command">resume</a> seamlessly from that point onward. The below animation shows this behavior in action:</p><figure><img alt="An animated GIF showing how resume can be used in Metaflow. The GIF shows how using `flow.py resume join` makes Metaflow clone previously executed steps and resumes the computation from the `join` step and continues executing till the end of the flow." src="https://cdn-images-1.medium.com/max/1024/1*AEDpnt-YULYV4mcyrwLk7g.gif"><figcaption>Using resume in Metaflow</figcaption></figure><p>In a sense, we can consider a @step similar to a notebook cell: it is the smallest unit of execution that updates state upon completion. It does have a few differences that address the issues with notebook cells:</p><ul><li><strong>The execution order is explicit and deterministic: </strong>no surprises due to out-of-order cell execution;</li><li><strong>The state is not hidden: </strong>state is explicitly stored as self. variables as shared state, which can be <a href="https://docs.metaflow.org/metaflow/client">discovered and inspected</a>;</li><li><strong>The state is versioned and persisted</strong> making results more reproducible.</li></ul><p>While <strong>Metaflow</strong>’s resume feature can approximate the incremental and iterative development approach of notebooks, it restarts execution from the selected step onward, introducing more latency between iterations. In contrast, a <strong>notebook</strong> allows near-instant feedback by letting users tweak and rerun individual cells while seamlessly reusing data from earlier cells held in memory.</p><p>The new spin command in Metaflow 2.19 addresses this gap. Similar to executing a single notebook cell, it quickly executes a single Metaflow @step — with all the state carried over from the parent step. As a result, users can develop and debug Metaflow steps as easily as a cell in a notebook.</p><p>The effect becomes clear when considering the three complementary execution modes — run, resume, and spin — side by side, mapping them to the corresponding notebook behavior:</p><figure><img alt="Diagram showing the various modes of execution in Metaflow: Run, Resume and Spin" src="https://cdn-images-1.medium.com/max/1024/1*DgRIxOu-7keiFoHia9JMrg.png"><figcaption>Run, Resume and Spin “modes”</figcaption></figure><p>Another major difference isn’t just what gets executed, but what gets recorded. Both run and resume create a full, versioned run with complete metadata and artifacts, while spin skips tracking altogether. It’s built for fast, throw-away iterations during development.</p><p>The one-minute clip below illustrates a typical iterative development workflow that alternates between run and spin. In this example, we are building a flow that reads a dataset from a Parquet file and trains a separate model for each product category, focusing on computer-related categories.</p><a href="https://medium.com/media/36b60cb7c79bf5c77ddd1f60cc97ae8b/href">https://medium.com/media/36b60cb7c79bf5c77ddd1f60cc97ae8b/href</a><p>As shown in the video, we start by creating a flow from scratch and running a minimal version of it to persist test artifacts — in this case, a Parquet dataset. From there, we can use spin to iterate on one step at a time, incrementally building out the flow, for example, by adding the parallel training steps demonstrated in the clip.</p><p>Once the flow has been iterated on locally, it can be seamlessly deployed to production orchestrators like Maestro or <a href="https://docs.metaflow.org/production/scheduling-metaflow-flows/scheduling-with-argo-workflows">Argo</a>, and <a href="https://docs.metaflow.org/scaling/remote-tasks/requesting-resources">scaled up</a> on compute platforms such as AWS Batch, Titus, Kubernetes and more. Thus, the experience is as smooth as developing in a notebook, but the outcome is a production-ready, scalable workflow, implemented as an idiomatic Python project!</p><h3>Spin up smooth development in VSCode/Cursor</h3><p>Instead of typing run and spin manually in the terminal, we can bind them to keyboard shortcuts. For example, <a href="https://github.com/outerbounds/metaflow-dev-vscode">the simple metaflow-dev VS Code extension</a> (works with Cursor as well) maps Ctrl+Opt+R to run and Ctrl+Opt+S to spin. Just hack away, hit Ctrl+Opt+S, and the extension will save your file and spin the step you are currently editing.</p><p>One area where spin truly shines is in creating mini-dashboards and reports with <a href="https://docs.metaflow.org/metaflow/visualizing-results">Metaflow Cards</a>. Visualization is another strong point of notebooks but the combination of spin and cards makes Metaflow a very compelling alternative for developing real-time and post-execution visualizations. Developing cards is inherently iterative and visual (much like building web pages) where you want to tweak code and see the results instantly. This workflow is readily available with the combination of VSCode/Cursor, which includes a built-in web-view, <a href="https://docs.metaflow.org/metaflow/visualizing-results/effortless-task-inspection-with-default-cards#using-local-card-viewer">the local card viewer</a>, and spin.</p><p>To see the trio of tools — along with the VS Code extension — in action, in this short clip we add observability to the train step that we built in the earlier example:</p><a href="https://medium.com/media/e2596ec88a7e65b4bad1b6ba7a2a167c/href">https://medium.com/media/e2596ec88a7e65b4bad1b6ba7a2a167c/href</a><p>A major benefit of Metaflow Cards is that we don’t need to deploy any extra services, data streams, and databases for observability. Just develop visual outputs as above, deploy the flow, and wehave a complete system in production with reporting and visualizations included.</p><h3>Spin to the next level: injecting inputs, inspecting outputs</h3><p>Spin does more than just run code — it also lets us take full control of a spun @step’s inputs and outputs, enabling a range of advanced patterns.</p><p>In contrast to notebooks, we can spin any arbitrary @step in a flow using state from any past run, making it easy to test functions with different inputs. For example, if we have multiple models produced by separate runs, we could spin an inference step, supplying a different model run each time.</p><p>We can also override artifact values or inject arbitrary Python objects — similar to a notebook cell — for spin. Simply specify a Python module with an ARTIFACTS dictionary:</p><pre>ARTIFACTS = {<br>  "model": "kmeans",<br>  "k": 15<br>}</pre><p>and point spin at the module:</p><pre>spin train --artifacts-module artifacts.py</pre><p>By default spin doesn’t persist artifacts, but we can easily change this by adding --persist. Even in this case, artifacts are not persisted in the usual Metaflow datastore but to a directory-specific location which you can easily clean up after testing. We can access the results with <a href="https://docs.metaflow.org/metaflow/client">the Client API</a> as usual — just specify the directory you want to inspect with inspect_spin:</p><pre>from metaflow import inspect_spin<br><br>inspect_spin(".")<br>Flow("TrainingFlow").latest_run["train"].task["model"].data</pre><p>Being able to inspect and modify a step’s inputs and outputs on the fly unlocks a powerful use case:<strong> unit testing individual steps</strong>. We can use spin programmatically through <a href="https://docs.metaflow.org/metaflow/managing-flows/runner">the Runner API</a> and assert the results:</p><pre>from metaflow import Runner<br><br>with Runner("flow.py").spin("train", persist=True) as spin:<br>  assert spin.task["model"].data == "kmeans"</pre><h3>Making AI agents spin</h3><p>In addition to speeding up development for humans, spin turns out to be surprisingly handy for coding agents too. There are two major advantages to teaching AI how to spin:</p><ol><li><strong>It accelerates the development loop</strong>. Agents don’t naturally understand what’s slow, or why speed matters, so they need to be nudged to favor faster tools over slower ones.</li><li><strong>It helps surface errors faster </strong>and contextualizes them to a specific piece of code, increasing the chance that the agent is able to fix errors by itself.</li></ol><p>Metaflow users are already <a href="https://claude.com/product/claude-code">using</a> Claude Code; spin makes this even easier. In the example below, we added the following section in a CLAUDE.md file:</p><pre>## Developing Metaflow code<br>Follow this incremental development workflow that ensures quick iterations<br>and correct results. You must create a flow incrementally, step by step<br>following this process:<br>1. Create a flow skeleton with empty `@step`s.<br>2. Add a data loading step.<br>3. `run` the flow.<br>4. Populate the next step and use `spin` to test it with the correct inputs.<br>5. `run` the flow to record outputs from the new step.<br>5. Iterate on (4–5) until all steps have been implemented and work correctly.<br>6. `run` the whole flow to ensure final correctness.<br><br>To test a flow, run the flow as follows<br>```<br>python flow.py - environment=pypi run<br>```<br><br>Do this once before running `spin`.<br>As you are building the flow, you `spin` to test steps quickly.<br>For instance<br>```<br>python flow.py - environment=pypi spin train<br>```</pre><p>Just based on these quick instructions, the agent is able to use spin effectively. Take a look at the following inspirational example that one-shots Claude to create a flow, along the lines of our earlier examples, which trains a classifier to predict product categories:</p><a href="https://medium.com/media/706e83950e73a05ca41b4fb463702690/href">https://medium.com/media/706e83950e73a05ca41b4fb463702690/href</a><p>In the video, we can see Claude using spin around the 45-second mark to test a preprocess step. The step initially fails due to a classic data science pitfall: during testing, Claude samples only a small subset of data, causing some classes to be underrepresented. The first spin surfaces the issue, which Claude then fixes by switching to stratified sampling — and finally does another spin to confirm the fix, before proceeding to complete the task.</p><h3>The inner loop of end-to-end ML/AI</h3><p>To circle back to where we started, our motivation for adding spin — and for creating Metaflow in the first place — is to accelerate development cycles so we can deliver more joy to our subscribers, faster. Ultimately, we believe there’s no single magic feature that makes this possible. It takes all parts of an ML/AI platform working together coherently — spin included.</p><p>From this perspective, it’s useful to place spin in the context of other Metaflow features. It’s designed for the innermost loop of model and business-logic development, with the added benefit of supporting unit testing during deployment, as shown in the overall blueprint of the Metaflow toolchain below.</p><figure><img alt="Metaflow tool-chain." src="https://cdn-images-1.medium.com/max/1024/1*9cd6SHFrW7A4iWMZHYkVkw.png"><figcaption>Metaflow tool-chain</figcaption></figure><p>In this diagram, the solid blue boxes represent different Metaflow commands, while the blue text denotes decorators and other features. In particular, note the <em>Shared Functionality</em> box — another key focus area for us over the past year — which includes <a href="https://netflixtechblog.com/introducing-configurable-metaflow-d2fb8e9ba1c6">configuration management</a> and <a href="https://docs.metaflow.org/metaflow/composing-flows/introduction">custom decorators</a>. These capabilities let domain-specific teams and platform providers tailor Metaflow to their own use cases. Following our ethos of composability, all of these features integrate seamlessly with spin as well.</p><p>Another key design philosophy of Metaflow is to let projects start small and simple, adding complexity only when it becomes necessary. So don’t be overwhelmed by the diagram above. To get started, install Metaflow easily with</p><pre>pip install metaflow</pre><p>and take your first baby @steps for a spin! Check out the <a href="https://docs.metaflow.org/metaflow/authoring-flows/introduction">docs</a> and for questions, support, and feedback, join the friendly <a href="http://chat.metaflow.org/">Metaflow Community Slack</a>.</p><h3>Acknowledgments</h3><p>We would like to thank our partners at <a href="https://outerbounds.com/">Outerbounds</a>, and particularly <a href="https://www.linkedin.com/in/villetuulos/">Ville Tuulos</a>, <a href="https://www.linkedin.com/in/savingoyal/">Savin Goyal</a>, and <a href="https://www.linkedin.com/in/madhur-tandon/">Madhur Tandon</a>, for their collaboration on this feature, from initial ideation to review, testing and documentation. We would also like to acknowledge the rest of the Model Development and Management team (<a href="https://www.linkedin.com/in/maria-alder/">Maria Alder</a>, <a href="https://www.linkedin.com/in/david-j-berg/">David J. Berg</a>, <a href="https://www.linkedin.com/in/shaojingli/">Shaojing Li</a>, <a href="https://www.linkedin.com/in/rui-lin-483a83111/">Rui Lin</a>, <a href="https://www.linkedin.com/in/nissanpow/">Nissan Pow</a>, <a href="https://www.linkedin.com/in/chaoying-wang/">Chaoying Wang</a>, <a href="https://www.linkedin.com/in/reginalw/">Regina Wang</a>, <a href="https://www.linkedin.com/in/shuishiyang/">Seth Yang</a>, <a href="https://www.linkedin.com/in/zitingyu/">Darin Yu</a>) for their input and comments.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=b2d5b95c63eb" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/supercharging-the-ml-and-ai-development-experience-at-netflix-b2d5b95c63eb">Supercharging the ML and AI Development Experience at Netflix</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/supercharging-the-ml-and-ai-development-experience-at-netflix-b2d5b95c63eb</link>
      <guid>https://netflixtechblog.com/supercharging-the-ml-and-ai-development-experience-at-netflix-b2d5b95c63eb</guid>
      <pubDate>Tue, 04 Nov 2025 21:33:00 +0100</pubDate>
    </item>
    <item>
      <title><![CDATA[Post-Training Generative Recommenders with Advantage-Weighted Supervised Finetuning]]></title>
      <description><![CDATA[<p>Author: <a href="https://keertanavc.github.io/">Keertana Chidambaram</a>, <a href="https://www.linkedin.com/in/qiuling-xu-a445b815a">Qiuling Xu</a>, <a href="https://www.linkedin.com/in/markhsiao/">Ko-Jen Hsiao</a>, <a href="https://www.linkedin.com/in/moumitab/">Moumita Bhattacharya</a></p><p>(*The work was done when Keertana interned at Netflix.)</p><h3>Introduction</h3><p>This blog focuses on post-training generative recommender systems. Generative recommenders (GRs) represent a new paradigm in the field of recommendation systems (e.g. <a href="https://github.com/meta-recsys/generative-recommenders">HSTU</a>, <a href="https://arxiv.org/abs/2502.18965">OneRec</a>). These models draw inspiration from recent advancements in transformer architectures used for language and vision tasks. They approach the recommendation problem, including both ranking and retrieval, as a sequential transduction task. This perspective enables generative training, where the model learns by imitating the next event in a sequence of user activities, thereby effectively modeling user behavior over time.</p><p>However, a key challenge with simply replicating observed user patterns is that it may not always lead to the best possible recommendations. User interactions are influenced by a variety of factors — such as trends, or external suggestions — and the system’s view of these interactions is inherently limited. For example, if a user tries a popular show but later indicates it wasn’t a good fit, a model that only imitates this behavior might continue to recommend similar content, missing the chance to enhance the user’s experience.</p><p>This highlights the importance of incorporating user preferences and feedback, rather than solely relying on observed behavior, to improve recommendation quality. In the context of recommendation systems, we benefit from a wealth of user feedback, which includes explicit signals such as ratings and reviews, as well as implicit signals like watch time, click-through rates, and overall engagement. This abundance of feedback serves as a valuable resource for improving model performance.</p><p>Given the recent success of reinforcement learning techniques in post-training large language models, such as DPO and GRPO, this study investigates whether similar methods can be applied to generative recommenders. Ultimately, our goal is to identify both the opportunities and challenges in using these techniques to enhance the quality and relevance of recommendations.</p><p>Unlike language models, post-training generative recommenders presents unique challenges. One of the most significant is the difficulty of obtaining counterfactual feedback in recommendation scenarios. The recommendation feedback is generated on-policy — that is, it reflects users’ real-time interactions with the system as they naturally use it. Since a typical user sequence can span weeks or even years of activity, it is impractical to ask users to review or provide feedback on hypothetical, counterfactual experiences. As a result, the absence of counterfactual data makes it challenging to apply post-training methods such as PPO or DPO, which require feedback from counterfactual user sequences.</p><p>Furthermore, post-training methods typically rely on a reward model — either implicit or explicit — to guide optimization. The quality of reward models heavily influences the effectiveness of post-training. In the context of recommendation systems, however, reward signals tend to be much noisier. For instance, if we use watch time as an implicit reward, it may not always accurately reflect user satisfaction: a viewer might stop watching a favorite show simply due to time constraints, while finishing a lengthy show doesn’t necessarily indicate genuine enjoyment.</p><p>To address these post-training challenges, we introduce a novel algorithm called Advantage-Weighted Supervised Fine-tuning (A-SFT). Our analysis first demonstrates that reward models in recommendation systems often exhibit higher uncertainty due to the issues discussed above. Rather than relying solely on these uncertain reward models, A-SFT combines supervised fine-tuning with the advantage function to more effectively guide post-training optimization. This approach proves especially effective when the reward model has high variance but still provides valuable directional signals. We benchmark A-SFT against four other representative methods, and our results show that A-SFT achieves better alignment between the pre-trained generative recommendation model and the reward model.</p><p>In Figure 1, we conceptualize the pros and cons of different post-training paradigms. For example, Online Reinforcement Learning is most useful when the reward model has a good generalization ability, and behavior cloning is suitable when no reward models are available. Using these algorithms under fitting use cases is the key to a successful post-training. For example, over-exploitation of noisy reward models will hurt task performance, as guidance from the reward models can be simply noise. Conversely, not leveraging a good reward model leaves out potential improvements. We find A-SFT fits the sweet point between offline reinforcement learning and behavior cloning, where it benefits from the directional signals in those noisy estimations and is less dependent on the reward accuracy.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*N9QNNLspEJJCxQjtNl9HZw.png"></figure><p>Figure 1: The landscape of RL algorithms based on the reward models’ accuracy</p><h3>Challenges in Post-training for Recommendation</h3><p>Reinforcement Learning from Human Feedback (RLHF) is the most popular framework for post-training large language models. In this framework, human annotators evaluate and rank different outputs generated by a model. This feedback is then used to train a reward model that predicts how well a model output aligns with human preferences. This reward model then serves as a proxy for human judgment during reinforcement learning, guiding the model to generate outputs that are more likely to be preferred by humans.</p><p>While traditional RLHF methods like PPO or DPO are effective for aligning LLMs, there are several challenges in applying them directly to large-scale recommendation systems:</p><ol><li>Lack of Counter-factual Observations</li></ol><p>As in typical RLHF settings, collecting real-time feedback from a diverse user base across a wide range of items is both costly and impractical. The data in recommendation are generated by the real-time user interests. Any third-party annotators or even the user themselves lack the practical means to evaluate an alternative reality. For example, it is impractical to ask the Netflix users to evaluate hundreds of unseen movies. Consequently, we lack a live environment in which to perform reinforcement learning.</p><p>2. Noisy Reward Models</p><p>In addition to the limited counter-factual data, the recommendation task itself has a higher randomness by its nature. The recommendation data has less structure than language data. Users choose to watch some shows not because there is a grammar rule that nouns need to follow by the verbs. In fact, the users’ choices usually exhibit a level of permutation invariance, where swapping the order of events in the user history still makes a valid activity sequence. This randomness in the behaviors makes learning a good reward model extremely difficult. Often the reward models we learnt still have a large margin of errors.</p><p>Here is an ablation study we did on the reward model performance with O(Millions) users and O(Billions) of tokens. The reward model uses an open-sourced HSTU architecture in the convenience of reproducing this study. We adopt the standard RLHF approach of training a reward model using offline, human-collected feedback. We start by creating a proxy reward, scored on a scale from 1 to 5 in the convenience of understanding. This reward model is co-trained as a shallow reward head on top of the generative recommender. It predicts the reward for the most recently selected title based on a user’s interaction history. To evaluate its effectiveness, we compare the model’s performance against two simple baselines: (1) predicting the next reward as the average reward the user has given in their past interactions, and (2) predicting it as the average reward that all users have assigned to that particular title in previous interactions.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*tmYdEXqVjSkne__-k6gSig.png"></figure><p>Table 1: Reward model performance metrics</p><p>We observe that the model’s predictions do not significantly outperform the simple baselines. This result is intuitive, as a user’s historical interactions typically cover only a small subset of titles, making it difficult to accurately predict their responses to the vast number of unexplored titles in the catalogue. We expect this to be a potential issue for any large recommendation systems where the ratio between explored and unexplored titles is very small.</p><p>3. Lack of Logged Policy</p><p>In recommendation systems, the policy that generated the logged data is typically unknown and cannot be directly estimated. Offline reinforcement learning methods often rely on Inverse Propensity Scoring (IPS) to debias such data by reweighting interactions according to the logging policy’s action probabilities. However, estimating the logging policy accurately is challenging and prone to error, which can introduce additional biases, and IPS itself is known to suffer from high variance. Consequently, offline RL approaches that depend on IPS are ill-suited for our setting.</p><h3>Advantage Weighted Supervised Fine Tuning</h3><p>Given the three challenges outlined above, we propose a new algorithm Advantage-Weighted SFT (A-SFT). It leverages a combination of supervised fine-tuning and advantage reweighting from reinforcement learning. The key observation is as follows. Despite the reward estimation for each individual event having a high uncertainty, we find the estimations of rewards contain directional signals between high-reward and low-reward events. These signals could help better align the model during post-training.</p><p>A central factor in this study is the generalization ability of the reward model. Better generalization enables more accurate predictions of user preferences for unseen titles, thereby making exploration more effective. For reward models with moderate to high generalization power, both online RL methods such as PPO and offline RL methods such as CQL can perform effectively. However, in our setting, reward model generalization is worse than the language counterparts’, which makes these algorithms less appropriate. In addition, the use of techniques like inverse propensity scoring (IPS) introduces a heightened risk of high-variance estimates, prompting us to exclude algorithms such as off-policy REINFORCE.</p><p>Our proposed method A-SFT does not rely on IPS. With no need of prior knowledge of the logging policy, it can be generally applied to cases where observation of the environments are limited or biased. This is particularly useful to the recommendation setting due to the user feedback loop and distribution shifts with time. Without knowing the logging policy, A-SFT still provides means to control the policy deviation between the current policy and logging policy by tuning the parameter. This design provides essential means to control the learnt bias from uncertain reward models. We show that A-SFT outperforms baseline behavior cloning by directly optimizing observed rewards.</p><p>The advantage-weighted SFT algorithm is as follows:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*eP-_FyLRs6vrGnwyIp_j_A.png"></figure><p>For the results presented in this blog post, we treat the recommendation problem as a contextual bandit, i.e. given a history of user interactions as the context, can we recommend a high reward next title recommendation for the user?</p><h3>Benchmarks</h3><p>We compared representative algorithms including PPO, IPO, DPO, CQL and SFT as the baselines:</p><ol><li><strong>Reward weighted Behavior Cloning</strong>: This benchmark algorithm modifies supervised fine-tuning (SFT) by weighting the loss with the raw rewards of the chosen item instead of weighing the loss with advantage as in the proposed algorithm.</li><li><strong>Rejection Sampling Direct Preference Optimization / Identity Preference Optimization (RS DPO/IPO)</strong>: this is a variant of DPO/IPO where, for each user history x, ​we generate contrasting response pairs by training an ensemble of reward models to estimate confidence intervals for the reward of multiple potential responses y. If the lower bound of the reward confidence interval for one response​ is less than the upper bound for another response, then this pair is used to train DPO/IPO.</li><li><strong>Conservative Q-Learning (CQL)</strong>: This is a standard offline algorithm that learns a conservative Q function, penalizing overestimation of Q-values, particularly in regions of the state-action space with little or no reward data.</li><li><strong>Proximal Policy Optimization (PPO)</strong>: This is a standard RLHF (Reinforcement Learning from Human Feedback) algorithm that uses reward models as an online environment. PPO learns an advantage function and optimizes the policy to maximize expected reward while maintaining proximity to the initial policy.</li></ol><p>We sampled a separate test set of O(Millions) users. This test set is collected on a future date after the training.</p><h3>Offline Evaluation Results</h3><p>We evaluate our algorithm on a dataset of high-reward user trajectories. For sake of simplicity, we consider a trajectory to have a high reward if the accumulated reward is higher than the median of the population. We present the following metrics for the held out test dataset:</p><ol><li><strong>NDCG@k</strong>: This measures the ranking quality of the recommended items up to position k. It accounts for the position of relevant items in the recommendation list, assigning higher scores when relevant items appear higher in the ranking. The gain is discounted logarithmically at lower ranks, and the result is normalized by the ideal ranking (i.e., the best possible ordering of items).</li><li><strong>HR@k</strong>: This measures the proportion of test cases in which the ground-truth chosen item y appears in the top k recommendations. It is a binary metric per test case (hit or miss) and is averaged over all test cases.</li><li><strong>MRR</strong>: MRR evaluates the ranking quality by measuring the reciprocal of the rank at which the chosen item appears in the recommendation list. The metric is averaged across all test cases.</li><li><strong>Reward Model as A Judge</strong>: We use the reward model to evaluate the policy for future user events. We propose to use an ensemble of reward models for the evaluation to increase confidence. The result is based on the discounted reward generated for a few steps. The standard deviation is less than 4%.</li></ol><p>We measure the percentage improvement in each metric compared to the baseline, Reward Weighted Behavior Cloning(BC). We notice that advantage weighted SFT shows the largest improvement in metrics, outweighing BC as well as reward model dependent algorithms like CQL, PPO, DPO and IPO.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*Z8wcCOETlobx8T_oVuDRiA.png"></figure><p>Our experiments show that advantage weighted SFT is a simple but promising approach for post-training generative recommenders as it deals with the issue of poor reward model generalizations and lack of IPS. More specifically, we find PPO, IPO and DPO achieve a good reward score, but also causes the overfitting from the reward model. Conservative Q-Learning achieves more robust improvements but does not fully capture the potential signals in the reward modeling. A-SFT achieved both better recommendation metrics and reward scores.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=61a538d717a9" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/post-training-generative-recommenders-with-advantage-weighted-supervised-finetuning-61a538d717a9">Post-Training Generative Recommenders with Advantage-Weighted Supervised Finetuning</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/post-training-generative-recommenders-with-advantage-weighted-supervised-finetuning-61a538d717a9</link>
      <guid>https://netflixtechblog.com/post-training-generative-recommenders-with-advantage-weighted-supervised-finetuning-61a538d717a9</guid>
      <pubDate>Sun, 26 Oct 2025 00:01:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Behind the Streams: Real-Time Recommendations for Live Events Part 3]]></title>
      <description><![CDATA[<p>By: <a href="https://www.linkedin.com/in/krisrange/">Kris Range</a>, <a href="https://www.linkedin.com/in/gulatiankush/">Ankush Gulati</a>, <a href="https://www.linkedin.com/in/jimpisaacs/">Jim Isaacs</a>, <a href="https://www.linkedin.com/in/jennifer-s-0019a516/">Jennifer Shin</a>, <a href="https://www.linkedin.com/in/jeremy-kelly-526a30180/">Jeremy Kelly</a>, <a href="https://www.linkedin.com/in/jason-t-26850b26/">Jason Tu</a></p><p><em>This is part 3 in a series called “Behind the Streams”. Check out </em><a href="https://netflixtechblog.com/behind-the-streams-live-at-netflix-part-1-d23f917c2f40"><em>part 1</em></a><em> and </em><a href="https://netflixtechblog.com/building-a-reliable-cloud-live-streaming-pipeline-for-netflix-8627c608c967"><em>part 2</em></a><em> to learn more.</em></p><p>Picture this: It’s seconds before the biggest fight night in Netflix history. Sixty-five million fans are waiting, devices in hand, hearts pounding. The countdown hits zero. What does it take to get everyone to the action on time, every time? At Netflix, we’re used to on-demand viewing where everyone chooses their own moment. But with live events, millions are eager to join in at once. Our job: make sure our members never miss a beat.</p><p>When Live events break streaming records <a href="https://about.netflix.com/en/news/60-million-households-tuned-in-live-for-jake-paul-vs-mike-tyson">¹</a> <a href="https://about.netflix.com/en/news/netflix-nfl-christmas-gameday-reaches-65-million-us-viewers">²</a> <a href="https://about.netflix.com/en/news/over-41-million-global-viewers-on-netflix-watch-terence-crawford-defeat">³</a>, our infrastructure faces the ultimate stress test. Here’s how we engineered a discovery experience for a global audience excited to see a knockout.</p><h3>Why are Live Events Different?</h3><p>Unlike Video on Demand (VOD), members want to catch live events as they happen. There’s something uniquely exciting about being part of the moment. That means we only have a brief window to recommend a Live event at just the right time. Too early, excitement fades; too late, the moment is missed. Every second counts.</p><p>To capture that excitement, we enhanced our recommendation delivery systems to serve real-time suggestions, providing members richer and more compelling signals to hit play in the moment when it matters most. The challenge? Sending dynamic, timely updates concurrently to over a hundred million devices worldwide without creating a <a href="https://en.wikipedia.org/wiki/Thundering_herd_problem">thundering herd effect</a> that would overwhelm our cloud services. Simply scaling up linearly isn’t efficient and reliable. For popular events, it could also divert resources from other critical services. We needed a smarter and more scalable solution than just adding more resources.</p><h3>Orchestrating the moment: Real-time Recommendations</h3><p>With millions of devices online and live event schedules that can shift in real time, the challenge was to keep everyone perfectly in sync. We set out to solve this by building a system that doesn’t just react, but adapts by dynamically updating recommendations as the event unfolds. We identified the need to balance three constraints:</p><ul><li><strong>Time</strong>: the duration required to coordinate an update.</li><li><strong>Request throughput</strong><em>: </em>the capacity of our cloud services to handle requests.</li><li><strong>Compute cardinality</strong>: the variety of requests necessary to serve a unique update.</li></ul><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*-z6A8FriBAbJW5BcwMrZTA.png"><figcaption>Visualizing constraints for real-time updates</figcaption></figure><p>We solved this constraint optimization problem by splitting the real-time recommendations into two phases: <strong>prefetching</strong> and <strong>real-time broadcasting</strong>. First, we prefetch the necessary data ahead of time, distributing the load over a longer period to avoid traffic spikes. When the Live event starts or ends, we broadcast a low cardinality message to all connected devices, prompting them to use the prefetched data locally. The timing of the broadcast also adapts when event times shift to preserve accuracy with the production of the Live event. By combining these two phases, we’re able to keep our members’ devices in sync and solve the thundering herd problem. To maximize device reach, especially for those with unstable networks, we use “at least once” broadcasts to ensure every device gets the latest updates and can catch up on any previously missed broadcasts as soon as they’re back online.</p><p>The first phase optimizes <strong>request throughput </strong>and <strong>compute cardinality</strong> by prefetching materialized recommendations, displayed title metadata, and artwork for a Live event. As members naturally browse their devices before the event, this data is prepopulated and stored locally in device cache, awaiting the notification trigger to serve the recommendations instantaneously. By distributing these requests naturally over time ahead of the event, we can eliminate any related traffic spikes and avoid the need for large-scale, real-time system scaling.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*QmB99SosLqs-JEo0gp1wwg.png"><figcaption>A phased approach, smoothing traffic requests over time with a real-time low-cardinality broadcast</figcaption></figure><p>The second phase optimizes<strong> request throughput </strong>and<strong> time </strong>to update<strong> </strong>devices by broadcasting a low-cardinality, real-time message to all connected devices at critical moments in a Live event’s lifecycle. Each broadcast payload includes a <strong>state key</strong> and a <strong>timestamp</strong>. The state key indicates the current stage of the Live event, allowing devices to use their pre-fetched data to update cached responses locally without additional server requests. The timestamp ensures that if a device misses a broadcast due to network issues, it can catch up by replaying missed updates upon reconnecting. This mechanism guarantees devices receive updates at least once, significantly increasing delivery reliability even on unstable networks.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*h6CIrfYnpR24NS5hgvxEGA.png"><figcaption>A phased approach optimizes each constraint to ensure we can deliver for the big moment!</figcaption></figure><blockquote>Moment in Numbers: During peak load, we have successfully delivered updates at multiple stages of our events to over 100 million devices in under a minute.</blockquote><h3>Under the Hood: How It Works</h3><p>With the big picture in mind, let’s examine how these pieces interact in practice.</p><p>In the diagram below, the Message Producer microservice centralizes all of the business logic. It continuously monitors live events for setup and timing changes. When it detects an update, it schedules broadcasts to be sent at precisely the right moment. The Message Producer also standardizes communication by providing a concise GraphQL schema for both device queries and broadcast payloads.</p><p>Rather than sending broadcasts directly to devices via WebSocket, the Message Producer hands them off to the Message Router. The Message Router is part of a robust two-tier pub/sub architecture built on proven technologies like <a href="https://netflixtechblog.com/pushy-to-the-limit-evolving-netflixs-websocket-proxy-for-the-future-b468bc0ff658">Pushy</a> (our WebSocket proxy), Apache Kafka, and <a href="https://netflixtechblog.com/introducing-netflixs-key-value-data-abstraction-layer-1ea8a0a11b30">Netflix’s KV key-value store</a>. The Message Router tracks subscriptions at the Pushy node granularity, while Pushy nodes map the subscriptions to individual connections, creating a low-latency fanout that minimizes compute and bandwidth requirements.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*Kc1l_Wnc4i08xFA8JnRLCA.png"></figure><p>Devices interface with our GraphQL <a href="https://netflix.github.io/dgs/">Domain Graph Service (DGS)</a>. These schemas offer multiple query interfaces for prefetching, allowing devices to tailor their requests to the specific experience being presented. Each response adheres to a consistent API that resolves to a map of stage keys, enabling fast lookups and keeping business logic off the device. Our broadcast schema specifies WebSocket connection parameters, the current event stage, and the timestamp of the last broadcast message. When a device receives a broadcast, it injects the payload directly into its cache, triggering an immediate update and re-render of the interface.</p><h3>Balancing the Moment: Throughput Management</h3><p>In addition to building the new technology to support real-time recommendations, we also evaluated our existing systems for potential traffic hotspots. Using high-watermark traffic projections for live events, we generated synthetic traffic to simulate game-day scenarios and observed how our online services handled these bursts. Through this process, several common patterns emerged:</p><p><strong>Breaking the Cache Synchrony</strong></p><p>Our game-day simulations revealed that while our approach mitigated the immediate thundering herd risks driven by member traffic during the events, live events introduced unexpected mini thundering herds in our systems hours before and after the actual events. The surge of members joining just in time for these events led to concentrated cache expirations and recomputations, which created traffic spikes well outside the event window that we did not anticipate. This was not a problem for VOD content because the member traffic patterns are a lot smoother. We found that fixed TTLs caused cache expirations and refresh-traffic spikes to happen all at once. To address this, we added jitter to server and client cache expirations to spread out refreshes and smooth out traffic spikes.</p><p><strong>Adaptive Traffic Prioritization</strong></p><p>While our services already leverage traffic prioritization and partitioning based on factors such as request type and device type, live events introduced a distinct challenge. These events generated brief traffic bursts that were intensely spiky and placed significant strain on our systems. Through simulations, we recognized the need for an additional event-driven layer of traffic management.</p><p>To tackle this, we improved our traffic sharding strategies by using event-based signals. This enabled us to route live event traffic to dedicated clusters with more aggressive scaling policies. We also added a dynamic traffic prioritization ruleset that activates whenever we see high requests per second (RPS) to ensure our systems can handle the surge smoothly. During these peaks, we aggressively deprioritize non-critical server-driven updates so that our systems can devote resources to the most time-sensitive computations. This approach ensures smooth performance and reliability when demand is at its highest.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/729/1*PYhFbFxK5PbOEAnYtTfD4Q.jpeg"><figcaption>Snapshot of non-critical traffic volume decline (in %) for a member-facing service during a live event — achieved via aggressive de-prioritization</figcaption></figure><h3>Looking Ahead</h3><p>When we set out to build a seamlessly scalable scheduled viewing experience, our goal was to create a dynamic and richer member experience for live content. Popular live events like the Crawford v. Canelo fight and the NFL Christmas games truly put our systems to the test. Along the way, we also uncovered valuable learnings that continue to shape our work. Our attempts to deprioritize traffic to other non-critical services caused unexpected call patterns and spikes in traffic elsewhere. Similarly, in hindsight, we also learned that the high traffic volume from popular events caused excessive non-essential logging and was putting unnecessary pressure on our ingestion pipelines.</p><p>None of this work would have been possible without our stunning colleagues at Netflix who collaborated across multiple functions to architect, build, and test these approaches, ensuring members can easily access events at the right moment: UI Engineering, Cloud Gateway, Data Science &amp; Engineering, Search and Discovery, Evidence Engineering, Member Experience Foundations, Content Promotion and Distribution, Operations and Reliability, Device Playback, Experience and Design and Product Management.</p><p>As Netflix’s content offering expands to include new formats like live titles, free-to-air linear content, and games, we’re excited to build on what we’ve accomplished and look ahead to even more possibilities. Our roadmap includes extending the capabilities we developed for scheduled live viewing to these emerging formats. We’re also focused on enhancing our engineering tooling for greater visibility into operations, message delivery, and error handling to help us continue to deliver the best possible experience for our members.</p><h3>Join Us for What’s Next</h3><p>We’re just scratching the surface of what’s possible as we bring new live experiences to members around the world. If you are looking to solve interesting technical challenges in a <a href="https://jobs.netflix.com/culture">unique culture</a>, then <a href="https://jobs.netflix.com/">apply</a> for a role that captures your curiosity.</p><p><em>Look out for future blog posts in our “Behind the Streams” series, where we’ll explore the systems that ensure viewers can watch live streams once they manage to find and play them.</em></p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=e027cb313f8f" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/behind-the-streams-real-time-recommendations-for-live-events-e027cb313f8f">Behind the Streams: Real-Time Recommendations for Live Events Part 3</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/behind-the-streams-real-time-recommendations-for-live-events-e027cb313f8f</link>
      <guid>https://netflixtechblog.com/behind-the-streams-real-time-recommendations-for-live-events-e027cb313f8f</guid>
      <pubDate>Tue, 21 Oct 2025 02:53:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[How and Why Netflix Built a Real-Time Distributed Graph: Part 1 — Ingesting and Processing Data…]]></title>
      <description><![CDATA[<h3>How and Why Netflix Built a Real-Time Distributed Graph: Part 1 — Ingesting and Processing Data Streams at Internet Scale</h3><p>Authors: <a href="https://www.linkedin.com/in/ataruc/">Adrian Taruc</a> and <a href="https://www.linkedin.com/in/jamesdalydalton/">James Dalton</a></p><p><em>This is the first entry of a multi-part blog series describing how we built a Real-Time Distributed Graph (RDG). In Part 1, we will discuss the motivation for creating the RDG and the architecture of the data processing pipeline that populates it.</em></p><h3>Introduction</h3><p>The Netflix product experience historically consisted of a single core offering: streaming video on demand. Our members logged into the app, browsed, and watched titles such as Stranger Things, Squid Game, and Bridgerton. Although this is still the core of our product, our business has changed significantly over the last few years. For example, we introduced ad-supported plans, live programming events (e.g., <a href="https://www.netflix.com/title/81764952">Jake Paul vs. Mike Tyson</a> and <a href="https://www.netflix.com/tudum/articles/nfl-games-on-netflix">NFL Christmas Day Games</a>), and <a href="https://about.netflix.com/en/news/let-the-games-begin-a-new-way-to-experience-entertainment-on-mobile">mobile games</a> as part of a Netflix subscription. This evolution of our business has created a new class of problems where we have to analyze member interactions with the app across different business verticals. Let’s walk through a simple example scenario:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*TPFlIvYqGC3L2x1A-KqkyQ.png"></figure><ol><li>Imagine a Netflix member logging into the app on their smartphone and beginning to watch an episode of Stranger Things.</li><li>Eventually, they decide to watch on a bigger screen, so they log into the app on a smart TV in their home and continue watching the same episode.</li><li>Finally, after completing the episode, they log into the app on their tablet and play the game “Stranger Things: 1984”.</li></ol><p>We want to know that these three activities belong to the same member, despite occurring at different times and across various devices. In a traditional data warehouse, these events would land in at least two different tables and may be processed at different cadences. But in a graph system, they become connected almost instantly. Ultimately, analyzing member interactions in the app across domains empowers Netflix to create more personalized and engaging experiences.</p><p>In the early days of our business expansion, discovering these relationships and contextual insights was extremely difficult. Netflix is famous for adopting a microservices architecture — hundreds of microservices developed and maintained by hundreds of individual teams. Some notable benefits of microservices are:</p><ol><li><strong>Service Decomposition</strong>: The overall platform is separated into smaller services, each responsible for a specific business capability. This modularity allows for independent service development, deployment, and scaling.</li><li><strong>Data Isolation</strong>: Each service manages its own data, reducing interdependencies. This allows teams to choose the most suitable data schemas and storage technologies for their services.</li></ol><p><strong>However, these benefits also led to drawbacks for our data science and engineering partners.</strong> In practice, the separation of business concerns and service development ultimately resulted in a separation of data. Manually stitching data together from our data warehouse and siloed databases was an onerous task for our partners. Our data engineering team recognized we needed a solution to process and store our enormous swath of interconnected data while enabling fast querying to discover insights. Although we could have structured the data in various ways, we ultimately settled on a graph representation. We believe a graph offers key advantages, specifically:</p><ul><li><strong>Relationship-Centric Queries:</strong> Graphs enable fast “hops” across multiple nodes and edges without expensive joins or manual denormalization that would be required in table-based data models.</li><li><strong>Flexibility as Relationships Grow:</strong> As new connections and entities emerge, graphs can quickly adapt without significant schema changes or re-architecture.</li><li><strong>Pattern and Anomaly Detection:</strong> Our stakeholders’ use cases often require identifying hidden relationships, cycles, or groupings in the data — capabilities much more naturally expressed and efficiently executed using graph traversals than siloed point lookups.</li></ul><p>This is why we set out to build a Real-Time Distributed Graph, or “RDG” for short.</p><h3>Ingestion and Processing</h3><p>Three main layers in the system power the RDG:</p><ol><li><strong>Ingestion and Processing</strong> — receive events from disparate upstream data sources and use them to generate graph nodes and edges.</li><li><strong>Storage</strong> — write nodes and edges to persistent data stores.</li><li><strong>Serving</strong> — expose ways for internal clients to query graph nodes and edges.</li></ol><p><strong>The rest of this post will focus on the first layer, while subsequent posts in this blog series will cover the other layers.</strong> The diagram below depicts a high-level overview of the ingestion and processing pipeline:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*Jy0eVxvB-AzFNijNfRb9ZA.png"></figure><p>Building and updating the RDG in real-time requires continuously processing vast volumes of incoming data. Batch processing systems and traditional data warehouses cannot offer the low latency needed to maintain an up-to-date graph that supports real-time applications. We opted for a stream processing architecture, enabling us to update the graph’s data as events happen, thus minimizing delay and ensuring the system reflects the latest member interactions within the Netflix app.</p><h3>Kafka as the Ingestion Backbone</h3><p>Member actions in the Netflix app are published to our API Gateway, which then writes them as records to <a href="https://kafka.apache.org/">Apache Kafka</a> topics. Kafka is the mechanism through which internal data applications can consume these events. It provides durable, replayable streams that downstream processors, such as <a href="https://flink.apache.org/">Apache Flink</a> jobs, can consume in real-time.</p><p>Our team’s applications consume several different Kafka topics, each generating up to roughly <strong>1 million messages per second</strong>. Topic records are encoded in the Apache Avro format, and Avro schemas are persisted in an internal centralized schema registry. In order to strike a balance between maintaining data availability and managing the financial expenses of storage infrastructure, we tailor retention policies for each topic according to its throughput and record size. We also persist topic records to <a href="https://iceberg.apache.org/">Apache Iceberg</a> data warehouse tables, which allows us to backfill data in scenarios where older data is no longer available in the Kafka topics.</p><h3>Processing Data with Apache Flink</h3><p>The event records in the Kafka streams are ingested by Flink jobs. We chose Flink because of its strong capabilities around near-real-time event processing. There is also robust internal platform support for Flink within Netflix, which allows jobs to integrate with Kafka and various storage backends seamlessly. At a high level, the anatomy of an RDG Flink job looks like this:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*0G-yrGzB_ZbgaRlPbextTQ.png"></figure><p>For the sake of simplicity, the diagram above depicts a basic flow in which a member logs into their Netflix account and begins watching an episode of Stranger Things. Reading the diagram from left to right:</p><ul><li>The actions of logging into the app and watching the Stranger Things episode are ultimately written as events to Kafka topics.</li><li>The Flink job consumes event records from the upstream Kafka topics.</li><li>Next, we have a series of Flink processor functions that:</li></ul><ol><li>Apply filtering and projections to remove noise based on the individual fields that are present — or in some cases, not present — in the events.</li><li>Enrich events with additional metadata, which are stored and accessed by the processor functions via side inputs.</li><li>Transform events into graph primitives — nodes representing entities (e.g., member accounts and show/movie titles), and edges representing relationships or interactions between them. In this example, the diagram only shows a few nodes and an edge to keep things simple. However, in reality, we create and update up to a few dozen different nodes and edges, depending on the member actions that occurred within the Netflix app.</li><li>Buffer, detect, and deduplicate overlapping updates that occur to the same nodes and edges within a small, configurable time window. This step reduces the data throughput we publish downstream. It is implemented using stateful process functions and timers.</li><li>Publish nodes and edges records to <a href="https://netflixtechblog.com/data-mesh-a-data-movement-and-processing-platform-netflix-1288bcab2873">Data Mesh</a>, an abstraction layer that connects data applications and storage systems. We write a total (nodes + edges) of <strong>more than 5 million records per second</strong> to Data Mesh, which handles persisting the records to various data stores that other internal services can query.</li></ol><h3>From One Job to Many: Scaling Flink the Hard Way</h3><p>Initially, we tried having just one Flink job that consumed all the Kafka source topics. However, this quickly became a big operational headache since different topics can have different data volumes and throughputs at different times during the day. Consequently, tuning the monolithic Flink job became extremely difficult — we struggled to find CPU, memory, job parallelism, and checkpointing interval configurations that ensured job stability.</p><p>Instead, we pivoted to having a 1:1 mapping from the Kafka source topic to the consuming Flink job. Although this led to additional operational overhead due to more jobs to develop and deploy, each job has been much simpler to maintain, analyze, and tune.</p><p>Similarly, each node and edge type is written to a separate Kafka topic. This means we have significantly more Kafka topics to manage. However, we decided the tradeoff of having bespoke tuning and scaling per topic was worth it. We also designed the graph data model to be as generic and flexible as possible, so adding new types of nodes and edges would be an infrequent operation.</p><h3>Acknowledgements</h3><p>We would be remiss if we didn’t give a special shout-out to our stunning colleagues who work on the internal Netflix data platform. Building the RDG was a multi-year effort that required us to design novel solutions, and the investments and foundations from our platform teams were critical to its successful creation. You make the lives of Netflix data engineers much easier, and the RDG would not exist without your diligent collaboration!</p><p>—</p><p>Thanks for reading the first season of the RDG blog series; stay tuned for Season 2, where we will go over the storage layer containing the graph’s various nodes and edges.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=80113e124acc" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/how-and-why-netflix-built-a-real-time-distributed-graph-part-1-ingesting-and-processing-data-80113e124acc">How and Why Netflix Built a Real-Time Distributed Graph: Part 1 — Ingesting and Processing Data…</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/how-and-why-netflix-built-a-real-time-distributed-graph-part-1-ingesting-and-processing-data-80113e124acc</link>
      <guid>https://netflixtechblog.com/how-and-why-netflix-built-a-real-time-distributed-graph-part-1-ingesting-and-processing-data-80113e124acc</guid>
      <pubDate>Fri, 17 Oct 2025 20:42:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[100X Faster: How We Supercharged Netflix Maestro’s Workflow Engine]]></title>
      <description><![CDATA[<p>By <a href="https://www.linkedin.com/in/jheua/">Jun He</a>, <a href="https://www.linkedin.com/in/yingyi-zhang-a0a164111/">Yingyi Zhang</a>, <a href="https://www.linkedin.com/in/spearsem/">Ely Spears</a></p><h3>TL;DR</h3><p>We recently upgraded the Maestro engine to go beyond scalability and improved its performance by <strong>100X</strong>! The overall overhead is reduced from seconds to milliseconds. We have updated the Maestro open source project with this improvement! Please visit the <a href="https://github.com/Netflix/maestro">Maestro GitHub repository</a> to get started. If you find it useful, please <a href="https://github.com/Netflix/maestro">give us a star</a>.</p><h3>Introduction</h3><p>In our previous <a href="https://netflixtechblog.com/maestro-netflixs-workflow-orchestrator-ee13a06f9c78">blog post</a>, we introduced Maestro as a horizontally scalable workflow orchestrator designed to manage large-scale Data/ML workflows at Netflix. Over the past two and a half years, Maestro has achieved its design goal and successfully supported massive workflows with hundreds of thousands of jobs, managing millions of executions daily. As the adoption of Maestro increases at Netflix, new use cases have emerged, driven by Netflix’s evolving business needs, such as Live, Ads, and Games. To meet these needs, some of the workflows are now scheduled on a sub-hourly basis. Additionally, Maestro is increasingly being used for low-latency use cases, such as ad hoc queries, beyond traditional daily or hourly scheduled ETL data pipeline use cases.</p><p>While Maestro excels in orchestrating various heterogeneous workflows and managing user end-to-end development experiences, users have experienced noticeable speedbumps (i.e. ten seconds overhead) from the Maestro engine during workflow executions and development, affecting overall efficiency and productivity. Although being fully scalable to support Netflix-scale use cases, the processing overhead from Maestro internal engine state transitions and lifecycle activities have become a bottleneck, particularly during development cycles. Users have expressed the need for a high performance workflow engine to support iterative development use cases.</p><p>To visualize our end users’ needs for the workflow orchestrator, we create a 5-layer structure graph shown below. Before the change, Maestro reached level 4 but faced challenges to satisfy the user’s needs in level 5. With the new engine design, Maestro is able to power the users to work with their highest capacity and spark joy for end users during their development over the Maestro.</p><figure><img alt="Figure 1. A 5-layer structure showing needs for the workflow orchestrator" src="https://cdn-images-1.medium.com/max/562/1*q871tK1C7Y8VSAXVOmu4Ig.png"><figcaption>Figure 1. A 5-layer structure showing needs for the workflow orchestrator.</figcaption></figure><p>In this blog post, we will share our new engine details, explain our design trade-off decisions, and share learnings from this redesign work.</p><h3>Architectural Evolution of Maestro</h3><h4>Before the change</h4><p>To understand the improvements, we will first revisit the original architecture of Maestro to understand why the overhead is high. The system was divided into three main layers, as illustrated in the diagram below. In the sections that follow we will explain each layer and the role it played in our performance optimization.</p><figure><img alt="Figure 2. The architecture diagram before the evolution." src="https://cdn-images-1.medium.com/max/1024/1*198QCvklU8o6aUrDdaBGRA.png"><figcaption>Figure 2. The architecture diagram before the evolution.</figcaption></figure><p><strong>Maestro API and Step Runtime Layer</strong></p><p>This layer offers seamless integrations with other Netflix services (e.g., compute engines like Spark and Trino). Using Maestro, thousands of practitioners build production workflows using a paved path to access platform services . They can focus primarily on their business logic while relying on Maestro to manage the lifecycle of jobs and workflows plus the integration with data platform services and required integrations such as for authentication, monitoring and alerting. This layer functioned efficiently without introducing significant overhead.</p><p><strong>Maestro Engine Layer</strong></p><p>The Maestro engine serves several crucial functions:</p><ul><li>Managing the lifecycle of workflows, their steps and maintaining their state machines</li><li>Supporting all user actions (e.g., start, restart, stop, pause) on workflow and step entities</li><li>Translating complex Maestro workflow graphs into parallel flows, where each flow is an array of sequentially chained flow tasks, translating every step into a flow task, and then executing transformed flows using the internal flow engine</li><li>Acting as a middle layer to maintain isolation between the Maestro step runtime layer and the underlying flow engine layer</li><li>Implementing required data access patterns and writing Maestro data into the database</li></ul><p>In terms of speed, this layer had acceptable overhead but faced edge cases (e.g. a step might be concurrently executed by two workers at the same time, causing race conditions) due to lacking a strong guarantee from the internal flow engine and the external distributed job queue.</p><p><strong>Maestro Internal Flow Engine Layer</strong></p><p>The Maestro internal flow engine performed <strong>2</strong> primary functions:</p><ul><li>Calling task’s execution functions at a given interval.</li><li>Starting the next tasks in an array of sequential task flows (not a graph), if applicable.</li></ul><p>This foundational layer was based on Netflix OSS Conductor 2.x (<a href="https://github.com/Netflix/conductor/releases/tag/v3.0.0">deprecated since Apr 2021</a>), which requires a dedicated set of separate database tables and distributed job queues.</p><p>The existing implementation of this layer introduces an impactful overhead (e.g. a few seconds to tens of seconds overall delays). The lack of strong guarantees (e.g. exactly once publishing) from this layer leads to race conditions which cause stuck jobs or lost executions.</p><h4>Options to consider</h4><p>We have evaluated three options to address those existing issues:</p><ul><li>Option 1: Implement an internal flow engine optimized for Maestro specific use cases</li><li>Option 2: Upgrade Conductor library to 4.0, which addresses the overheads and offers other improvements and enhancements compared with Conductor 2.X.</li><li>Option 3: Use Temporal as the internal flow engine</li></ul><p>One aspect that influenced our assessment of option two is that Conductor 2 provided a final callback capability in the state machine that was contributed specifically for Maestro’s use case to ensure database synchronization between the Conductor and Maestro engine states. It would require porting this functionality to Conductor 4 though it had been dropped given no other Conductor use cases besides Maestro relied on this. By rewriting the flow engine it would allow removal of several complex internal databases and database synchronization requirements which was attractive for simplifying operational reliability. Given Maestro did not need the full set of state engine features offered by Conductor, this motivated us to consider a flow engine rewrite as a higher priority.</p><p>The decision for Temporal was more straightforward. Temporal is optimized towards facilitating inter-process orchestration and would involve calling an external service to interact with the Temporal flow engine. Given Maestro is operating greater than a million tasks per day, many of which are long running, we felt it was an unnecessary source of risk to couple the DAG engine execution with an external service call. If our requirements went beyond lightweight state transition management we might reconsider because Temporal is a very robust control plane orchestration system, but for our needs it introduced complexity and potential reliability weak spots when there was no direct need for the advanced feature set that it offered.</p><p>After considering Option 2 and Option 3, we developed more conviction that Maestro’s architecture could be greatly simplified by not using a full DAG evaluation engine and having to maintain the state machine for two systems (Maestro and Conductor/Temporal). Therefore, we have decided to go with Option 1.</p><h4>After the change</h4><p>To address these issues, we completely rewrote the Maestro internal flow engine layer to satisfy Maestro’s specific needs and optimize its performance. This new flow engine is lightweight with minimal dependencies, focusing on excelling in the two primary functions mentioned <a href="https://netflixtechblog.com/100x-faster-how-we-supercharged-netflix-maestros-workflow-engine-028e9637f041#4bd7">above</a>. We also replaced existing distributed job queues with internal ones to provide a strong guarantee.</p><p>The new engine is <strong>highly performant, efficient, scalable, and fault-tolerant</strong>. It is the foundation for all upper components of Maestro and provides the following guarantees to avoid race conditions:</p><ul><li>A single step should only be executed by a single worker at any given time</li><li>Step state should never be rolled back</li><li>Steps should always eventually run to a terminal state</li><li>The internal flow state should be eventually consistent with the Maestro workflow state</li><li>External API and user actions should not cause race conditions on the workflow execution</li></ul><p>Here is the new architecture diagram after the change, which is much simpler with less dependencies:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*ZGutO7STwBr81NQ2qAWjYQ.png"><figcaption>Figure 3. The architecture diagram after the evolution.</figcaption></figure><h3>New Flow Engine Optimization</h3><p>The new flow engine significantly boosts speed by maintaining state in memory. It ensures consistency by using Maestro engine’s database as the source of truth for workflow and step states. During bootstrapping, the flow engine rebuilds its in-memory state from the database, improving performance and simplifying the overall architecture. This is in contrast to the previous design in which multiple databases had to be reconciled against one another (Conductor’s tables and Maestro’s tables) or else suffer race conditions and rare orphaned job status.</p><p>The flow engine operates on in-memory flow states, resembling a <a href="https://docs.aws.amazon.com/whitepapers/latest/database-caching-strategies-using-redis/caching-patterns.html#write-through">write through caching pattern</a>. Updates to workflow or step state in the database also update the in-memory flow state. If in-memory state is lost, the flow engine rebuilds it from the database, ensuring eventual consistency and resolving race conditions.</p><p>This design delivers lower latency and higher throughput, avoids inconsistencies from dual persistence, simplifies the architecture, and keeps the in‑memory view eventually consistent with the database.</p><h4>Maintaining Scalability While Gaining Speed</h4><p>With the new engine, we significantly boost performance by collocating flows and their tasks on the same node throughout their lifecycle. Therefore, states of a flow and its tasks will stay in a single node’s memory without persisting to the database. This stickiness and locality bring great performance benefits but inevitably impact scalability since tasks are no longer reassigned to a new worker of the whole cluster in each polling cycle.</p><p>To maintain horizontal scalability, we introduced a flow group concept to partition running flows into groups. In this way, each Maestro flow engine instance only needs to maintain ownership of groups rather than individual flows, reducing maintenance costs (e.g., heartbeat) and simplifying reconciliation by allowing each Maestro node to load flows for a group in batches. Each Maestro node claims ownership of a group of flows through a flow group actor and manages their entire lifecycle via child flow actors. If ownership is lost due to node failure or long JVM GC, another node can claim the group to resume flow executions by reconciling internal state from Maestro database. The following diagram illustrates the ownership maintenance.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*nvUSa5zxZOWi4G8EIh9bQA.png"><figcaption>Figure 4. Ownership maintenance sequence diagram.</figcaption></figure><h4>Flow Partitioning</h4><p>To efficiently distribute traffic, Maestro assigns a consistent group ID to flows/workflows by a simple stable ID assignment method, as shown in the diagram’s Partitioning Function box. We chose this simpler partitioning strategy over advanced ones, e.g. consistent hashing, primarily due to execution and reconciliation costs and consistency challenges in a distributed system.</p><p>Since Maestro decomposes workflows into hierarchical internal flows (e.g., foreach), parent flows need to interact with child flows across different groups. To enable this, the maximal group number from the parent, denoted as N’ in the diagram, is passed down to all child flows. This allows child flows, such as subworkflows or foreach iterations, to recompute their own group IDs and also ensures that a parent flow can always determine the group ID of its child flows using only their workflow identifiers.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*PdP7OEWMFDpbNTZZNuKe7Q.png"><figcaption>Figure 5. Flow group partitioning mechanism diagram.</figcaption></figure><p>After a flow’s group ID is determined, the flow operator routes the flow request to the appropriate node. Each node owns a specific range of group IDs. For example, in the diagram, Node 1 owns groups 0, 1, and 2, while Node 3 owns groups 6, 7, and 8. The groups then contain the individual flows (e.g., Flow A, Flow B).</p><p>In this design, the group size is configurable and nodes can also have different group size configurations. The following diagram shows a flow group partitioning example while the maximal group number is changed during the engine execution without impacting any existing workflows.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*cHk1MpAAcjfKC0GvRL57gA.png"><figcaption>Figure 6. A flow group partitioning example.</figcaption></figure><p>In short, Maestro flow engine shares the group info across the parent and child workflows to provide a flexible and stable partitioning mechanism to distribute work across the cluster.</p><h4>Queue Optimization</h4><p>We replaced both external distributed job queues in the existing system with internal ones, preserving the same fault‑tolerance and recovery guarantees while reducing latency and boosting throughput.</p><p>For the internal flow engine, the queue is a simple in‑memory Java blocking queue. It requires no persistence and can be rebuilt from Maestro state during reconciliation.</p><p>For the Maestro engine, we implemented a database‑backed in‑memory queue that provides <strong>exactly‑once publishing and at‑least‑once delivery guarantees</strong>, addressing multiple edge cases that previously required manual state correction.</p><p>This design is similar to the<a href="https://docs.aws.amazon.com/prescriptive-guidance/latest/cloud-design-patterns/transactional-outbox.html"> transactional outbox pattern</a>. In the same transaction that updates Maestro tables, a row is inserted into the `maestro_queue` table. Upon transaction commit, the job is immediately pushed to a queue worker on the same node, eliminating polling latency. After successful processing, the worker deletes the row from the database. A periodic sweeper re-enqueues any rows whose timeout has expired, ensuring another worker picks them up if a worker stalls or a node fails.</p><p>This design handles failures cleanly. If the transaction fails, both data and message roll back atomically, no partial publishing. If a worker or node fails after commit, the timeout mechanism ensures the job is retried elsewhere. On restart, a node rebuilds its in‑memory queue from the queue table, providing at-least-once delivery guarantee.</p><p>To enhance scalability and avoid contention across event types, each event type is assigned a `queue_id`. Job messages are then partitioned by `queue_id`, optimizing performance and maintaining system efficiency under high load.</p><h3>From Stateless Worker Model to Stateful Actor Model</h3><p>Maestro previously used a shared-nothing stateless worker model with a polling mechanism. When a task started, its identifier was enqueued to a distributed task queue. A worker from the flow engine would pick the task identifier from the queue, load the complete states of the whole workflow (including the flow itself and every task), execute the task interface method once, write the updated task data back to the database, and put the task back in the queue with a polling delay. The worker would then forget this task and start polling the next one.</p><p>That architecture was simple and horizontally scalable (excluding database scalability considerations), but it had drawbacks. The process introduced considerable overhead due to polling intervals and state loading. The time spent in one polling cycle on distributed queues, loading complete states, and other DB queries was significant.</p><p>As Maestro engine decomposes complex workflow graphs into multiple flows, actions might involve multiple flows spanning multiple polling cycles, adding up to significant overhead (around ten seconds in the worst cases). Also, this design didn’t offer strong execution guarantees mainly because the distributed job queue could only provide at-least-once guarantees. Tasks might be dequeued and dispatched to multiple workers, workers might reset states in certain race conditions, or load stale states of other tasks and make incorrect decisions. For example, after a long garbage-collection pause or network hiccup, two workers can pick up the same task: one sets the task status as completed and then unblocks the downstream steps to move forward. However, the other worker, working off stale state, resets the task status back to running, leaving the whole workflow in a conflicting state.</p><p>In the new design, we developed a stateful actor model, keeping internal states in memory. All tasks of a workflow are collocated in the same Maestro node, providing the best performance as states are in the same JVM.</p><h4>Actor-Based Model</h4><p>The new flow engine fits well into an actor model. We also deliberately designed it to allow sharing certain local states (read-only) between parent, child, and sibling actors. This optimization gains performance benefits without losing thread safety due to Maestro’s use cases. We used Java 21’s virtual thread support to implement it with minimal dependencies.</p><p>The new actor-based flow engine is fully message/event-driven and can take actions immediately when events are received, eliminating polling interval delays. To maintain compatibility with the existing polling-based logic, we developed a wakeup mechanism. This model requires flow actors and their child task actors to be collocated in the same JVM for communication over the in-memory queue. Since the Maestro engine already decomposes large-scale workflow instances into many small flows, each flow has a limited number of tasks that fit well into memory.</p><p>Below is a high-level overview of the Maestro execution flow based on the actor model.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*acx3i0wXvowwOm0t0ufOHA.png"><figcaption>Figure 7. The high level overview of the Maestro execution.</figcaption></figure><ul><li>When a workflow starts or during reconciliation, the flow engine inserts (if not existing) or loads the Maestro workflow and step instance from the database, transforming it into the internal flow and task state. This state remains in JVM memory until evicted (e.g., when the workflow instance reaches a terminal state).</li><li>A virtual thread is created for each entity (workflow instance or step attempt) as an actor to handle all updates or actions for this entity, ensuring thread safety and eliminating distributed locks and potential race conditions.</li><li>Each virtual thread actor contains an in-memory state, a thread-safe blocking queue, and a state machine to update states, ensuring thread safety and high efficiency.</li><li>Actors are organized hierarchically, with flow actors managing all their task actors. Flow actors and their task actors are kept in the same JVM for locality benefits, with the ability to relocate flow instances to other nodes if needed.</li><li>An event can wake up a virtual thread by pushing a message to the actor’s job queue, enabling Maestro to move toward an event-driven approach alongside the current polling-based approach.</li><li>A reconciliation process transforms the Maestro data model into the internal flow data.</li></ul><h4>Virtual Thread Based Implementation</h4><p>We chose Java virtual threads to implement various actors (e.g. group actors and flow actors), which simplified the actor model implementation. With a smaller amount of code, we developed a fully functional and highly performant event-driven distributed flow engine. Virtual threads fit very well in use cases like state machine transitions within actors. They are lightweight enough to be created in a large number without Out-Of-Memory risks.</p><p>However, virtual threads can potentially deadlock. They’re not suitable for executing user-provided logic or complex step runtime logic that might depend on external libraries or services outside our control. To address this, we separate flow engine execution from task execution logic by adding a separate worker thread pool (not virtual threads) to run actual step runtime business logic like launching containers or making external API calls. Flow/task actors can <a href="https://github.com/Netflix/maestro/blob/main/maestro-flow/src/main/java/com/netflix/maestro/flow/engine/ExecutionContext.java#L96-L100">wait indefinitely for the future of the thread poll executor to complete</a> but don’t perform actual execution, allowing us to benefit from virtual threads while avoiding deadlock issues.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*ulNZD0dkW4Tj50OaQ0e8aw.png"><figcaption>Figure 8. Virtual thread and worker thread separation.</figcaption></figure><h4>Providing Strong Execution Guarantees</h4><p>To provide strong execution guarantees, we implemented a generation ID-based solution to ensure that a single flow or task is executed by only one actor at any time, with states that never roll back and eventually reach a terminal state.</p><p>When a node claims a new group or a group with an expired heartbeat, it updates the database table row and increments the group generation ID. During node bootstrap, the group actor updates all its owned flows’ generation IDs while rebuilding internal flow states. When creating a new flow, the group actor verifies that the database generation ID matches its in-memory generation ID, otherwise rejecting the creation and reporting a retryable error to the caller. Please check <a href="https://github.com/Netflix/maestro/blob/main/maestro-flow/src/main/java/com/netflix/maestro/flow/dao/MaestroFlowDao.java">the source code</a> for the implementation details.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*UexwsCGLLE-lrzV3HVzCEg.png"><figcaption>Figure 9. An example sequence diagram showing how generation id provides a strong guarantee.</figcaption></figure><p>Additionally, the new flow engine supports both event-driven execution and polling-based periodic reconciliation. Event-driven support allows us to extend polling intervals for state reconciliation at a very low cost, while polling-based reconciliation relaxes event delivery requirements to at-most-once.</p><h3>Testing, Validation and Rollout</h3><p>Migrating hundreds of thousands of Netflix data processing jobs to a new workflow engine required meticulous planning and execution to avoid data corruption, unexpected traffic patterns, and edge cases that could hinder performance gains. We adopted a principled approach to ensure a smooth transition:</p><ol><li><strong>Realistic Testing:</strong> Our testing mirrored real-world use cases as closely as possible.</li><li><strong>Balanced Approach:</strong> We balanced the need for rapid delivery with comprehensive testing.</li><li><strong>Minimal User Disruption:</strong> The goal was for users to be unaware of the underlying changes.</li><li><strong>Clear Communication:</strong> For cases requiring user involvement, clear communication was provided.</li></ol><h4>Maestro Test Framework</h4><p>To achieve our testing goals, we developed an adaptable testing framework for Maestro. This framework addresses the limitations of static unit and integration tests by providing a more dynamic and comprehensive approach, mimicking organic production traffic. It complements existing tests to instill confidence when rolling out major changes, such as new DAG engines.</p><p>The framework is designed to sample real user workflows, disconnecting business logic from external side effects like data reads or writes. This allows us to run workflow graphs of various shapes and sizes, reflecting the diverse use cases across Netflix. While system integrations are handled through deployment pipeline integration tests, the ability to exercise a wide variety of workflow topologies (e.g., parallel executions, for-each jobs, conditional branching and parameter passing between jobs) was crucial for ensuring the new flow engine’s correctness and performance.</p><p>The prototype workflow for the test framework focuses on auto-testing parameters, involving two main steps:</p><p><strong>1. Caching Production Workflows:</strong></p><ul><li>Successful production instances are queried from a historical Maestro feed table over a specified period.</li><li>Run parameters, initiator, and instance IDs are extracted and organized into an instance data map.</li><li>YAML definitions and subworkflow IDs are pulled from S3 storage.</li><li>Both workflow definitions and instance data are cached on S3 for subsequent steps.</li></ul><p><strong>2. Pushing, Running, and Monitoring Workflows:</strong></p><ul><li>Cached workflow definitions and instance data are loaded.</li><li>Notebook-based jobs are replaced with custom notebooks, and certain job types (e.g., vanilla container runtime jobs, templated data movement jobs) and signal triggers are converted to a special no-op job type or skipped.</li><li>Abstract job types like Write-Audit-Publish are expressed as a single step template but are translated to multiple reified nodes of the DAG when executed. These are auto-translated into several custom notebook job types to replace the generated nodes.</li><li>Workflows and subworkflows are pushed, with only non-subworkflows being run using original production instance information.</li><li>1. In the parent workflow, each sub-workflow is replaced with a special no-op placeholder so that the overall topology is preserved but without executing any side-effects of child workflows and avoid cases using dynamic runtime parameter logic.</li><li>2. Each sub-workflow is then separately treated like a top-level parent workflow not initiated from its parent, to exercise the actual workflow steps of the sub-workflow.</li><li>The custom notebook internally compares all passed parameters for each job.</li><li>Workflow instances are monitored until termination (success or failure).</li><li>An email detailing failed workflow instances is generated.</li></ul><p>Future phases of the test framework aim to expand support for native steps, more templates, Titus and Metaflow workflows, and include more robust signal testing. Further integration with the ecosystem, including dedicated Genie clusters for no-op jobs and DGS for our internal workflow UI feature verification, is also being explored.</p><h4>Rollout Plan</h4><p>Our rollout strategy prioritized minimal user disruption. We determined that an entire workflow, from its root instance, must reside in either the old or new flow engine, preventing mixed operations that could lead to complex failure modes and manual data reconciliation.</p><p>To facilitate this, we established a parallel infrastructure for the new workflow engine and leveraged our orchestrator gateway API to hide any routing or redirection logic from users. This approach provided excellent isolation for managing the migration. Initially, specific workflows could explicitly opt in via a system flag, allowing us to observe their execution and gain confidence. By scaling up traffic to the parallel infrastructure in direct proportion to what was scaled down from the original infrastructure, the dual infrastructure cost increase was negligible.</p><p>Once confident, we transitioned to a percentage-based cutover. In the event of a sustained failure in the new engine, our team could roll back a workflow by removing it from the new engine’s database and restarting it in the original stack. However, one consequence of rollback was that failed workflows had to restart from the beginning, recomputing previously successful steps, to ensure all artifacts were generated from a consistent flow engine.</p><p>Leveraging Maestro’s 10-day workflow timeout, we migrated users without disruption. Existing executions would either complete or time out. Upon restarting (due to failure/timeout) or triggering a new instance (due to success), the workflow would be picked up by the new engine. This effectively allowed us to gradually “drain” traffic from the old engine to the new one with no user involvement.</p><p>While the plan generally proceeded as expected with limited edge cases, we did encounter a few challenges:</p><ul><li><strong>Stuck Workflows:</strong> Around 50 workflows with defunct or incorrect ownership information entered a stuck state. In some cases, a backlog of queued instances behind a stuck instance created a race condition in which a new instance would be started immediately when an old instance was terminated, perpetually keeping the workflow on the old engine. For these, we proactively contacted users to negotiate manual stop-and-restart times, forcing them onto the new engine.</li><li><strong>Configuration Discrepancies:</strong> A significant lesson learned was the importance of meticulous record-keeping and management of parallel infrastructure components. We discovered alerts, system flags, and feature flags configured for one stack but not the other. This led to a failure in a partner team’s system that dynamically rolled out a Python migration by analyzing workflow configurations. The absence of a required feature flag in the new engine stack caused the process to be silently skipped, resulting in incorrect Python version configurations for about 40 workflows. Although quickly remediated, this caused user inconvenience as affected workflows needed to be restarted and verified for no lingering data corruption issues. This issue also highlighted limitations in the testing framework since runtime configuration based on external API calls to the configuration service were not exercised in simulated workflow executions.</li></ul><p>Despite these challenges, the migration was a success. We migrated over 60,000 active workflows generating over a million data processing tasks daily with almost no user involvement. By observing the flow engine’s lifecycle management latency, we validated a reduction in step launch overhead from around 5 seconds to 50 milliseconds. Workflow start overhead (incurred once per each workflow execution) also improved from 200 milliseconds to 50 milliseconds. Aggregating this over a million daily step executions translates to saving approximately 57 days of flow engine overhead per day, leading to a snappier user experience, more timely workflow status for data practitioners and greater overall task throughput for the same infrastructure scale.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/826/1*63TFS2uVWjXxlZGpHhTpoA.jpeg"></figure><figure><img alt="" src="https://cdn-images-1.medium.com/max/874/1*yxfbCcs1Y2h4dpxJ5hZkDA.jpeg"></figure><figure><img alt="" src="https://cdn-images-1.medium.com/max/948/1*6Sq7ozdrxqeHI56SvejJkw.png"></figure><p>We additionally realized significant benefits internally with reduced maintenance effort due to the new flow engine’s simplified set of database components. We were able to delete nearly 40TB of obsolete tables related to the previous stateless flow engine and saw a 90% reduction in internal database query traffic which had previously been a significant source of system alerts for the team.</p><h3>Conclusion</h3><p>The architectural evolution of Maestro represents a significant leap in performance, reducing overhead from seconds to milliseconds. This redesign with a stateful actor model not only enhances speed by 100X but also maintains scalability and reliability, ensuring Maestro continues to meet the diverse needs of Netflix’s data and ML workflows.</p><p>Key takeaways from this evolution include:</p><ul><li><strong>Performance matters:</strong> Even in a system designed for scale, the speed of individual operations significantly impacts user experience and productivity.</li><li><strong>Simplicity wins:</strong> Reducing dependencies and simplifying architecture not only improved performance but also enhanced reliability and maintainability.</li><li><strong>Strong guarantees are essential:</strong> Providing strong execution guarantees eliminates race conditions and edge cases that previously required manual intervention.</li><li><strong>Locality optimizations pay off:</strong> Collocating related flows and tasks in the same JVM dramatically reduces overhead from the Maestro engine.</li><li><strong>Modern language features help:</strong> Java 21’s virtual threads enabled an elegant actor-based implementation with minimal code complexity and dependencies.</li></ul><p>We’re excited to share these improvements with the open-source community and look forward to seeing how Maestro continues to evolve. The performance gains we’ve achieved open new possibilities for low-latency workflow orchestration use cases while continuing to support the massive scale that Netflix and other organizations require.</p><p>Visit the <a href="https://github.com/Netflix/maestro">Maestro GitHub repository</a> to explore these improvements. If you have any questions, thoughts, or comments about Maestro, please feel free to create a <a href="https://github.com/Netflix/maestro/issues">GitHub issue</a> in the Maestro repository. We are eager to hear from you. If you are passionate about solving large scale orchestration problems, please <a href="https://explore.jobs.netflix.net/careers?query=Data%20Platform&amp;Teams=Engineering&amp;domain=netflix.com&amp;sort_by=relevance">join us</a>.</p><h3>Acknowledgements</h3><p>Special thanks to Big Data Orchestration team members for general contributions to Maestro and diligent review, discussion and incident response required to make this project successful: Davis Shepherd, Natallia Dzenisenka, Praneeth Yenugutala, Brittany Truong, Jonathan Indig, Deepak Ramalingam, Binbing Hou, Zhuoran Dong, Victor Dusa, and Gabriel Ikpaetuk — and and internal partners Yun Li and Romain Cledat.</p><p>Thank you to Anoop Panicker and Aravindan Ramkumar from our partner organization that leads Conductor development in Netflix. They helped us understand issues in Conductor 2.X that initially motivated the rearchitecture and helped provide context on later versions of Conductor that defined some of the core trade-offs for the decision to implement a custom DAG engine in Maestro.</p><p>We’d also like to thank our partners on the Data Security &amp; Infrastructure and Engineering Support teams who helped identify and rapidly fix the configuration discrepancy error encountered during production rollout: Amer Hesson, Ye Ji, Sungmin Lee, Brandon Quan, Anmol Khurana, and Manav Garekar.</p><p>A special thanks also goes out to partners from the Data Experience team including Jeff Bothe, Justin Wei, and Andrew Seier. The flow engine speed improvement was actually so dramatic that it broke some integrations with our internal workflow UI that reported state transition durations. Our partners helped us catch and fix UI regressions before they shipped to avoid impact to users.</p><p>We also thank Prashanth Ramdas, Anjali Norwood, Eva Tse, Charles Zhao, Sumukh Shivaprakash, Joey Lynch, Harikrishna Menon, Marcelo Mayworm, Charles Smith and other leaders for their constructive feedback and guidance on the Maestro project.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=028e9637f041" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/100x-faster-how-we-supercharged-netflix-maestros-workflow-engine-028e9637f041">100X Faster: How We Supercharged Netflix Maestro’s Workflow Engine</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/100x-faster-how-we-supercharged-netflix-maestros-workflow-engine-028e9637f041</link>
      <guid>https://netflixtechblog.com/100x-faster-how-we-supercharged-netflix-maestros-workflow-engine-028e9637f041</guid>
      <pubDate>Mon, 29 Sep 2025 16:10:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Building a Resilient Data Platform with Write-Ahead Log at Netflix]]></title>
      <description><![CDATA[<p>By <a href="https://www.linkedin.com/in/prudhviraj9">Prudhviraj Karumanchi</a>, <a href="https://www.linkedin.com/in/samuelfu/">Samuel Fu</a>, <a href="https://www.linkedin.com/in/sriram-rangarajan-35169715/">Sriram Rangarajan</a>, <a href="https://www.linkedin.com/in/vidhya-arvind-11908723">Vidhya Arvind</a>, <a href="https://www.linkedin.com/in/yunwang-io/">Yun Wang</a>, <a href="https://www.linkedin.com/in/john-l-693b7915a/">John Lu</a></p><h3>Introduction</h3><p>Netflix operates at a massive scale, serving hundreds of millions of users with diverse content and features. Behind the scenes, ensuring data consistency, reliability, and efficient operations across various services presents a continuous challenge. At the heart of many critical functions lies the concept of a Write-Ahead Log (WAL) abstraction. At Netflix scale, every challenge gets amplified. Some of the key challenges we encountered include:</p><ul><li>Accidental data loss and data corruption in databases</li><li>System entropy across different datastores (e.g., writing to Cassandra and Elasticsearch)</li><li>Handling updates to multiple partitions (e.g., building secondary indices on top of a NoSQL database)</li><li>Data replication (in-region and across regions)</li><li>Reliable retry mechanisms for<strong> </strong>real time data pipeline at scale</li><li>Bulk deletes to database causing OOM on the Key-Value nodes</li></ul><p>All the above challenges either resulted in production incidents or outages, consumed significant engineering resources, or led to bespoke solutions and technical debt. During one particular incident, a developer issued an ALTER TABLE command that led to data corruption. Fortunately, the data was fronted by a cache, so the ability to extend cache TTL quickly together with the app writing the mutations to Kafka allowed us to recover. Absent the resilience features on the application, there would have been permanent data loss. As the data platform team, we needed to provide resilience and guarantees to protect not just this application, but all the critical applications we have at Netflix.</p><p>Regarding the retry mechanisms for real time data pipelines, Netflix operates at a massive scale where failures (network errors, downstream service outages, etc.) are inevitable. We needed a reliable and scalable way to retry failed messages, without sacrificing throughput.</p><p>With these problems in mind, we decided to build a system that would solve all the aforementioned issues and continue to serve the future needs of Netflix in the online data platform space. Our Write-Ahead Log (WAL) is a distributed system that captures data changes, provides strong durability guarantees, and reliably delivers these changes to downstream consumers. This blog post dives into how Netflix is building a generic WAL solution to address common data challenges, enhance developer efficiency, and power high-leverage capabilities like secondary indices, enable cross-region replication for non-replicated storage engines, and support widely used patterns like delayed queues.</p><h3>API</h3><p>Our API is intentionally simple, exposing just the essential parameters. WAL has one main API endpoint, <em>WriteToLog</em>, abstracting away the internal implementation and ensuring that users can onboard easily.</p><pre>rpc WriteToLog (WriteToLogRequest) returns (WriteToLogResponse) {...}<br><br>/**<br>  * WAL request message<br>  * namespace: Identifier for a particular WAL<br>  * lifecycle: How much delay to set and original write time <br>  * payload: Payload of the message<br>  * target: Details of where to send the payload <br>  */<br>message WriteToLogRequest {<br>  string namespace = 1;<br>  Lifecycle lifecycle = 2;<br>  bytes payload = 3;<br>  Target target = 4;<br>}<br><br>/**<br>  * WAL response message<br>  * durable: Whether the request succeeded, failed, or unknown<br>  * message: Reason for failure<br>  */<br>message WriteToLogResponse {<br>  Trilean durable = 1;<br>  string message = 2;<br>}</pre><p>A <em>namespace</em> defines where and how data is stored, providing logical separation while abstracting the underlying storage systems. Each <em>namespace</em> can be configured to use different queues: Kafka, SQS, or combinations of multiple. <em>Namespace</em> also serves as a central configuration of settings, such as backoff multiplier or maximum number of retry attempts, and more. This flexibility allows our Data Platform to route different use cases to the most suitable storage system based on performance, durability, and consistency needs.</p><p>WAL can assume different <em>personas</em> depending on the namespace configuration.</p><h4><strong>Persona #1 (<em>Delayed Queues</em>)</strong></h4><p>In the example configuration below, the Product Data Systems (PDS) <em>namespace</em> uses SQS as the underlying message queue, enabling delayed messages. PDS uses Kafka extensively, and failures (network errors, downstream service outages, etc.) are inevitable. We needed a reliable and scalable way to retry failed messages, without sacrificing throughput. That’s when PDS started leveraging WAL for delayed messages.</p><pre>"persistenceConfigurations": {<br>  "persistenceConfiguration": [<br>  {<br>    "physicalStorage": {<br>      "type": "SQS",<br>    },<br>    "config": {<br>      "wal-queue": [<br>        "dgwwal-dq-pds"<br>      ],<br>      "wal-dlq-queue": [<br>        "dgwwal-dlq-pds"<br>      ],<br>      "queue.poll-interval.secs": 10,<br>      "queue.max-messages-per-poll": 100<br>    }<br>  }<br>  ]<br>}</pre><h4><strong>Persona #2 (<em>Generic Cross-Region Replication</em>)</strong></h4><p>Below is the namespace configuration for cross-region replication of <a href="https://netflixtechblog.com/caching-for-a-global-netflix-7bcc457012f1">EVCache</a> using WAL, which replicates messages from a source region to multiple destinations. It uses Kafka under the hood.</p><pre>"persistence_configurations": {<br>  "persistence_configuration": [<br>  {<br>    "physical_storage": {<br>      "type": "KAFKA"<br>    },<br>    "config": {<br>      "consumer_stack": "consumer",<br>      "context": "This is for cross region replication for evcache_foobar",<br>      "target": {<br>        "euwest1": "dgwwal.foobar.cluster.eu-west-1.netflix.net",<br>        "type": "evc-replication",<br>        "useast1": "dgwwal.foobar.cluster.us-east-1.netflix.net",<br>        "useast2": "dgwwal.foobar.cluster.us-east-2.netflix.net",<br>        "uswest2": "dgwwal.foobar.cluster.us-west-2.netflix.net"<br>      },<br>      "wal-kafka-dlq-topics": [],<br>      "wal-kafka-topics": [<br>        "evcache_foobar"<br>      ],<br>      "wal.kafka.bootstrap.servers.prefix": "kafka-foobar"<br>    }<br>  }<br>  ]<br>}</pre><h4><strong>Persona #3 (Handling <em>multi-partition mutations</em>)</strong></h4><p>Below is the namespace configuration for supporting <em>mutateItems</em> API in <a href="https://netflixtechblog.com/introducing-netflixs-key-value-data-abstraction-layer-1ea8a0a11b30">Key-Value</a>, where multiple write requests can go to different partitions and have to be eventually consistent. A key detail in the below configuration is the presence of Kafka and durable_storage. These data stores are required to facilitate two phase commit semantics, which we will discuss in detail below.</p><pre>"persistence_configurations": {<br>  "persistence_configuration": [<br>  {<br>    "physical_storage": {<br>      "type": "KAFKA"<br>    },<br>    "config": {<br>      "consumer_stack": "consumer",<br>      "contacts": "unknown",<br>      "context": "WAL to support multi-id/namespace mutations for dgwkv.foobar",<br>      "durable_storage": {<br>        "namespace": "foobar_wal_type",<br>        "shard": "walfoobar",<br>        "type": "kv"<br>      },<br>      "target": {},<br>      "wal-kafka-dlq-topics": [<br>        "foobar_kv_multi_id-dlq"<br>      ],<br>      "wal-kafka-topics": [<br>        "foobar_kv_multi_id"<br>      ],<br>      "wal.kafka.bootstrap.servers.prefix": "kaas_kafka-dgwwal_foobar7102"<br>    }<br>  }<br>  ]<br>}</pre><p>An important note is that requests to WAL support at-least once semantics due to the underlying implementation.</p><h3>Under the Hood</h3><p>The core architecture consists of several key components working together.</p><p><strong>Message Producer and Message Consumer separation:</strong> The message producer receives incoming messages from client applications and adds them into the queue, while the message consumer processes messages from the queue and sends them to the targets. Because of this separation, other systems can bring their own pluggable producers or consumers, depending on their use cases. WAL’s control plane allows for a pluggable model, which, depending on the use-case, allows us to switch between different message queues.</p><p><strong>SQS and Kafka with a dead letter queue by default</strong>: Every WAL <em>namespace</em> has its own message queue and gets a dead letter queue (DLQ) by default, because there can be transient errors and hard errors. Application teams using <a href="https://netflixtechblog.com/introducing-netflixs-key-value-data-abstraction-layer-1ea8a0a11b30">Key-Value</a> abstraction simply need to toggle a flag to enable WAL and get all this functionality without needing to understand the underlying complexity.</p><ul><li><strong>Kafka-backed namespaces</strong>: handle standard message processing</li><li><strong>SQS-backed namespaces</strong>: support delayed queue semantics (we added custom logic to go beyond the standard defaults enforced in terms of delay, size limits, etc)</li><li><strong>Complex multi-partition scenarios:</strong> use queues and durable storage</li></ul><p><strong>Target Flexibility</strong>: The messages added to WAL are pushed to the target datastores. Targets can be Cassandra databases, Memcached caches, Kafka queues, or upstream applications. Users can specify the target via namespace configuration and in the API itself.</p><figure><img alt="Architecture of WAL" src="https://cdn-images-1.medium.com/max/1024/1*tfnrbP7oD_r9iEesLhACpA.png"><figcaption>Architecture of WAL</figcaption></figure><h3>Deployment Model</h3><p>WAL is deployed using the <a href="https://netflixtechblog.medium.com/data-gateway-a-platform-for-growing-and-protecting-the-data-tier-f1ed8db8f5c6">Data Gateway infrastructure</a>. This means that WAL deployments automatically come with mTLS, connection management, authentication, runtime and deployment configurations out of the box.</p><p>Each data gateway abstraction (including WAL) is deployed as a <em>shard</em>. A <em>shard</em> is a physical concept describing a group of hardware instances. Each use case of WAL is usually deployed as a separate <em>shard</em>. For example, the Ads Events service will send requests to WAL <em>shard A</em>, while the Gaming Catalog service will send requests to WAL <em>shard </em>B, allowing for separation of concerns and avoiding noisy neighbour problems.</p><p>Each <em>shard</em> of WAL can have multiple <em>namespaces</em>. A <em>namespace</em> is a logical concept describing a configuration. Each request to WAL has to specify its <em>namespace</em> so that WAL can apply the correct configuration to the request. Each <em>namespace</em> has its own configuration of queues to ensure isolation per use case. If the underlying queue of a WAL <em>namespace</em> becomes the bottleneck of throughput, the operators can choose to add more queues on the fly by modifying the <em>namespace</em> configurations. The concept of <em>shards</em> and <em>namespaces</em> is shared across all Data Gateway Abstractions, including <a href="https://netflixtechblog.com/introducing-netflixs-key-value-data-abstraction-layer-1ea8a0a11b30">Key-Value</a>, <a href="https://netflixtechblog.com/netflixs-distributed-counter-abstraction-8d0c45eb66b2">Counter</a>, <a href="https://netflixtechblog.com/introducing-netflix-timeseries-data-abstraction-layer-31552f6326f8">Timeseries</a>, etc. The <em>namespace</em> configurations are stored in a globally replicated Relational SQL database to ensure availability and consistency.</p><figure><img alt="Deployment model of WAL" src="https://cdn-images-1.medium.com/max/1024/1*Gh2O_tTvxZxlbRmKn9Atag.png"><figcaption>Deployment model of WAL</figcaption></figure><p>Based on certain CPU and network thresholds, the Producer group and the Consumer group of each <em>shard</em> will (separately) automatically scale up the number of instances to ensure the service has low latency, high throughput and high availability. WAL, along with other abstractions, also uses the <a href="https://netflixtechblog.medium.com/performance-under-load-3e6fa9a60581">Netflix adaptive load shedding libraries</a> and Envoy to automatically shed requests beyond a certain limit. WAL can be deployed to multiple regions, so each region will deploy its own group of instances.</p><h3>Solving different flavors of problems with no change to the core architecture</h3><p>The WAL addresses multiple data reliability challenges with no changes to the core architecture:</p><p><strong>Data Loss Prevention:</strong> In case of database downtime, WAL can continue to hold the incoming mutations. When the database becomes available again, replay mutations back to the database. The tradeoff is eventual consistency rather than immediate consistency, and no data loss.</p><p><strong>Generic Data Replication:</strong> For systems like EVCache (using Memcached) and RocksDB that do not support replication by default, WAL provides systematic replication (both in-region and across-region). The target can be another application, another WAL, or another queue — it’s completely pluggable through configuration.</p><p><strong>System Entropy and Multi-Partition Solutions: </strong>Whether dealing with writes across two databases (like Cassandra and Elasticsearch) or mutations across multiple partitions in one database, the solution is the same — write to WAL first, then let the WAL consumer handle the mutations. No more asynchronous repairs needed; WAL handles retries and backoff automatically.</p><p><strong>Data Corruption Recovery:</strong> In case of DB corruptions, restore to the last known good backup, then replay mutations from WAL omitting the offending write/mutation.</p><p>There are some major differences between using WAL and directly using Kafka/SQS. WAL is an abstraction on the underlying queues, so the underlying technology can be swapped out depending on use cases with no code changes. WAL emphasizes an easy yet effective API that saves users from complicated setups and configurations. We leverage the control plane to pivot technologies behind WAL when needed without app or client intervention.</p><h3>WAL usage at Netflix</h3><h4>Delay Queue</h4><p>The most common use case for WAL is as a Delay Queue. If an application is interested in sending a request at a certain time in the future, it can offload its requests to WAL, which guarantees that their requests will land after the specified delay.</p><p>Netflix’s Live Origin processes and delivers Netflix live stream video chunks, storing its video data in a Key-Value <em>abstraction</em> backed by Cassandra and EVCache. When Live Origin decides to delete certain video data after an event is completed, it issues delete requests to the Key-Value abstraction. However, the large amount of delete requests in a short burst interfere with the more important real-time read/write requests, causing performance issues in Cassandra and timeouts for the incoming live traffic. To get around this, Key-Value issues the delete requests to WAL first, with a random delay and jitter set for each delete request. WAL, after the delay, sends the delete requests back to Key-Value. Since the deletes are now a flatter curve of requests over time, Key-Value is then able to send the requests to the datastore with no issues.</p><figure><img alt="Requests being spread out over time through delayed requests" src="https://cdn-images-1.medium.com/max/1024/1*7JV6kc5QMyfviAJIjVXqew.png"><figcaption>Requests being spread out over time through delayed requests</figcaption></figure><p>Additionally, WAL is used by many services that utilize Kafka to stream events, including Ads, Gaming, Product Data Systems, etc. Whenever Kafka requests fail for any reason, the client apps will send WAL a request to retry the kafka request with a delay. This abstracts away the backoff and retry layer of Kafka for many teams, increasing developer efficiency.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*xnAYzUCsEQ18qyCOnUmuPg.png"><figcaption>Backoff and delayed retries for clients producing to Kafka</figcaption></figure><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*sbEVgICN9qbEecEPYiSHWQ.png"><figcaption>Backoff and delayed retries for clients consuming from Kafka</figcaption></figure><h4>Cross-Region Replication</h4><p>WAL is also used for global cross-region replication. The architecture of WAL is generic and allows any datastore/applications to onboard for cross-region replication. Currently, the largest use case is <a href="https://netflixtechblog.com/caching-for-a-global-netflix-7bcc457012f1">EVCache</a>, and we are working to onboard other storage engines.</p><p>EVCache is deployed by clusters of Memcached instances across multiple regions, where each cluster in each region shares the same data. Each region’s client apps will write, read, or delete data from the EVCache cluster of the same region. To ensure global consistency, the EVCache client of one region will replicate write and delete requests to all other regions. To implement this, the EVCache client that originated the request will send the request to a WAL corresponding to the EVCache cluster and region.</p><p>Since the EVCache client acts as the message producer group in this case, WAL only needs to deploy the message consumer groups. From there, the multiple message consumers are set up to each target region. They will read from the Kafka topic, and send the replicated write or delete requests to a Writer group in their target region. The Writer group will then go ahead and replicate the request to the EVCache server in the same region.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*twUIRRILpBEXb085wCdP7A.png"><figcaption>EVCache Global Cross-Region Replication Implemented through WAL</figcaption></figure><p>The biggest benefits of this approach, compared to our legacy architecture, is being able to migrate from multi-tenant architecture to single tenant architecture for the most latency sensitive applications. For example, Live Origin will have its own dedicated Message Consumer and Writer groups, while a less latency sensitive service can be multi-tenant. This helps us reduce the blast radius of the issues and also prevents noisy neighbor issues.</p><h4>Multi-Table Mutations</h4><p>WAL is used by <a href="https://netflixtechblog.com/introducing-netflixs-key-value-data-abstraction-layer-1ea8a0a11b30">Key-Value</a> service to build the MutateItems API. WAL enables the API’s multi-table and multi-id mutations by implementing 2-phase commit semantics under the hood. For this discussion, we can assume that Key-Value service is backed by Cassandra, and each of its <em>namespaces</em> represents a certain table in a Cassandra DB.</p><p>When a Key-Value client issues a MutateItems request to Key-Value server, the request can contain multiple PutItems or DeleteItems requests. Each of those requests can go to different ids and <em>namespaces</em>, or Cassandra tables.</p><pre>message MutateItemsRequest {<br> repeated MutationRequest mutations = 1;<br> message MutationRequest {<br>  oneof mutation {<br>    PutItemsRequest put = 1;<br>    DeleteItemsRequest delete = 2;<br>  }<br> }<br>}</pre><p>The MutateItems request operates on an eventually consistent model. When the Key-Value server returns a success response, it guarantees that every operation within the MutateItemsRequest will eventually complete successfully. Individual put or delete operations may be partitioned into smaller chunks based on request size, meaning a single operation could spawn multiple chunk requests that must be processed in a specific sequence.</p><p>Two approaches exist to ensure Key-Value client requests achieve success. The synchronous approach involves client-side retries until all mutations complete. However, this method introduces significant challenges; datastores might not natively support transactions and provide no guarantees about the entire request succeeding. Additionally, when more than one replica set is involved in a request, latency occurs in unexpected ways, and the entire request chain must be retried. Also, partial failures in synchronous processing can leave the database in an inconsistent state if some mutations succeed while others fail, requiring complex rollback mechanisms or leaving data integrity compromised. The asynchronous approach was ultimately adopted to address these performance and consistency concerns.</p><p>Given Key-Value’s stateless architecture, the service cannot maintain the mutation success state or guarantee order internally. Instead, it leverages a Write-Ahead Log (WAL) to guarantee mutation completion. For each MutateItems request, Key-Value forwards individual put or delete operations to WAL as they arrive, with each operation tagged with a sequence number to preserve ordering. After transmitting all mutations, Key-Value sends a completion marker indicating the full request has been submitted.</p><p>The WAL producer receives these messages and persists the content, state, and ordering information to a durable storage. The message producer then forwards only the completion marker to the message queue. The message consumer retrieves these markers from the queue and reconstructs the complete mutation set by reading the stored state and content data, ordering operations according to their designated sequence. Failed mutations trigger re-queuing of the completion marker for subsequent retry attempts.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*voSQEpItvosZjPqeTpBseg.png"><figcaption>Architecture of Multi-Table Mutations through WAL</figcaption></figure><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*EVFnGkr57z5cSB5kNI91iA.png"><figcaption>Sequence diagram for Multi-Table Mutations through WAL</figcaption></figure><h3>Closing Thoughts</h3><p>Building Netflix’s generic Write-Ahead Log system has taught us several key lessons that guided our design decisions:</p><p><strong>Pluggable Architecture is Core: </strong>The ability to support different targets, whether databases, caches, queues, or upstream applications, through configuration rather than code changes has been fundamental to WAL’s success across diverse use cases.</p><p><strong>Leverage Existing Building Blocks: </strong>We had control plane infrastructure, Key-Value abstractions, and other components already in place. Building on top of these existing abstractions allowed us to focus on the unique challenges WAL needed to solve.</p><p><strong>Separation of Concerns Enables Scale:</strong> By separating message processing from consumption and allowing independent scaling of each component, we can handle traffic surges and failures more gracefully.</p><p><strong>Systems Fail — Consider Tradeoffs Carefully: </strong>WAL itself has failure modes, including traffic surges, slow consumers, and non-transient errors. We use abstractions and operational strategies like data partitioning and backpressure signals to handle these, but the tradeoffs must be understood.</p><h3>Future work</h3><ul><li>We are planning to add secondary indices in Key-Value service leveraging WAL.</li><li>WAL can also be used by a service to guarantee sending requests to multiple datastores. For example, a database and a backup, or a database and a queue at the same time etc.</li></ul><h3>Acknowledgements</h3><p>Launching WAL was a collaborative effort involving multiple teams at Netflix, and we are grateful to everyone who contributed to making this idea a reality. We would like to thank the following teams for their roles in this launch.</p><ul><li>Caching team — Additional thanks to <a href="https://www.linkedin.com/in/shihhaoyeh/">Shih-Hao Yeh</a>, <a href="https://www.linkedin.com/in/akashdeepgoel/">Akashdeep Goel</a> for contributing to cross region replication for KV, EVCache etc. and owning this service.</li><li>Product Data System team — <a href="https://www.linkedin.com/in/carlos-jmh/">Carlos Matias Herrero</a>, <a href="https://www.linkedin.com/in/bbremen/">Brandon Bremen</a> for contributing to the delay queue design and being early adopters of WAL giving valuable feedback.</li><li>KeyValue and Composite abstractions team — <a href="https://www.linkedin.com/in/rummadis/">Raj Ummadisetty</a> for feedback on API design and mutateItems design discussions. <a href="https://www.linkedin.com/in/rajiv-shringi/">Rajiv Shringi</a><strong> </strong>for feedback on API design.</li><li>Kafka and Real Time Data Infrastructure teams — <a href="https://www.linkedin.com/in/nickmahilani/">Nick Mahilani</a> for feedback and inputs on integrating the WAL client into Kafka client. <a href="https://www.linkedin.com/in/sundaram-ananthanarayanan-97b8b545/">Sundaram Ananthanarayan</a> for design discussions around the possibility of leveraging Flink for some of the WAL use cases.</li><li><a href="https://jolynch.github.io/">Joseph Lynch</a> for providing strategic direction and organizational support for this project.</li></ul><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=127b6712359a" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/building-a-resilient-data-platform-with-write-ahead-log-at-netflix-127b6712359a">Building a Resilient Data Platform with Write-Ahead Log at Netflix</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/building-a-resilient-data-platform-with-write-ahead-log-at-netflix-127b6712359a</link>
      <guid>https://netflixtechblog.com/building-a-resilient-data-platform-with-write-ahead-log-at-netflix-127b6712359a</guid>
      <pubDate>Fri, 26 Sep 2025 20:57:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Scaling Muse: How Netflix Powers Data-Driven Creative Insights at Trillion-Row Scale]]></title>
      <description><![CDATA[<p>By <a href="https://www.linkedin.com/in/andrew-pierce-34443a7/">Andrew Pierce</a>, <a href="https://www.linkedin.com/in/chris-thrailkill-a268914/">Chris Thrailkill</a>, <a href="https://www.linkedin.com/in/victor-chiapaikeo-974a501b/">Victor Chiapaikeo</a></p><p>At Netflix, we prioritize getting timely data and insights into the hands of the people who can act on them. One of our key internal applications for this purpose is Muse. Muse’s ultimate goal is to help Netflix members discover content they’ll love by ensuring our promotional media is as effective and authentic as possible. It achieves this by equipping creative strategists and launch managers with data-driven insights showing which artwork or video clips resonate best with global or regional audiences and flagging outliers such as potentially misleading (clickbait-y) assets. These kinds of applications fall under Online Analytical Processing (OLAP), a category of systems designed for complex querying and data exploration. However, enabling Muse to support new, more advanced filtering and grouping capabilities while maintaining high performance and data accuracy has been a challenge. Previous posts have touched on <a href="https://netflixtechblog.com/artwork-personalization-c589f074ad76">artwork personalization</a> and our <a href="https://netflixtechblog.com/introducing-impressions-at-netflix-e2b67c88c9fb">impressions architecture</a>. <strong>In this post, we’ll discuss some steps we’ve taken to evolve the Muse data serving layer to enable new capabilities while maintaining high performance and data accuracy.</strong></p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*22EDwImi4b5D3tZuY8KXjQ.png"><figcaption>Muse application</figcaption></figure><h3>An Evolving Architecture</h3><p>Like many early analytics applications, Muse began as a simple dashboard powered by batch data pipelines (Spark¹) and a modest Druid² cluster. As the application evolved, so did user demands. Users wanted new features like outlier detection and notification delivery, media comparison and playback, and advanced filtering, all while requiring lower latency and supporting ever-growing datasets (in the order of trillions of rows a year). One of the most challenging requirements was enabling dynamic analysis of promotional media performance by “audience” affinities: internally defined, algorithmically inferred labels representing collections of viewers with similar tastes. Answering questions like “Does specific promotional media resonate more with Character Drama fans or Pop Culture enthusiasts?” required augmenting already voluminous impression and playback data. Supporting filtering and grouping by these many-to-many audience relationships led to a combinatorial explosion in data volume, pushing the limits of our original architecture.</p><p>To address these complexities and support the evolving needs of our users, we undertook a significant evolution of Muse’s architecture. Today’s Muse is a React app that queries a GraphQL layer served with a set of Spring Boot GRPC microservices. In the remainder of this post, we’ll focus on steps we took to scale the data microservice, its backing ETL, and our Druid cluster. <strong>Specifically, we’ve changed the data model to rely on HyperLogLog (HLL) sketches, used </strong><a href="https://hollow.how/"><strong>Hollow</strong></a><strong> for access to in-memory, precomputed aggregates, and taken a series of steps to tune Druid. To ensure the accuracy of these changes, we relied heavily on internal debugging tools to validate pre- and post-changes.</strong></p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*-VJUGD9aJ3BNr-q0MgawIg.png"><figcaption><em>Muse’s Current Architecture</em></figcaption></figure><h4>Moving to HyperLogLog (HLL) Sketches for Distinct Counts</h4><p>Some of the most important metrics we track are impressions, the number of times an asset is shown to a user within a time window, and qualified plays, which links a playback event with a minimum duration back to a specific impression. Calculating these metrics requires counting distinct users. However, performing distinct counts in distributed systems is resource-intensive and challenging. For instance, to determine how many unique profiles have ever seen a particular asset, we need to compare each new set of profile ids with those from all days before it, potentially spanning months or even years.</p><p>For performance, we can trade accuracy. The <a href="https://datasketches.apache.org/">Apache Datasketches library</a> allows us to get distinct count estimates that are within a 1–2% error. This is tunable with a precision parameter called logK (0.8% in our case with logK of 17). We build sketches in two places:</p><ol><li>During Druid ingest: we use the <a href="https://druid.apache.org/docs/latest/development/extensions-core/datasketches-hll/#aggregators">HLLSketchBuild aggregator</a> with Druid <a href="https://druid.apache.org/docs/latest/ingestion/rollup/">rollup set to true</a> to reduce our data in preparation for fast distinct counting</li><li>During our Spark ETL: we persist precomputed aggregates like all-time impressions per asset in the form of HLL sketches. Each day, we merge a new HLL sketch into the existing one using a combination of <a href="https://spark.apache.org/docs/3.5.1/api/java/org/apache/spark/sql/functions.html#hll_union(org.apache.spark.sql.Column,org.apache.spark.sql.Column)">hll_union</a> and <a href="https://spark.apache.org/docs/3.5.1/api/java/org/apache/spark/sql/functions.html#hll_union_agg(org.apache.spark.sql.Column)">hll_union_agg</a> (<a href="https://www.databricks.com/blog/apache-spark-3-apache-datasketches-new-sketch-based-approximate-distinct-counting">functions added by our very own Ryan Berti</a>)</li></ol><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*29CBkjnyIn_1e9V1R0m8hg.png"><figcaption><em>We use Datasketches in our ETL and serving systems</em></figcaption></figure><p>HLL has been a huge performance boost for us both within the serving and ETL layer. Across our most common OLAP query patterns, we’ve seen latencies reduce by approx 50%. Nevertheless, running APPROX_COUNT_DISTINCT over large date ranges on the Druid cluster for very large titles exhausts limited threads, especially in high-concurrency situations. To further offload Druid query volume and preserve cluster threads, we’ve also relied extensively on the <a href="https://github.com/Netflix/hollow">Hollow library</a>.</p><h4>Hollow as a Read-Only Key Value Store for Precomputed Aggregates</h4><p>Our in-house Hollow³ infrastructure allows us to easily create Hollow feeds — essentially highly compressed and performant in-memory key/value stores — from Iceberg⁴ tables. In this setup, dedicated producer servers listen for changes to Iceberg tables, and when updates occur, they push the latest data to downstream consumers. On the consumer side, our Spring Boot applications listen to announcements from these producers and automatically refresh in-memory caches with the latest dataset.</p><p>This architecture has enabled us to migrate several data access patterns from Druid to Hollow, specifically ones with a limited number of parameter combinations per title. One of these was fetching distinct filter dimensions. For example, while most Netflix-branded titles are released globally, licensed titles often have rights restrictions that limit their availability to specific countries and time windows. As a result, a particular licensed title might only be available to members in Germany and Luxembourg.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1014/1*m4Tdj3YixQfPmnEaMGdspA.png"><figcaption><em>Distinct countries queried from a Hollow feed for the assets for Manta Manta</em></figcaption></figure><p>In the past, retrieving these distinct country values per asset required issuing a SELECT DISTINCT query to our Druid cluster. With Hollow, we maintain a feed of distinct dimension values, allowing us to perform stream operations like the one below directly on a cached dataset.</p><pre>/**<br> * Returns the possible filter values for a dimension such as countries<br> */<br>public List&lt;Dimension&gt; getDimensions(long movieId, String dimensionId) {<br>    // Access in-memory Hollow feed with near instant query time<br>    Map&lt;String, List&lt;Dimension&gt;&gt; dimensions = dimensionsHollowConsumer.lookup(movieId);<br>    return dimensions.getOrDefault(dimensionId, List.of()).stream()<br>        .sorted(Comparator.comparing(Dimension::getName))<br>        .toList();<br>}</pre><p>Although it adds complexity to our service by requiring more intricate request routing and a higher memory footprint, pre-computed aggregates have given us greater stability and performance. In the case of fetching distinct dimensions, we’ve observed query times drop from hundreds of milliseconds to just tens of milliseconds. More importantly, this shift has offloaded high concurrency demands from our Druid cluster, resulting in more consistent query performance. In addition to this use case, cached pre-computed aggregates also power features such as retrieving recently launched titles, accessing all-time asset metrics, and serving various pieces of title metadata.</p><h4>Tuning Druid</h4><p>Even with the efficiencies gained from HLL sketches and Hollow feeds, ensuring that our Druid cluster operates performantly has been an ongoing challenge. Fortunately, at Netflix, we are in the company of multiple <a href="https://www.apache.org/foundation/governance/pmcs">Apache Druid PMC members</a> like <a href="https://www.linkedin.com/in/maytasm/">Maytas Monsereenusorn</a> and <a href="https://www.linkedin.com/in/jessetuglu/">Jesse Tuğlu</a> who have helped us wring out every ounce of performance. Some of the key optimizations we’ve implemented include:</p><ul><li><strong>Increasing broker count relative to historical nodes:</strong> We aim for a broker-to-historical ratio close to the <a href="https://druid.apache.org/docs/latest/operations/basic-cluster-tuning/#number-of-brokers">recommended 1:15</a>, which helps improve query throughput.</li><li><strong>Tuning segment sizes:</strong> By targeting the <a href="https://druid.apache.org/docs/latest/operations/segment-optimization/">300–700 MB “sweet spot”</a> for segment sizes, primarily using the tuningConfig.targetRowsPerSegment parameter during ingestion — we ensure that each segment a single historical thread scans is not overly large.</li><li><strong>Leveraging Druid lookups for data enrichment:</strong> Since joins can be prohibitively expensive in Druid, we use <a href="https://druid.apache.org/docs/latest/querying/lookups/">lookups</a> at query time for any key column enrichment.</li><li><strong>Optimizing search predicates:</strong> We ensure that all search predicates operate on physical columns rather than virtual ones, creating necessary columns during ingestion with <a href="https://druid.apache.org/docs/latest/ingestion/ingestion-spec/#transforms">transformSpec.transforms</a>.</li><li><strong>Filtering and slimming data sources at ingest:</strong> By applying filters within <a href="https://druid.apache.org/docs/latest/ingestion/ingestion-spec/#filter">transformSpec.filter</a> and removing all unused columns in <a href="https://druid.apache.org/docs/latest/ingestion/ingestion-spec/#dimensionsspec">dimensionsSpec.dimensions</a>, we keep our data sources lean and improve the possibility of <a href="https://druid.apache.org/docs/latest/ingestion/rollup">higher rollup yield</a>.</li><li><strong>Use of multi-value dimensions:</strong> Exploiting the Druid <a href="https://druid.apache.org/docs/latest/querying/multi-value-dimensions/">multi-value dimension</a> feature was key to overcoming the “many-to-many” combinatorial quandary when integrating audience filtering and grouping functionality mentioned in the “An Evolving Architecture” section above.</li></ul><p>Together, these optimizations, combined with previous ones, have decreased our p99 Druid latencies by roughly 50%.</p><h4>Validation &amp; Rollout</h4><p>Rolling out these changes to our metrics system required a thorough validation and release strategy. Our approach prioritized both data integrity and user trust, leveraging a blend of automation, targeted tooling, and incremental exposure to production traffic. At the core of our strategy was a parallel stack deployment: both the legacy and new metric stacks operated side-by-side within the Muse Data microservice. This setup allowed us to validate data quality, monitor real-world performance, and mitigate risk by enabling seamless fallback at any stage.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/831/1*gmcpkg9ehK1iX4Z6tKsetw.png"></figure><p>We adopted a two-pronged validation process:</p><ul><li><strong>Automated Offline Validation: </strong>Using Jupyter Notebooks, we automated the sampling and comparison of key metrics across both the legacy and new stacks. Our sampling set included a representative mix: recently accessed titles, high-profile launches, and edge-case titles with unique handling requirements. This allowed us to catch subtle discrepancies in metrics early in the process. Iterative testing on this set guided fixes, such as tuning the HLL logK parameter and benchmarking end-to-end latency improvements.</li><li><strong>In-App Data Comparison Tooling: </strong>To facilitate rapid triage, we built a developer-facing comparison feature within our application that displays data from both the legacy and new metric stacks side by side. The tool automatically highlights any significant differences, making it easy to quickly spot and investigate discrepancies identified during offline validation or reported by users.</li></ul><p>We implemented several release best practices to mitigate risk and maintain stability:</p><ul><li><strong>Staggered Implementation by Application Segment: </strong>We developed and deployed the new metric stack in stages, focusing on specific application segments. This meant building out support for asset types like artwork and video separately and then further dividing by CEE phase (Explore, Exploit). By implementing changes segment by segment, we were able to isolate issues early, validate each piece independently, and reduce overall risk during the migration.</li><li><strong>Shadow Testing (“Dark Launch”):</strong> Prior to exposing the new stack to end users, we mirrored production traffic asynchronously to the new implementation. This allowed us to validate real-world latency and catch potential faults in a live environment, without impacting the actual user experience.</li><li><strong>Granular Feature Flagging: </strong>We implemented fine-grained feature flags to control exposure within each segment. This allowed us to target specific user groups or titles and instantly roll back or adjust the rollout scope if any issues were detected, ensuring rapid mitigation with minimal disruption.</li></ul><h3>Learnings and Next Steps</h3><p>Our journey with Muse tested the limits of several parts of the stack: the ETL layer, the Druid layer, and the data serving layer. While some choices, like leveraging Netflix’s in-house Hollow infrastructure, were influenced by available resources, simple principles like offloading query volume, pre-filtering of rows and columns before Druid rollup, and optimizing search predicates (along with a bit of HLL magic) went a long way in allowing us to support new capabilities while maintaining performance. Additionally, engineering best practices like producing side-by-side implementations and backwards-compatible changes enabled us to roll out revisions steadily while maintaining rigorous validation standards. Looking ahead, we’ll continue to build on this foundation by supporting a wider range of content types like Live and Games, incorporating synopsis data, deepening our understanding of how assets work together to influence member choosing, and incorporating new metrics to distinguish between “effective” and “authentic” promotional assets, in service of helping members find content that truly resonates with them.</p><p>¹ Apache Spark is an open-source analytics engine for processing large-scale data, enabling tasks like batch processing, machine learning, and stream processing.</p><p>² Apache Druid is a high-performance, real-time analytics database designed for quickly querying large volumes of data.</p><p>³ Hollow is a Java library for efficient in-memory storage and access to moderately sized, read-only datasets, making it ideal for high-performance data retrieval.</p><p>⁴ Apache Iceberg is an open-source table format designed for large-scale analytical datasets stored in data lakes. It provides a robust and reliable way to manage data in formats like Parquet or ORC within cloud object storage or distributed file systems.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=aa9ad326fd77" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/scaling-muse-how-netflix-powers-data-driven-creative-insights-at-trillion-row-scale-aa9ad326fd77">Scaling Muse: How Netflix Powers Data-Driven Creative Insights at Trillion-Row Scale</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/scaling-muse-how-netflix-powers-data-driven-creative-insights-at-trillion-row-scale-aa9ad326fd77</link>
      <guid>https://netflixtechblog.com/scaling-muse-how-netflix-powers-data-driven-creative-insights-at-trillion-row-scale-aa9ad326fd77</guid>
      <pubDate>Mon, 22 Sep 2025 23:24:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Empowering Netflix Engineers with Incident Management]]></title>
      <description><![CDATA[<p><em>By: </em><a href="https://www.linkedin.com/in/mollystruve/"><em>Molly Struve</em></a></p><p>Netflix’s mission to provide seamless entertainment to hundreds of millions of users globally demands exceptional reliability. At the heart of this reliability is how we handle incidents — those inevitable moments when something doesn’t go as expected.</p><p>Teams can respond quickly and more effectively when incidents are managed consistently across a company. A robust process for following up after incidents creates opportunities for learning and improving systems. This continuous improvement cycle is essential for maintaining the highly reliable systems on which our members depend.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*xiHZ0qDpJbANPSw1UwqLJA.png"></figure><p>Having a shared, consistent approach to incident management became critical as Netflix grew and expanded its business. This post delves into our journey to transform incident management from a centralized function into a widespread, accessible practice and the hard-won lessons we’ve learned along the way.</p><h3>The Past: Countless Missed Opportunities</h3><p>For most of Netflix’s past, incident management was the domain of our central Site Reliability Engineering team, called <a href="https://netflixtechblog.com/keeping-customers-streaming-the-centralized-site-reliability-practice-at-netflix-205cc37aa9fb">CORE</a> (Critical Operations and Reliability Engineering). CORE was focused on streaming and was the sole initiator of incidents. They used Jira and a single Slack channel for incident response. This approach worked in the early days, but we knew it wouldn’t scale as Netflix grew and diversified.</p><p>With thousands of microservices supporting critical functions beyond streaming, we knew plenty of things were breaking that we were not capturing. We had an internal post-incident write-up template called “OOPS,” which teams could use to write about operational surprises. The template saw limited adoption as many engineers didn’t know about it or understand its purpose or value. With countless smaller, everyday incidents going unnoticed, we were missing key opportunities to learn and improve.</p><h3>The Aspiration: A Paved Road to Incident Management</h3><p>Recognizing these limits, we embarked on a journey to democratize incident management. Our goal: open more incidents and engage more teams in the process. We envisioned a “paved road” for incident management — a process so intuitive and streamlined that anyone could easily declare and manage an incident, even at 3 AM. Creating a paved road required a shift: our central SRE team would no longer be the only ones declaring incidents. Instead, we’d empower teams across engineering to own their own incidents. Making this significant shift required both technological and cultural changes.</p><h3>Finding the Right Tool</h3><p>Scaling technical processes within an organization as diverse and intricate as Netflix is challenging. To enable every engineering team to manage incidents effectively, we needed a comprehensive incident management tool that was far more sophisticated than Jira and a single Slack channel. We knew any solution, whether built or bought, would need to meet four key requirements:</p><ul><li><strong>Intuitive user experience</strong> — Our number one priority was making sure the tool was so intuitive that anyone could use it with little to no training.</li><li><strong>Internal data integration capabilities</strong> — We needed the ability to hook in Netflix-specific data.</li><li><strong>Balanced customization with consistency</strong> — We wanted teams to have flexibility while maintaining shared standards.</li><li><strong>Approachable</strong> — A friendly and appealing tool that could help drive a cultural shift around incidents.</li></ul><p>The “build vs. buy” question was a significant consideration. While Netflix boasts a world-class engineering team, building an in-house solution meeting these requirements was impractical due to our ambitious timeline, the substantial investment needed, and ongoing ownership costs. Following Netflix’s engineering principle of “build only when necessary,” we evaluated external solutions against these criteria.</p><p>This evaluation process led us to adopt <a href="http://incident.io/">Incident.io</a>. While the platform checked all our boxes during selection, the four above requirements proved even more impactful than anticipated during Netflix’s incident management transformation.</p><h3>Tackling the Transformation</h3><p>Selecting the right tool was just the beginning. The real challenge was rolling it out across Netflix’s diverse engineering organization and achieving the cultural shift we envisioned. Here are four elements that helped make our goal a reality.</p><h4>Intuitive Design Drove Adoption and Cultural Transformation</h4><p>Tool usability was crucial to encourage teams to open incidents. It had to be easily understandable, even for engineers who aren’t incident management experts and only use it a few times a year. When introducing <a href="http://incident.io/">Incident.io</a>, we saw rapid organic adoption because the tool was easy to pick up without much guidance. Its intuitive design allowed users to discover features as they used it. Thanks to prioritizing usability, within four months, 20% of engineering teams were using the tooling, and six months later, we had over 50% adoption.</p><p>Beyond rapid adoption, the tool helped shift how Netflix engineers think about incidents. Incidents went from “big scary outages” to simply “any blip or issue that degrades or disrupts a service that deserves attention and learning.” The tool’s friendly, welcoming interface made incident management less intimidating and more accessible. Some engineers described the platform as “jolly” and mentioned that it actually made them <em>want</em> to open incidents. The approachable design lowered psychological barriers for engineers to declare incidents and made it feel like a natural, even positive, part of their workflow.</p><h4>Organizational Investment Supported Scalable Growth</h4><p>While having an intuitive tool was important, successfully empowering engineers to open incidents required deliberate organizational investment. We invested heavily in standardization, developing an incident management process lightweight enough to avoid overwhelming users yet structured enough to support complex incidents. Finding the right balance took time and active engagement with users to understand what was working and what wasn’t. To this day, we still make adjustments to refine and improve the process.</p><p>On the education front, we created lightweight docs, quick-reference cheatsheets, and short demo videos to accelerate adoption across Netflix’s diverse engineering organization. We took these resources on roadshows across engineering teams and proved that the barrier to entry for managing incidents was practically nonexistent. While most engineers bought in easily, we had our skeptics. Over time, we worked with these folks to understand their needs better and help them fit incident management into their daily routines and processes.</p><h4>Internal Integrations Reduce Cognitive Load</h4><p>Integrating our unique organizational context — like teams, software services, business domains, and even hardware devices — directly into the incident management platform was critical. Netflix-specific contextualization enables powerful automations, such as automatically looping in the right teams or pre-filling incident fields from alerts. These integrations significantly reduce cognitive load during an incident and empower engineers to focus on quick mitigation. Beyond individual incidents, integrations with internal data across multiple incidents enable us to identify and address systemic issues.</p><h4>Balanced Customization with Consistency Improved Response</h4><p>A flexible platform allowed us to create a tailored incident response experience while enforcing a shared language and standard metadata across all engineering teams. This balance proved crucial for response effectiveness: different teams can adapt workflows to their specific needs, but core elements like “impacted areas and domains” stay consistent. Incident responders can quickly understand any incident organization-wide because the structure and language remain familiar, enabling faster, more effective incident response.</p><h3>The Result: A New Era of Incident Management</h3><p>Our journey to democratize incident management has yielded massive wins across Netflix Engineering. We successfully transitioned from a centralized incident response model to empowering engineers to declare and manage incidents. The transformation has fostered a culture of renewed ownership and learning across engineering teams.</p><p>We’ve established new practices and are growing an incident management culture we’re genuinely proud of, but we’re not done yet. Our incident management processes continue to evolve and adapt to fit Netflix’s growing needs. Every day, we work to educate engineers and leaders on the tremendous value incidents provide. We’re excited to continue harnessing these incredible learning opportunities to improve our platform for our hundreds of millions of members.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=ebb967871de4" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/empowering-netflix-engineers-with-incident-management-ebb967871de4">Empowering Netflix Engineers with Incident Management</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/empowering-netflix-engineers-with-incident-management-ebb967871de4</link>
      <guid>https://netflixtechblog.com/empowering-netflix-engineers-with-incident-management-ebb967871de4</guid>
      <pubDate>Fri, 19 Sep 2025 18:48:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[From Facts & Metrics to Media Machine Learning: Evolving the Data Engineering Function at Netflix]]></title>
      <description><![CDATA[<p>By<em> </em><a href="https://www.linkedin.com/in/daomi/">Dao Mi</a>, <a href="https://www.linkedin.com/in/pabloadelgado/">Pablo Delgado</a>, <a href="https://www.linkedin.com/in/ryan-berti-4942aa83/">Ryan Berti</a>, <a href="https://www.linkedin.com/in/amanuel-kahsay-81ab29153/">Amanuel Kahsay</a>, <a href="https://www.linkedin.com/in/onwoke/">Obi-Ike Nwoke</a>, <a href="https://www.linkedin.com/in/chris-thrailkill-a268914/">Christopher Thrailkill</a>, and <a href="https://www.linkedin.com/in/patriciogarza/">Patricio Garza</a></p><p>At Netflix, data engineering has always been a critical function to enable the business’s ability to understand content, power recommendations, and drive business decisions. Traditionally, the function centered on building robust tables and pipelines to capture facts, derive metrics, and provide well modeled data products to their partners in analytics &amp; data science functions. But as Netflix’s studio and content production scaled, so too have the challenges — and opportunities — of working with complex media data.</p><p>Today, we’re excited to share how our team is formalizing a new specialization of data engineering at Netflix: <strong>Media ML Data Engineering</strong>. This evolution is embodied in our latest collaboration with our platform teams, the <strong>Media Data Lake</strong>, which is designed to harness the full potential of media assets (video, audio, subtitles, scripts, and more) and enable the latest advances in machine learning, including latest transformer model architecture. As part of this initiative, we’re intentionally applying data engineering best practices — ensuring that our approach is both innovative and grounded in proven methodologies.</p><h3>The Evolution: From Traditional Tables to Media Tables</h3><p><strong>Traditional data engineering</strong> at Netflix focused on building structured tables for metrics, dashboards, and data science models. These tables were primarily structured text or numerical fields, ideal for business intelligence, analytics and statistical modeling.</p><p>However, the nature of media data is fundamentally different:</p><ul><li>It’s <strong>multi-modal</strong> (video, audio, text, images).</li><li>It contains <strong>derived</strong> fields from media (embeddings, captions, transcriptions…etc)</li><li>It’s <strong>unstructured</strong> and massive in scale when parsed out.</li><li>It’s deeply <strong>intertwined</strong> with creative workflows and business asset lineage.</li></ul><p>As our studio operations (see below) expanded, we saw the need for a new approach — one that could provide centralized, standardized, and scalable access to all types of media assets and their metadata for both analytical and machine learning workflows.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*87cD7-YKwcl_Quvu"></figure><h3>The Rise of Media ML Data Engineering</h3><p>Enter <strong>Media ML Data Engineering</strong> — a new specialization at Netflix that bridges the gap between traditional data engineering and the unique demands of media-centric machine learning. This role sits at the intersection of data engineering, ML infrastructure, and media production. Our mission is to provide seamless access to media assets and derived data (including outputs from machine learning models) for researchers, data scientists, and other downstream data consumers.</p><h3>Key Responsibilities</h3><ul><li><strong>Centralized Media Data Access:</strong> Building, cataloging and maintaining the data and pipelines that populates the Media Data Lake, a data platform for storing and serving media assets and their metadata.</li><li><strong>Asset Standardization:</strong> Standardizing media assets across modalities (video, images, audio, text) to ensure consistency and quality for ML applications in partnership with domain engineering teams.</li><li><strong>Metadata Management:</strong> Unifying and enriching asset metadata, making it easier to track asset lineage, quality, and coverage.</li><li><strong>ML-Ready Data:</strong> Exposing large corpora of assets for early-stage algorithm exploration, benchmarking, and productionization.</li><li><strong>Collaboration:</strong> Partnering closely with domain experts, algorithm researchers, upstream content engineering teams and (machine learning &amp; data) platform colleagues to ensure our data meets real-world needs.</li></ul><p>This new role is essential for bridging the gap between creative media workflows and the technical demands of cutting-edge ML.</p><h3>Introducing the Media Data Lake</h3><p>To enable the next generation of media analytics and machine learning, we are building the <strong>Media Data Lake </strong>at Netflix — a data lake designed specifically for media assets at Netflix using <a href="https://lancedb.com/">LanceDB</a>. We have partnered with our data platform team on integrating LanceDB into our <a href="https://netflixtechblog.com/all?topic=big-data">Big Data Platform</a>.</p><h3>Architecture and Key Components</h3><ul><li><strong>Media Table:</strong> The core of the Media Data Lake, this structured dataset captures essential metadata and references to all media assets. It’s designed to be extensible, supporting both traditional metadata and outputs from ML models (including transformer-based embeddings, media understanding research and more).</li><li><strong>Data Model:</strong> We are developing a robust data model to standardize how media assets and their attributes are represented, making it easier to query and join across schemas.</li><li><strong>Data API:</strong> An pythonic interface that will provide programmatic access to the Media Table, supporting both interactive exploration and automated workflows.</li><li><strong>UI Components:</strong> Off-the-shelf UI interfaces enable teams to visually explore assets in the media data lake, accelerating discovery and iteration for ICs.</li><li><strong>Online and Offline System Architecture:</strong> Real-time access for lightweight queries and exploration of raw media assets; scalable large batch processing for ML training, benchmarking, and research.</li><li><strong>Compute</strong>: distributed batch inference layer capable of processing using GPUs and media data processing at scale using CPUs.</li></ul><h3>Starting Small with New Technology</h3><p>Our initial focus this past year has been on delivering a “data pond” — a mini-version of the Media Data Lake targeted at video/audio datasets for early stage model training, evaluation and research. All data for this phase comes from AMP, our internal <a href="https://netflixtechblog.com/elasticsearch-indexing-strategy-in-asset-management-platform-amp-99332231e541">asset management system</a> and <a href="https://netflixtechblog.com/scalable-annotation-service-marken-f5ba9266d428">annotation store</a>, and the scope is intentionally small to ensure a solid, extensible foundation could be built while introducing a new technology into the company. We are able to perform data exploration of the raw media assets to build up an intuitive understanding of the media via lightweight queries to AMP.</p><h3>Media Tables: The New Foundation for ML and Innovation</h3><p>One of the most exciting developments is the rise of <strong>media tables</strong> — structured datasets that not only capture traditional metadata, but also include the outputs of advanced ML models.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*4AL_ScuaNy4ZguBj"></figure><p>These media tables power a range of innovative applications, such as:</p><ul><li><strong>Translation &amp; Audio Quality Measures:</strong> Managing audio clips and features via text-to-speech models for engineering localization quality metrics.</li><li><strong>Media Fidelity Restoration:</strong> Research on restoration of videos to HDR for remastering and other image technology use-cases.</li><li><strong>Story Understanding and Content Embedding:</strong> Structuring narrative elements extracted from textual evidence and video of a title to increase operational efficiency in title launch preparation and ratings, e.g. detection of smoking, gore, NSFW scenes in our titles.</li><li><strong>Media Search:</strong> Leverage multi-modal vector search to find similar keyframes, shots, dialogue to facilitate research and experimentation.</li></ul><p>These tables built on top of LanceDB are designed to scale, support complex queries, and serve both research and other data science &amp; analytical needs.</p><h3>The Human Side: New Roles and Collaboration</h3><p>Media ML Data Engineering is a team sport. Our data engineers partner with domain experts, data scientists, ML researchers, upstream business ops and content engineering teams to ensure our data solutions are fit for purpose. We also work closely with our friendly platform teams to ensure technological breakthroughs that are beneficial beyond our small corner of the universe could become horizontal abstractions that benefit the rest of Netflix. This collaborative model enables rapid iteration, high data quality, innovative use cases and technology re-use.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*adTotLaJXhE7KC2-"></figure><h3>Looking Ahead</h3><p>The evolution from traditional data engineering to Media ML data engineering — anchored by our media data lake — is unlocking new frontiers for Netflix:</p><ul><li><strong>Richer, more accurate ML models</strong> trained on high-quality, standardized media data.</li><li><strong>Supercharge ML Model evaluations </strong>via quick iteration cycles on the data.</li><li><strong>Faster experimentation and productization</strong> of new AI-powered features.</li><li><strong>Deeper insights into our content and creative workflows</strong> via metrics constructed from Media ML algorithms inferred features.</li></ul><p>As we continue to grow the media data lake, be on the lookout for subsequent blog posts sharing our learnings and tools with the broader media ml &amp; data engineering community.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=6dcc91058d8d" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/from-facts-metrics-to-media-machine-learning-evolving-the-data-engineering-function-at-netflix-6dcc91058d8d">From Facts &amp; Metrics to Media Machine Learning: Evolving the Data Engineering Function at Netflix</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/from-facts-metrics-to-media-machine-learning-evolving-the-data-engineering-function-at-netflix-6dcc91058d8d</link>
      <guid>https://netflixtechblog.com/from-facts-metrics-to-media-machine-learning-evolving-the-data-engineering-function-at-netflix-6dcc91058d8d</guid>
      <pubDate>Thu, 21 Aug 2025 19:39:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[ML Observability: Bringing Transparency to Payments and Beyond]]></title>
      <description><![CDATA[<p>By <a href="http://www.linkedin.com/in/tang">Tanya Tang</a>, <a href="https://www.linkedin.com/in/dkmehrmann">Andrew Mehrmann</a></p><p>At Netflix, the importance of ML observability cannot be overstated. ML observability refers to the ability to monitor, understand, and gain insights into the performance and behavior of machine learning models in production. It involves tracking key metrics, detecting anomalies, diagnosing issues, and ensuring models are operating reliably and as intended. ML observability helps teams identify data drift, model degradation, and operational problems, enabling faster troubleshooting and continuous improvement of ML systems.</p><p>One specific area where ML observability plays a crucial role is in payment processing. At Netflix, we strive to ensure that technical or process-related payment issues never become a barrier for someone wanting to sign up or continue using our service. By leveraging ML to optimize payment processing, and using ML observability to monitor and explain these decisions, we can reduce payment friction. This ensures that new members can subscribe seamlessly and existing members can renew without hassle, allowing everyone to enjoy Netflix without interruption.</p><h3>ML Observability: A Primer</h3><p>ML Observability is a set of practices and tools to help ML practitioners and stakeholders alike gain a deeper, end to end understanding of their ML systems across all stages of its lifecycle, from development to deployment to ongoing operations. An effective ML Observability framework not only facilitates automatic detection and surfacing of issues but also provides detailed root cause analysis, acting as a guardrail to ensure ML systems perform reliably over time. This enables teams to iterate and improve their models rapidly, reduce time to detection for incidents, while also increasing the buy-in and trust of their stakeholders by providing rich context about the system’s’ behaviors and impact.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*nIX4fj34CFNPh1TCXWADSw.png"></figure><p>Some examples of these tools include long-term production performance monitoring and analysis, feature, target, and prediction drift monitoring, automated data quality checks, and model explainability. For example, a good observability system would detect aberrations in input data, the feature pipeline, predictions, and outcomes as well provide insight into the likely causes of model decisions and/or performance.</p><p>As an ML portfolio grows, ad-hoc monitoring becomes increasingly challenging. Greater complexity also raises the likelihood of interactions between different model components, making it unrealistic to treat each model as an isolated black box. At this stage, investing strategically in observability is essential — not only to support the current portfolio, but also to prepare for future growth.</p><h3>Stakeholder-Facing Observability Modules</h3><p>In order to reliably and responsibly evolve our payments processing system to be increasingly ML-driven, we invested heavily up-front in ML observability solutions. To provide confidence to our business stakeholders through this evolution, we looked beyond technical metrics such as precision and recall and placed greater emphasis on real-world outcomes like “how much traffic did we send down this route” and “where are the regions that ML is underperforming.”</p><p>Using this as a guidepost, we designed a collection of interconnected modules for machine learning observability: <strong>logging, monitoring, and explaining.</strong></p><figure><img alt="" src="https://cdn-images-1.medium.com/max/374/0*Q0lbG7mF6SCfH4kM"></figure><h3>Logging</h3><p>In order to support the monitoring and explaining we wanted to do, we first needed to log the appropriate data. This seems obvious and trivial, but as usual the devil is in the details: what fields exactly do we need to log and when? How does this work for simple models vs. more complex ones? What about models that are actually made of multiple models?</p><p>Consider the following, relatively straightforward model. It takes some input data, creates features, passes these to a model which creates some score between 0 and 1, and then that score is translated into a decision (say, whether to process a card as Debit or Credit).</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*eqnhrCw9fSO4t01l42GuhQ.png"></figure><p>There are several elements you may wish to log: a unique identifier for each record that is trained and scored, the raw data, the final features that fed the model, a unique identifier for the model, the feature importances for that model, the raw model score, the cutoffs used to map a score to a decision, timestamps for the decision as well as the model, etc.</p><p>To address this, we drafted an initial data schema that would enable our various ML observability initiatives. We identified the following logical entities to be necessary for the observability initiatives we were pursuing:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*QnB8neT1CZFqkceY0eVT4g.png"></figure><h3>Monitoring</h3><p>The goal of monitoring is twofold: 1) enable self-serve analytics and 2) provide opinionated views into key model insights. It can be helpful to think of this as “business analytics on your models,” as a lot of the key concepts (online analytical processing cubes, visualizations, metric definitions, etc) carry over. Following this analogy, we can craft key metrics that help us understand our models. There are several considerations when defining metrics, including whether your needs are understanding real-world model behavior versus offline model metrics, and whether your audience are ML practitioners or model stakeholders.</p><p>Due to our particular needs, our bias for metrics is toward online, stakeholder-focused metrics. Online metrics tell us what actually happened in the real world, rather than in an idealized counterfactual universe that might have its own biases. Additionally, our stakeholders’ focus is on business outcomes, so our metrics tend to be outcome-focused rather than abstract and technical model metrics.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/916/0*rXEcgYja4RCzbOHm"></figure><p>We focused on simple, easy to explain metrics:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/836/1*gcysTqQSi5Eh4a8UXNGIHQ.png"></figure><p>These metrics begin to suggest reasons for changing trends in the model’s behavior over time, as well as more generally how the model is performing. This gives us an overall view of model health and an intuitive approximation of what we think the model should have done. For example, if payment processor A starts receiving more traffic in a certain market compared to payment processor B, you might ask:</p><ul><li>Has the model seen something to make it prefer processor A?</li><li>Have more transactions become eligible to go to processor A?</li></ul><p>However, to truly explain specific decisions made by the model, especially which features are responsible for current trends, we need to use more advanced explainability tools, which will be discussed in the next section.</p><h3>Explaining</h3><p>Explainability means understanding the “why” behind ML decisions. This can mean the “why” in aggregate (e.g. why are so many of our transactions suddenly going down one particular route) or the “why” for a single instance (e.g. what factors led to this particular transaction being routed a particular way). This gives us the ability to approximate the previous status quo where we could inspect our static rules for insights about route volume.</p><p>One of the most effective tools we can leverage for ML explainability is SHAP (Shapley Additive exPlanations, <em>Lundberg &amp; Lee 2017</em>). At a high level, SHAP values are derived from cooperative game theory, specifically the Shapley values concept. The core idea is to fairly distribute the “payout” (in our case, the model’s prediction) among the “players” (the input features) by considering their contribution to the prediction.</p><h4>Key Benefits of SHAP:</h4><ol><li><strong>Model-Agnostic:</strong> SHAP can be applied to any ML model, making it a versatile tool for explainability.</li><li><strong>Consistent and Local Explanations:</strong> SHAP provides consistent explanations for individual predictions, helping us understand the contribution of each feature to a specific decision.</li><li><strong>Global Interpretability:</strong> By aggregating SHAP values across many predictions, we can gain insights into the overall behavior of the model and the importance of different features.</li><li><strong>Mathematical properties: </strong>SHAP satisfies important mathematical axioms such as efficiency, symmetry, dummy, and additivity. These properties allow us to compute explanations at the individual level and aggregate them for any ad-hoc groups that stakeholders are interested in, such as country, issuing bank, processor, or any combinations thereof.</li></ol><p>Because of the above advantages, we leverage SHAP as one core algorithm to unpack a variety of models and open the black box for stakeholders. Its well-documented <a href="https://shap.readthedocs.io/en/latest/example_notebooks/overviews/An%20introduction%20to%20explainable%20AI%20with%20Shapley%20values.html">Python interface</a> makes it easy to integrate into our workflows.</p><h4>Explaining Complex ML Systems</h4><p>For ML systems that score single events and use the output scores directly for business decisions, explainability is relatively straightforward, as the production decision is directly tied to the model’s output. However, in the case of a bandit algorithm, explainability can be more complex because the bandit policy may involve multiple layers, meaning the model’s output may not be the final decision used in production. For example, we may have a classifier model to predict the likelihood of transaction approval for each route, but we might want to penalize certain routes due to higher processing fees.</p><p>Here is an example of a plot we built to visualize these layers. The traffic that the model would have selected on its own is on the left, and different penalty or guardrail layers impact final volume as you move left to right. For example, the model originally allocated 22% traffic to processor W with Configuration A, however for cost and contractual considerations, the traffic was reduced to 19% with 3% being allocated to Processor W with Configuration B, and Processor Nc with Configuration B.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1000/0*F5saiA3Wzn7Gv39r"></figure><p>While individual event analysis is crucial, such as in fraud detection where false positives need to be scrutinized, in payment processing, stakeholders are more interested in explaining model decisions at a group level (e.g., decisions for one issuing bank from a country). This is essential for business conversations with external parties. SHAP’s mathematical properties allow for flexible aggregation at the group level while maintaining consistency and accuracy.</p><p>Additionally, due to the multi-candidate structure, when stakeholders inquire about why a particular candidate was chosen, they are often interested in the differential perspective — specifically, why another similar candidate was not selected. We leverage SHAP to segment populations into cohorts that share the same candidates and identify the features that make subtle but critical differences. For example, while Feature A might be globally important, if we compare two candidates that both have the same value for Feature A, the local differences become crucial. This facilitates stakeholders discussions and helps understand subtle differences among different routes or payment partners.</p><p>Earlier, we were alerted that our ML model consistently reduced traffic to a particular route every Tuesday. By leveraging our explanation system, we identified that two features — <em>route</em> and <em>day of the week</em> — were contributing negatively to the predictions on Tuesdays. Further analysis revealed that this route had experienced an outage on a previous Tuesday, which the model had learned and encoded into the <em>route</em> and <em>day of the week</em> features. This raises an important question: should outage data be included in model training? This discovery opens up discussions with stakeholders and provides opportunities to further enhance our ML system.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*D3sMvWaKsWvoFmSt"></figure><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*pgodI89ibzBtgySJ"></figure><p>The explanation system not only demystifies our machine learning models but also fosters transparency and trust among our stakeholders, enabling more informed and confident decision-making.</p><h3>Wrapping Up</h3><p>At Netflix, we face the challenge of routing thousands of payment transactions per minute in our mission to entertain the world. To help meet this challenge, we introduced an observability framework and set of tools to allow us to open the ML black box and understand the intricacies of how we route billions of dollars of transactions in hundreds of countries every year. This has led to a massive operational complexity reduction in addition to improved transaction approval rate, while also allowing us to focus on innovation rather than operations.</p><p>Looking ahead, we are generalizing our solution with a standardized data schema. This will simplify applying our advanced ML observability tools to other models across various domains. By creating a versatile and scalable framework, we aim to empower ML developers to quickly deploy and improve models, bring transparency to stakeholders, and accelerate innovation.</p><p><em>We also thank </em><a href="https://www.linkedin.com/in/ckarthiksriram/"><em>Karthik Chandrashekar</em></a><em>, </em><a href="https://www.linkedin.com/in/zainabmir/"><em>Zainab Mahar Mir</em></a><em>, </em><a href="https://www.linkedin.com/in/joshuakaroly/"><em>Josh Karoly</em></a><em> and </em><a href="https://www.linkedin.com/in/joshkornpublicpolicy/"><em>Josh Korn</em></a><em> for their helpful suggestions.</em></p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=33073e260a38" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/ml-observability-bring-transparency-to-payments-and-beyond-33073e260a38">ML Observability: Bringing Transparency to Payments and Beyond</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/ml-observability-bring-transparency-to-payments-and-beyond-33073e260a38</link>
      <guid>https://netflixtechblog.com/ml-observability-bring-transparency-to-payments-and-beyond-33073e260a38</guid>
      <pubDate>Mon, 18 Aug 2025 20:15:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Accelerating Video Quality Control at Netflix with Pixel Error Detection]]></title>
      <description><![CDATA[<p><em>By </em><a href="https://www.isikdogan.com/"><em>Leo Isikdogan</em></a><em>, Jesse Korosi, Zile Liao, Nagendra Kamath, Ananya Poddar</em></p><p>At Netflix, we support the filmmaking process that merges creativity with technology. This includes reducing manual workloads wherever possible. Automating tedious tasks that take a lot of time while requiring very little creativity allows our creative partners to devote their time and energy to what matters most: creative storytelling.</p><p>With that in mind, we developed a new method for quality control (QC) that automatically detects pixel-level artifacts in videos, reducing the need for manual visual reviews in the early stages of QC.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/798/1*VN8kAKM6Hr1vYUT7itlauA.gif"><figcaption><em>Examples of detected pixel errors.</em></figcaption></figure><h3>Why This Matters</h3><p>Netflix is deeply invested in ensuring our content creators’ stories are accurately carried from production to screen. As such, we invest manual time and energy in reviewing for technical errors that could distract from our members’ immersion in and enjoyment of these stories.</p><p>Teams spend a lot of time manually reviewing every shot to identify any issues that could cause problems down the line. One of the problems they look for is tiny bright spots caused by malfunctioning camera sensors (often called hot or lit pixels). Flagging those issues is a painstaking and error-prone process. They can be hard to catch even when every single frame in a shot is manually inspected. And if left undetected, they can surface unexpectedly later in production, leading to labor-intensive and costly fixes.</p><p>By automating these QC checks, we help production teams spot and address issues sooner, reduce tedious manual searches, and address issues before they accumulate.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*wgfs1OZHkGnf91WO"><figcaption><em>Proof of Concept: hours spent on full-frame manual QC vs. minutes with the automated workflow.</em></figcaption></figure><h3>Precision at the Pixel Level: Pixel Error Detection</h3><p>Pixel errors come in two main types:</p><ol><li>Hot (lit) pixels: single frame bright pixels</li><li>Dead (stuck) pixels: pixels that don’t respond to light</li></ol><p>Earlier work at Netflix addressed detecting dead pixels using techniques based on pixel intensity gradients and statistical comparisons [<a href="https://patents.google.com/patent/US11107206B2/en">1</a>, <a href="https://www.researchgate.net/publication/326104534_Shot_Change_and_Stuck_Pixel_Detection_of_Digital_Video_Assets">2</a>]. In this work, we focus on hot pixels, which are a lot harder to flag manually.</p><p>Hot pixels in a frame can occupy only a few pixels and appear for just a single frame. Imagine reviewing thousands of high-resolution video frames looking for hot pixels. To reduce manual effort, we built a highly efficient neural network to pinpoint pixel-level artifacts in real time. While detection of hot pixels is not entirely new in video production workflows, we do it at scale and with near-perfect recall rates.</p><p>Detecting artifacts at the pixel level requires the ability to identify small-scale, fine features in large images. It also requires leveraging temporal information to distinguish between actual pixel artifacts and naturally bright pixels with artifact-like features, such as small lights, catch lights, and other specular reflections.</p><p>Given those requirements, we designed a bespoke model for this task. Many mainstream computer vision models downsample inputs to reduce dimensionality, but pixel errors are sensitive to this. For example, if we downsample a 4K frame by 8x to 480p resolution, pixel-level errors almost entirely disappear. For that reason, our model processes large-scale inputs at full resolution rather than explicitly downsampling them in pre-processing.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/865/1*LJSpRqkvB76OcIgP5yE6tw.png"><figcaption><em>Why Downsampling Fails. Top: Original frame crop showing a prominent hot pixel. Bottom: Same area after 8x downscaling, causing the artifact to become nearly invisible before it even reaches the model.</em></figcaption></figure><p>The network analyzes a window of five consecutive frames at a time, giving it the temporal context it needs to tell the difference between a one-off sensor glitch and a naturally bright object that persists across frames.</p><p>For every frame, the model outputs a continuous-valued map of pixel error occurrences at the input resolution. During training, we directly optimize those error maps by minimizing dense, pixel-wise loss functions.</p><p>During inference, our algorithm binarizes the model’s outputs using a confidence threshold, then performs connected component labeling to find clusters of pixel errors. Finally, it calculates the centroids of those clusters to report (x, y) locations of the found pixel errors.</p><p>All of this processing happens in real-time on a single GPU.</p><h3>Building a Synthetic Pixel Error Generator</h3><p>Pixel errors are rare and make up a very small portion of videos, both temporally and spatially, in the context of the total volume of footage captured and the full resolution of a given frame. Therefore, they are hard to annotate manually. Initially, we had virtually no data to train our model. To overcome this, we developed a synthetic pixel error generator that closely mimicked real-world artifacts. We simulated two main types of pixel errors: symmetrical and curvilinear.</p><p><strong><em>Symmetrical:</em></strong> Most pixel errors are symmetrical along at least one axis.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/809/1*Wjwzn4Xa6Wb_xUyEY-2ENw.png"><figcaption><em>Symmetrical Artifacts. Left: Three real examples of hot pixels. Right: Three synthetically generated hot pixel samples.</em></figcaption></figure><p><strong><em>Curvilinear:</em></strong> Some pixel errors follow curvilinear structures.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/814/1*CPgXIysLwFoyj7uwLLjeSw.png"><figcaption><em>Curvilinear Artifacts. Left: Three real examples of hot pixels. Right: Three synthetically generated hot pixel samples.</em></figcaption></figure><p>To create realistic training samples, we superimposed these synthetic errors onto frames from the Netflix catalog. We added those artificial hot pixels to where they would be most visible: dark, still areas in the scenes. Instead of sampling (x, y) coordinates for the synthetic errors uniformly, we sampled them from a heatmap, with selection probabilities determined by the amount of motion and image intensity.</p><p>Synthetic data was essential for training our initial model. However, to close the domain gap and improve precision, we needed to run multiple tuning cycles on fresh, real-world footage.</p><p>After training an initial model solely on this synthetic data, we refined it iteratively with real-world data as follows:</p><ol><li>Inference: Run the model on previously unseen footage without any added synthetic hot pixels.</li><li>False Positive Elimination: Manually review detections and zero out labels for false positives, which is easier than labeling hot pixels from scratch.</li><li>Fine-tuning and Iteration: Fine-tune on the refined dataset and repeat until convergence.</li></ol><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*HZ9nl1ReoUPbbZp0qIH__Q.png"><figcaption><em>Synthetic-to-Real Training Pipeline</em></figcaption></figure><p>While false positives represent a small percentage of total input volume, they can still constitute a meaningful number of alerts in absolute terms given the scale of content processing. We continue to refine our model and reduce false positives through ongoing application on real-world datasets. This synthetic-to-real refinement loop steadily reduces false alarms while preserving high sensitivity.</p><h3>Looking Ahead</h3><p>What once required hours of painstaking manual review can now potentially be completed in minutes, freeing creative teams to focus on what matters most: the art of storytelling. As we continue refining these capabilities through ongoing real-world deployment, we’re inspired by the many ways production teams can gain more time to build amazing stories for audiences around the world. We are also working with our partners to better understand how pixel errors affect the viewing experience, which will help us further optimize our models.</p><img src="https://medium.com/_/stat?event=post.clientViewed&amp;referrerSource=full_rss&amp;postId=47ef7af7ca2e" width="1" height="1" alt=""><hr><p><a href="https://netflixtechblog.com/accelerating-video-quality-control-at-netflix-with-pixel-error-detection-47ef7af7ca2e">Accelerating Video Quality Control at Netflix with Pixel Error Detection</a> was originally published in <a href="https://netflixtechblog.com/">Netflix TechBlog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></description>
      <link>https://netflixtechblog.com/accelerating-video-quality-control-at-netflix-with-pixel-error-detection-47ef7af7ca2e</link>
      <guid>https://netflixtechblog.com/accelerating-video-quality-control-at-netflix-with-pixel-error-detection-47ef7af7ca2e</guid>
      <pubDate>Mon, 11 Aug 2025 23:29:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Behind the Streams: Live at Netflix. Part 1]]></title>
      <description><![CDATA[<div class="ac cb"><div class="ci bh hv hw hx hy"><div><div></div><p id="de94" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">By <a class="ag hb" href="https://www.linkedin.com/in/sfedov/" rel="noopener ugc nofollow" target="_blank">Sergey Fedorov</a>, <a class="ag hb" href="https://www.linkedin.com/in/phamchristopher/" rel="noopener ugc nofollow" target="_blank">Chris Pham</a>, <a class="ag hb" href="https://www.linkedin.com/in/flavioribeiro/" rel="noopener ugc nofollow" target="_blank">Flavio Ribeiro</a>, <a class="ag hb" href="https://www.linkedin.com/in/chrisnewton2/" rel="noopener ugc nofollow" target="_blank">Chris Newton</a>, and <a class="ag hb" href="https://www.linkedin.com/in/wei-wei-1571794/" rel="noopener ugc nofollow" target="_blank">Wei Wei</a></p><p id="13a1" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">Many great ideas at Netflix begin with a question, and three years ago, we asked one of our boldest yet: if we were to entertain the world through Live — a format almost as old as television itself — how would <em class="oq">we</em> do it?</p><p id="0e1c" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">What began with an engineering plan to pave the path towards our first Live comedy special, <a class="ag hb" href="https://www.netflix.com/title/80167499" rel="noopener ugc nofollow" target="_blank">Chris Rock: Selective Outrage</a>, has since led to hundreds of Live events ranging from the biggest <a class="ag hb" href="https://www.netflix.com/tudum/articles/greatest-roast-of-all-time-tom-brady-live" rel="noopener ugc nofollow" target="_blank">comedy shows</a> and <a class="ag hb" href="https://about.netflix.com/en/news/nfl-christmas-day-games-on-netflix-average-over-30-million-global-viewers" rel="noopener ugc nofollow" target="_blank">NFL Christmas Games</a> to record-breaking <a class="ag hb" href="https://about.netflix.com/en/news/jake-paul-vs-mike-tyson-over-108-million-live-global-viewers" rel="noopener ugc nofollow" target="_blank">boxing fights</a> and becoming the <a class="ag hb" href="https://about.netflix.com/en/news/netflix-to-become-new-home-of-wwe-raw-beginning-2025" rel="noopener ugc nofollow" target="_blank">home of WWE</a>.</p><p id="1d28" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">In our series <em class="oq">Behind the Streams</em> — where we take you through the technical journey of our biggest bets — we will do a multiple part deep-dive into the architecture of Live and what we learned while building it. Part one begins with the foundation we set for Live, and the critical decisions we made that influenced our approach.</p><h1 id="9094" class="or os io bf ot ou ov ow gl ox oy oz gn pa pb pc pd pe pf pg ph pi pj pk pl pm bk">| But First: What Makes Live Streaming Different?</h1><p id="5adf" class="pw-post-body-paragraph nv nw io nx b ny pn oa ob oc po oe of go pp oh oi gr pq ok ol gu pr on oo op hp bk">While Live as a television format is not new, the streaming experience we intended to build required capabilities we did not have at the time. Despite 15 years of on-demand streaming under our belt, Live introduced new considerations influencing architecture and technology choices:</p><figure class="pv pw px py pz qa ps pt paragraph-image"><div role="button" tabindex="0" class="qb qc fl qd bh qe"><div class="ps pt pu"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*o1u1pYFC7BDuJSrT%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*o1u1pYFC7BDuJSrT%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*o1u1pYFC7BDuJSrT%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*o1u1pYFC7BDuJSrT%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*o1u1pYFC7BDuJSrT%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*o1u1pYFC7BDuJSrT%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*o1u1pYFC7BDuJSrT%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*o1u1pYFC7BDuJSrT 640w, https://miro.medium.com/v2/resize:fit:720/0*o1u1pYFC7BDuJSrT 720w, https://miro.medium.com/v2/resize:fit:750/0*o1u1pYFC7BDuJSrT 750w, https://miro.medium.com/v2/resize:fit:786/0*o1u1pYFC7BDuJSrT 786w, https://miro.medium.com/v2/resize:fit:828/0*o1u1pYFC7BDuJSrT 828w, https://miro.medium.com/v2/resize:fit:1100/0*o1u1pYFC7BDuJSrT 1100w, https://miro.medium.com/v2/resize:fit:1400/0*o1u1pYFC7BDuJSrT 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh fw qf c" width="700" height="408" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="qg ff qh ps pt qi qj bf b bg ab du">References: 1. <a class="ag hb" rel="noopener ugc nofollow" href="https://netflixtechblog.com/distributing-content-to-open-connect-3e3e391d4dc9" target="_blank" data-discover="true">Content Pre-Positioning on Open Connect</a>, 2.<a class="ag hb" href="https://www.infoq.com/presentations/load-balancing-netflix/" rel="noopener ugc nofollow" target="_blank">Load-Balancing Netflix Traffic at Global Scale</a></figcaption></figure><p id="6b2b" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">This means that we had a lot to build in order to make Live work well on Netflix. That starts with making the right choices regarding the fundamentals of our Live Architecture.</p><h1 id="c6a9" class="or os io bf ot ou ov ow gl ox oy oz gn pa pb pc pd pe pf pg ph pi pj pk pl pm bk">| Key Pillars of Netflix Live Architecture</h1><p id="f409" class="pw-post-body-paragraph nv nw io nx b ny pn oa ob oc po oe of go pp oh oi gr pq ok ol gu pr on oo op hp bk">Our Live Technology needed to extend the same promise to members that we’ve made with on-demand streaming: <strong class="nx ip">great quality</strong> on as <strong class="nx ip">many devices</strong> as possible <strong class="nx ip">without interruptions</strong>. Live is one of many entertainment formats on Netflix, so we also needed to seamlessly blend Live events into the user experience, all while scaling to over 300 million global subscribers.</p><p id="ca8d" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">When we started, we had <strong class="nx ip">nine months</strong> until the first launch. While we needed to execute quickly, we also wanted to <strong class="nx ip">architect for future growth</strong> in both <strong class="nx ip">magnitude</strong> and <strong class="nx ip">multitude</strong> of events. As a key principle, we leveraged our unique position of building support for a single product — Netflix — and having control over the full Live lifecycle, from Production to Screen.</p><figure class="pv pw px py pz qa ps pt paragraph-image"><div role="button" tabindex="0" class="qb qc fl qd bh qe"><div class="ps pt qk"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*eMwIzmDUqURUESqu%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*eMwIzmDUqURUESqu%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*eMwIzmDUqURUESqu%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*eMwIzmDUqURUESqu%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*eMwIzmDUqURUESqu%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*eMwIzmDUqURUESqu%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*eMwIzmDUqURUESqu%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*eMwIzmDUqURUESqu 640w, https://miro.medium.com/v2/resize:fit:720/0*eMwIzmDUqURUESqu 720w, https://miro.medium.com/v2/resize:fit:750/0*eMwIzmDUqURUESqu 750w, https://miro.medium.com/v2/resize:fit:786/0*eMwIzmDUqURUESqu 786w, https://miro.medium.com/v2/resize:fit:828/0*eMwIzmDUqURUESqu 828w, https://miro.medium.com/v2/resize:fit:1100/0*eMwIzmDUqURUESqu 1100w, https://miro.medium.com/v2/resize:fit:1400/0*eMwIzmDUqURUESqu 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh fw qf c" width="700" height="261" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="acf2" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk"><strong class="nx ip">Dedicated Broadcast Facilities to Ingest Live Content from Production</strong></p><p id="9608" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">Live events can happen anywhere in the world, but not every location has Live facilities or great connectivity. To ensure secure and reliable live signal transport, we leverage distributed and highly connected broadcast operations centers, with specialized equipment for signal ingest and inspection, closed-captioning, graphics and advertisement management. We prioritized <strong class="nx ip">repeatability</strong>, conditioning engineering to launch live events consistently, reliably, and cost-effectively, leaning into <strong class="nx ip">automation</strong> wherever possible. As a result, we have been able to reduce the event-specific setup to the transmission between production and the Broadcast Operations Center, reusing the rest across events.</p><p id="f6b4" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk"><strong class="nx ip">Cloud-based Redundant Transcoding and Packaging Pipelines</strong></p><p id="92de" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">The feed received at the Broadcast Center contains a fully produced program, but still needs to be encoded and packaged for streaming on devices. We chose a Cloud-based approach to allow for <strong class="nx ip">dynamic scaling</strong>, <strong class="nx ip">flexibility</strong> in configuration, and <strong class="nx ip">ease of integration</strong> with our Digital Rights Management (DRM), content management, and content delivery services already deployed in the cloud. We leverage AWS MediaConnect and AWS MediaLive to acquire feeds in the cloud and transcode them into various video quality levels with bitrates tailored per show. We built a <strong class="nx ip">custom packager</strong> to better integrate with our delivery and playback systems. We also built a <strong class="nx ip">custom Live Origin</strong> to ensure strict read and write SLAs for Live segments.</p><p id="7d57" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk"><strong class="nx ip">Scaling Live Content Delivery to millions of viewers with Open Connect CDN</strong></p><p id="2393" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">In order for the produced media assets to be streamed, they need to be transferred from a few AWS locations, where Live Origin is deployed, to hundreds of millions of devices worldwide. We leverage Netflix’s CDN, <a class="ag hb" href="https://www.theverge.com/22787426/netflix-cdn-open-connect" rel="noopener ugc nofollow" target="_blank">Open Connect</a>, to scale Live asset delivery. Open Connect servers are placed close to the viewers at <strong class="nx ip">over 6K locations</strong> and connected to AWS locations via a <strong class="nx ip">dedicated Open Connect Backbone network</strong>.</p><figure class="pv pw px py pz qa ps pt paragraph-image"><div role="button" tabindex="0" class="qb qc fl qd bh qe"><div class="ps pt ql"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*drjqrhSVFS7jfQzOBoG7dA.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*drjqrhSVFS7jfQzOBoG7dA.png" /><img alt="" class="bh fw qf c" width="700" height="433" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="qg ff qh ps pt qi qj bf b bg ab du">18K+<em class="qm"> servers in 6K+ locations, in Internet Exchanges, or embedded into ISP networks</em></figcaption></figure><figure class="pv pw px py pz qa ps pt paragraph-image"><div role="button" tabindex="0" class="qb qc fl qd bh qe"><div class="ps pt qn"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*KWo6BzotdnxqoveWCvY12Q.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*KWo6BzotdnxqoveWCvY12Q.png" /><img alt="" class="bh fw qf c" width="700" height="349" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="qg ff qh ps pt qi qj bf b bg ab du"><em class="qm">Open Connect Backbone connects servers in Internet Exchange locations to 5 AWS regions</em></figcaption></figure><p id="11ff" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">By enabling Live delivery on Open Connect, we build on top of $1B+ in Netflix investments over the last 12 years focused on scaling the network and optimizing the performance of delivery servers. By sharing capacity across on-demand and Live viewership we improve utilization, and by caching past Live content on the same servers used for on-demand streaming, we can easily enable catch-up viewing.</p><p id="e6ba" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk"><strong class="nx ip">Optimizing Live Playback for Device Compatibility, Scale, Quality, and Stability</strong></p><p id="e965" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">To make Live accessible to the majority of our customers without upgrading their streaming devices, we settled on using <strong class="nx ip">HTTPS</strong>-based Live Streaming. While UDP-based protocols can provide additional features like ultra-low latency, HTTPS has ubiquitous support among devices and compatibility with delivery and encoding systems. Furthermore, we use <strong class="nx ip">AVC</strong> and <strong class="nx ip">HEVC</strong> video codecs, transcode with multiple quality levels up <strong class="nx ip">from SD to 4K</strong>, and use a <strong class="nx ip">2-second segment</strong> duration to balance compression efficiency, infrastructure load, and latency. While prioritizing streaming quality and playback stability, we have also achieved industry standard latency from camera to device, and continue to improve it.</p><p id="104e" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">To configure playback, the device player receives a playback manifest at the play start. The manifest contains items like the encoding bitrates and CDN servers players should use. We deliver the manifest from the cloud instead of the CDN, as it allows us to personalize the configuration for each device. To reference segments of the stream, the manifest includes a segment template that is used by devices to map a wall-clock time to URLs on the CDN. Using a segment template vs periodic polling for manifest updates minimizes network dependencies, CDN server load, and overhead on resource-constrained devices, like smart TVs, thus improving both scalability and stability of our system. While streaming, the player monitors network performance and dynamically chooses the bitrate and CDN server, maximizing streaming quality while minimizing rebuffering.</p><p id="e72a" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk"><strong class="nx ip">Run Discovery and Playback Control Services in the Cloud</strong></p><p id="e33b" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">So far, we have covered the streaming path from Camera to Device. To make the stream fully work, we also need to orchestrate across all systems, and ensure viewers can find and start the Live event. This functionality is performed by <strong class="nx ip">dozens of Cloud services</strong>, with functions like playback configuration, personalization, or metrics collection. These services tend to receive disproportionately <strong class="nx ip">higher loads around Live event start time</strong>, and Cloud deployment provides flexibility in dynamically scaling compute resources. Moreover, as Live demand tends to be localized, we are able to balance load across <strong class="nx ip">multiple AWS regions</strong>, better utilizing our global footprint. Deployment in the cloud also allows us to build a user experience where we embed Live content into a broader selection of entertainment options in the UI, like on-demand titles or Games.</p><p id="8c6e" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk"><strong class="nx ip">Centralize Real-time Metrics in the Cloud with Specialized Tools and Facilities</strong></p><p id="bce4" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">With control over ingest, encoding pipelines, the Open Connect CDN, and device players, we have nearly <strong class="nx ip">end-to-end observability</strong> into the Live workflow. During Live, we collect system and user metrics in real-time (e.g., where members see the title on Netflix and their quality of experience), alerting us to poor user experiences or degraded system performance. Our real-time monitoring is built using a mix of internally developed tools, such as <a class="ag hb" href="https://netflix.github.io/atlas-docs/" rel="noopener ugc nofollow" target="_blank">Atlas</a>, <a class="ag hb" href="https://netflix.github.io/mantis/" rel="noopener ugc nofollow" target="_blank">Mantis</a>, and <a class="ag hb" rel="noopener ugc nofollow" href="https://netflixtechblog.com/lumen-custom-self-service-dashboarding-for-netflix-8c56b541548c" target="_blank" data-discover="true">Lumen</a>, and open-source technologies, such as Kafka and <a class="ag hb" rel="noopener ugc nofollow" href="https://netflixtechblog.com/how-netflix-uses-druid-for-real-time-insights-to-ensure-a-high-quality-experience-19e1e8568d06" target="_blank" data-discover="true">Druid</a>, processing up to <strong class="nx ip">38 million events per second</strong> during some of our largest live events while providing critical metrics and operational insights in a matter of seconds. Furthermore, we set up dedicated <strong class="nx ip">“Control Center” facilities</strong>, which bring key metrics together to the operational team that monitors the event in real-time.</p><h1 id="6604" class="or os io bf ot ou ov ow gl ox oy oz gn pa pb pc pd pe pf pg ph pi pj pk pl pm bk">| Our key learnings so far</h1><p id="0598" class="pw-post-body-paragraph nv nw io nx b ny pn oa ob oc po oe of go pp oh oi gr pq ok ol gu pr on oo op hp bk">Building new functionality always brings fresh challenges and opportunities to learn, especially with a system as complex as Live. Even after three years, we’re still learning every day how to deliver Live events more effectively. Here are a few key highlights:</p><p id="3f57" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk"><strong class="nx ip">Extensive testing: </strong>Prior to Live we heavily relied on the predictable flow of on-demand traffic for pre-release canaries or A/B tests to validate deployments. But Live traffic was not always available, especially not at the scale representative of a big launch. As a result, we spent considerable effort to:</p><ol class=""><li id="3982" class="nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op qo qp qq bk">Generate internal “test streams,” which engineers use to run <strong class="nx ip">integration</strong>, <strong class="nx ip">regression</strong>, or <strong class="nx ip">smoke tests</strong> as part of the development lifecycle.</li><li id="892f" class="nv nw io nx b ny qr oa ob oc qs oe of go qt oh oi gr qu ok ol gu qv on oo op qo qp qq bk">Build synthetic <strong class="nx ip">load testing</strong> capabilities to stress test cloud and CDN systems. We use 2 approaches, allowing us to generate up to <strong class="nx ip">100K starts-per-second</strong>:<br /> — Capture, modify, and replay past Live production traffic, representing a diversity of user devices and request patterns.<br /> — Virtualize Netflix devices and generate traffic against CDN or Cloud endpoints to test the impact of the latest changes across all systems.</li><li id="b6ec" class="nv nw io nx b ny qr oa ob oc qs oe of go qt oh oi gr qu ok ol gu qv on oo op qo qp qq bk">Run automated <strong class="nx ip">failure injection</strong>, forcing missing or corrupted segments from the encoding pipeline, loss of a cloud region, network drop, or server timeouts.</li></ol><p id="82d4" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk"><strong class="nx ip">Regular practice: </strong>Despite rigorous pre-release testing, nothing beats a production environment, especially when operating at scale. We learned that having a regular schedule with diverse Live content is essential to making improvements while balancing the risks of member impact. We run<a class="ag hb" rel="noopener ugc nofollow" href="https://netflixtechblog.com/decision-making-at-netflix-33065fa06481" target="_blank" data-discover="true"> A/B tests</a>, perform <a class="ag hb" rel="noopener ugc nofollow" href="https://netflixtechblog.com/chap-chaos-automation-platform-53e6d528371f" target="_blank" data-discover="true">chaos testing</a>, operational exercises, and train operational teams for upcoming launches.</p><p id="4319" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk"><strong class="nx ip">Viewership predictions:</strong> We use prediction-based techniques to pre-provision Cloud and CDN capacity, and share forecasts with our ISP and Cloud partners ahead of time so they can plan network and compute resources. Then we complement them with reactive scaling of cloud systems powering sign-up, log-in, title discovery, and playback services to account for viewership exceeding our predictions. We have found success with forward-looking real-time viewership predictions <em class="oq">during</em> a live event, allowing us to take steps to mitigate risks earlier, before more members are impacted.</p><p id="3d52" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk"><strong class="nx ip">Graceful degradation: </strong>Despite our best efforts, we can (and did!) find ourselves in a situation where viewership exceeded our predictions and provisioned capacity. In this case, we developed a number of levers to continue streaming, even if it means gradually removing some nice-to-have features. For example, we use <a class="ag hb" rel="noopener ugc nofollow" href="https://netflixtechblog.com/enhancing-netflix-reliability-with-service-level-prioritized-load-shedding-e735e6ce8f7d" target="_blank" data-discover="true">service-level prioritized load shedding</a> to prioritize live traffic over non-critical traffic (like pre-fetch). Beyond that, we can lighten the experience, like dialing down personalization, disabling bookmarks, or lowering the maximum streaming quality. Our load tests include scenarios where we under-scale systems to validate desired behavior.</p><p id="255c" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk"><strong class="nx ip">Retry storms: </strong>When systems reach capacity, our key focus is to avoid cascading issues or further overloading systems with retries.Beyond system retries, users may retry manually — we’ve seen a 10x increase in traffic load due to stream restarts after viewing interruptions of as little as 30 seconds. We spent considerable time understanding device retry behavior in the presence of issues like network timeouts or missing segments. As a result, we implemented strategies like server-guided backoff for device retries, absorbing spikes via prioritized traffic shedding at Cloud Edge Gateway, and re-balancing traffic between cloud regions.</p><p id="ace2" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk"><strong class="nx ip">Contingency planning: </strong>“<em class="oq">Everyone has a plan until they get punched in the mouth</em>” is very relevant for Live. When something breaks, there is practically no time for troubleshooting. For large events, we set up <strong class="nx ip">in-person launch rooms</strong> with engineering owners of critical systems. For quick detection and response, we developed a small set of metrics as early indicators of issues, and have extensive runbooks for common operational issues. We don’t learn on launch day; instead, launch teams practice failure response via <strong class="nx ip">Game Day exercises ahead of time</strong>. Finally, our runbooks extend beyond engineering, covering escalation to executive leadership and coordination across functions like Customer Service, Production, Communications, or Social.</p><p id="018a" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">Our commitment to enhancing the member experience doesn’t end at the “Thanks for Watching!” screen. Shortly after each live stream, we dive into metrics to identify areas for improvement. Our Data &amp; Insights team conducts comprehensive analyses, A/B tests, and consumer research to ensure the next event is even more delightful for our members. We leverage insights on member behavior, preferences, and expectations to refine the Netflix product experience and optimize our Live technology — like reducing latency by ~10 seconds through A/B tests, without affecting quality or stability.</p><h1 id="1acd" class="or os io bf ot ou ov ow gl ox oy oz gn pa pb pc pd pe pf pg ph pi pj pk pl pm bk">| What’s next on our Live journey?</h1><p id="de44" class="pw-post-body-paragraph nv nw io nx b ny pn oa ob oc po oe of go pp oh oi gr pq ok ol gu pr on oo op hp bk">Despite three years of effort, we are far from done! In fact, we are just getting started, actively building on the learnings shared above to deliver more joy to our members with Live events. To support the growing number of Live titles and new formats, like <a class="ag hb" href="https://www.netflix.com/tudum/articles/womens-world-cup-netflix" rel="noopener ugc nofollow" target="_blank">FIFA WWC in 2027</a>, we keep building our broadcast and delivery infrastructure and are actively working to further improve the Live experience.</p><p id="2732" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">In this post, we’ve provided a broad overview and have barely scratched the surface. In the upcoming posts, we will dive deeper into key pillars of our Live systems, covering our encoding, delivery, playback, and user experience investments in more detail.</p><p id="95fd" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">Getting this far would not have been possible without the hard work of dozens of teams across Netflix, who collaborate closely to design, build, and operate Live systems: Operations and Reliability, Encoding Technologies, Content Delivery, Device Playback, Streaming Algorithms, UI Engineering, Search and Discovery, Messaging, Content Promotion and Distribution, Data Platform, Cloud Infrastructure, Tooling and Productivity, Program Management, Data Science &amp; Engineering, Product Management, Globalization, Consumer Insights, Ads, Security, Payments, Live Production, Experience and Design, Product Marketing and Customer Service, amongst many others.</p></div></div><div class="qa"><div class="ac cb"><div class="my qw mz qx na qy cf qz cg ra ci bh"><div class="pv pw px py pz ac lu"><figure class="mu qa rb rc rd re rf paragraph-image"><div role="button" tabindex="0" class="qb qc fl qd bh qe"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*yokgEf1StGt84gHMxQqtrQ.jpeg" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*yokgEf1StGt84gHMxQqtrQ.jpeg" /><img alt="" class="bh fw qf c" width="500" height="3024" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></figure><figure class="mu qa rb rc rd re rf paragraph-image"><div role="button" tabindex="0" class="qb qc fl qd bh qe"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*C2TVMPryZ6NwH1-DgDy7Rw.jpeg" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*C2TVMPryZ6NwH1-DgDy7Rw.jpeg" /><img alt="" class="bh fw qf c" width="500" height="3024" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></figure></div><div class="ac lu"><figure class="mu qa rb rc rd re rf paragraph-image"><div role="button" tabindex="0" class="qb qc fl qd bh qe"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*q4FIjjQDicGkzlz-ktzhKA.jpeg" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*q4FIjjQDicGkzlz-ktzhKA.jpeg" /><img alt="" class="bh fw qf c" width="500" height="3024" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></figure><figure class="mu qa rb rc rd re rf paragraph-image"><div role="button" tabindex="0" class="qb qc fl qd bh qe"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*Zj_RUWUUlkBs3JiEul6AjQ.jpeg" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*Zj_RUWUUlkBs3JiEul6AjQ.jpeg" /><img alt="" class="bh fw qf c" width="500" height="3024" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></figure></div></div></div></div></div>]]></description>
      <link>https://netflixtechblog.com/behind-the-streams-live-at-netflix-part-1-d23f917c2f40</link>
      <guid>https://netflixtechblog.com/behind-the-streams-live-at-netflix-part-1-d23f917c2f40</guid>
      <pubDate>Tue, 15 Jul 2025 18:04:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Netflix Tudum Architecture: from CQRS with Kafka to CQRS with RAW Hollow]]></title>
      <description><![CDATA[<div class="ac cb"><div class="ci bh hv hw hx hy"><div><div></div><p id="10a0" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">By <a class="ag hb" href="https://www.linkedin.com/in/eugeneemelyanov/" rel="noopener ugc nofollow" target="_blank">Eugene Yemelyanau</a>, <a class="ag hb" href="https://www.linkedin.com/in/jake-grice/" rel="noopener ugc nofollow" target="_blank">Jake Grice</a></p><figure class="ot ou ov ow ox oy oq or paragraph-image"><div role="button" tabindex="0" class="oz pa fl pb bh pc"><div class="oq or os"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*5J0sqUrM77vCIDsM8wq_Dw.jpeg" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*5J0sqUrM77vCIDsM8wq_Dw.jpeg" /><img alt="" class="bh fw pd c" width="700" height="368" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><h1 id="ffb3" class="pe pf io bf pg ph pi pj gl pk pl pm gn pn po pp pq pr ps pt pu pv pw px py pz bk"><em class="qa">Introduction</em></h1><p id="2f88" class="pw-post-body-paragraph nv nw io nx b ny qb oa ob oc qc oe of go qd oh oi gr qe ok ol gu qf on oo op hp bk"><a class="ag hb" href="http://tudum.com" rel="noopener ugc nofollow" target="_blank"><em class="qg">Tudum.com</em></a><em class="qg"> is Netflix’s official fan destination, enabling fans to dive deeper into their favorite Netflix shows and movies. Tudum offers exclusive first-looks, behind-the-scenes content, talent interviews, live events, guides, and interactive experiences. “Tudum” is named after the sonic ID you hear when pressing play on a Netflix show or movie. Attracting over 20 million members each month, Tudum is designed to enrich the viewing experience by offering additional context and insights into the content available on Netflix.</em></p><h1 id="655e" class="pe pf io bf pg ph pi pj gl pk pl pm gn pn po pp pq pr ps pt pu pv pw px py pz bk">Initial architecture</h1><p id="3da2" class="pw-post-body-paragraph nv nw io nx b ny qb oa ob oc qc oe of go qd oh oi gr qe ok ol gu qf on oo op hp bk">At the end of 2021, when we envisioned Tudum’s implementation, we considered architectural patterns that would be maintainable, extensible, and well-understood by engineers. With the goal of building a flexible, configuration-driven system, we looked to <strong class="nx ip">server-driven UI</strong> (SDUI) as an appealing solution. SDUI is a design approach where the server dictates the structure and content of the UI, allowing for dynamic updates and customization without requiring changes to the client application. Client applications like web, mobile, and TV devices, act as rendering engines for SDUI data. After our teams weighed and vetted all the details, the dust settled and we landed on an approach similar to Command Query Responsibility Segregation (<a class="ag hb" href="https://www.geeksforgeeks.org/cqrs-command-query-responsibility-segregation/" rel="noopener ugc nofollow" target="_blank">CQRS</a>). At Tudum, we have two main use cases that CQRS is perfectly capable of solving:</p><ul class=""><li id="9a85" class="nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op qh qi qj bk"><strong class="nx ip">Tudum’s editorial team</strong> brings exclusive interviews, first-look photos, behind the scenes videos, and many more forms of fan-forward content, and compiles it all into pages on the <a class="ag hb" href="http://tudum.com" rel="noopener ugc nofollow" target="_blank">Tudum.com</a> website. This content comes onto Tudum in the form of individually published pages, and content elements within the pages. In support of this, Tudum’s architecture includes a write path to store all of this data, including internal comments, revisions, version history, asset metadata, and scheduling settings.</li><li id="a323" class="nv nw io nx b ny qk oa ob oc ql oe of go qm oh oi gr qn ok ol gu qo on oo op qh qi qj bk"><strong class="nx ip">Tudum visitors</strong> consume published pages. In this case, Tudum needs to serve personalized experiences for our beloved fans, and accesses only the latest version of our content.</li></ul></div></div><div class="oy"><div class="ac cb"><div class="my qp mz qq na qr cf qs cg qt ci bh"><figure class="ot ou ov ow ox oy qv qw paragraph-image"><div role="button" tabindex="0" class="oz pa fl pb bh pc"><div class="oq or qu"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*i_LBGZ4i7QWeiDLES88HoA.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*i_LBGZ4i7QWeiDLES88HoA.png" /><img alt="" class="bh fw pd c" width="1000" height="779" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="qx ff qy oq or qz ra bf b bg ab du">Initial Tudum data architecture</figcaption></figure></div></div></div><div class="ac cb"><div class="ci bh hv hw hx hy"><p id="6413" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">The high-level diagram above focuses on storage &amp; distribution, illustrating how we leveraged Kafka to separate the write and read databases. The write database would store internal page content and metadata from our CMS. The read database would store read-optimized page content, for example: CDN image URLs rather than internal asset IDs, and movie titles, synopses, and actor names instead of placeholders. This content ingestion pipeline allowed us to regenerate all consumer-facing content on demand, applying new structure and data, such as global navigation or branding changes. The Tudum Ingestion Service converted internal CMS data into a read-optimized format by applying page templates, running validations, performing data transformations, and producing the individual content elements into a Kafka topic. The Data Service Consumer, received the content elements from Kafka, stored them in a high-availability database (Cassandra), and acted as an API layer for the Page Construction service and other internal Tudum services to retrieve content.</p><p id="704d" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">A key advantage of decoupling read and write paths is the ability to scale them independently. It is a well-known architectural approach to connect both write and read databases using an event driven architecture. As a result, content edits would <strong class="nx ip"><em class="qg">eventually</em></strong> appear on <a class="ag hb" href="http://tudum.com" rel="noopener ugc nofollow" target="_blank">tudum.com</a>.</p><h1 id="38cc" class="pe pf io bf pg ph pi pj gl pk pl pm gn pn po pp pq pr ps pt pu pv pw px py pz bk">Challenges with eventual consistency</h1><p id="05f4" class="pw-post-body-paragraph nv nw io nx b ny qb oa ob oc qc oe of go qd oh oi gr qe ok ol gu qf on oo op hp bk">Did you notice the emphasis on “<strong class="nx ip"><em class="qg">eventually</em></strong>?” A major downside of this architecture was the delay between making an edit and observing that edit reflected on the website. For instance, when the team publishes an update, the following steps must occur:</p><ol class=""><li id="d25e" class="nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op rb qi qj bk">Call the REST endpoint on the 3rd party CMS to save the data.</li><li id="3462" class="nv nw io nx b ny qk oa ob oc ql oe of go qm oh oi gr qn ok ol gu qo on oo op rb qi qj bk">Wait for the CMS to notify the Tudum Ingestion layer via a webhook.</li><li id="3b1e" class="nv nw io nx b ny qk oa ob oc ql oe of go qm oh oi gr qn ok ol gu qo on oo op rb qi qj bk">Wait for the Tudum Ingestion layer to query all necessary sections via API, validate data and assets, process the page, and produce the modified content to Kafka.</li><li id="140d" class="nv nw io nx b ny qk oa ob oc ql oe of go qm oh oi gr qn ok ol gu qo on oo op rb qi qj bk">Wait for the Data Service Consumer to consume this message from Kafka and store it in the database.</li><li id="d298" class="nv nw io nx b ny qk oa ob oc ql oe of go qm oh oi gr qn ok ol gu qo on oo op rb qi qj bk">Finally, after some <strong class="nx ip">cache refresh delay</strong>, this data would <strong class="nx ip"><em class="qg">eventually</em></strong> become available to the Page Construction service. Great!</li></ol><p id="4a60" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">By introducing a highly-scalable eventually-consistent architecture we were missing the ability to quickly render changes after writing them — an important capability for internal previews.</p><p id="6b8e" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">In our performance profiling, we found the source of delay was our Page Data Service which acted as a facade for an underlying <a class="ag hb" rel="noopener ugc nofollow" href="https://netflixtechblog.com/introducing-netflixs-key-value-data-abstraction-layer-1ea8a0a11b30" target="_blank" data-discover="true">Key Value Data Abstraction</a> database. Page Data Service utilized a <strong class="nx ip">near cache</strong> to accelerate page building and reduce read latencies from the database.</p><p id="9814" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">This cache was implemented to optimize the N+1 key lookups necessary for page construction by having a complete data set in memory. When engineers hear “<em class="qg">slow reads</em>,” the immediate answer is often “<em class="qg">cache</em>,” which is exactly what our team adopted. The KVDAL near cache can refresh in the background on every app node. Regardless of which system modifies the data, the cache is updated with each refresh cycle. If you have 60 keys and a refresh interval of 60 seconds, the near cache will update one key per second. This was problematic for previewing recent modifications, as these changes were only reflected with each cache refresh. As Tudum’s content grew, cache refresh times increased, further extending the delay.</p><h1 id="0410" class="pe pf io bf pg ph pi pj gl pk pl pm gn pn po pp pq pr ps pt pu pv pw px py pz bk">RAW Hollow</h1><p id="76e2" class="pw-post-body-paragraph nv nw io nx b ny qb oa ob oc qc oe of go qd oh oi gr qe ok ol gu qf on oo op hp bk">As this pain point grew, a new technology was being developed that would act as our silver bullet. <a class="ag hb" href="https://hollow.how/raw-hollow-sigmod.pdf" rel="noopener ugc nofollow" target="_blank">RAW Hollow</a> is an innovative in-memory, co-located, compressed object database developed by Netflix, designed to handle small to medium datasets with support for strong read-after-write consistency. It addresses the challenges of achieving consistent performance with low latency and high availability in applications that deal with less frequently changing datasets. Unlike traditional SQL databases or fully in-memory solutions, RAW Hollow offers a unique approach where the entire dataset is distributed across the application cluster and resides in the memory of each application process.</p><p id="3cf2" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">This design leverages compression techniques to scale datasets up to 100 million records per entity, ensuring extremely low latencies and high availability. RAW Hollow provides eventual consistency by default, with the option for strong consistency at the individual request level, allowing users to balance between high availability and data consistency. It simplifies the development of highly available and scalable stateful applications by eliminating the complexities of cache synchronization and external dependencies. This makes RAW Hollow a robust solution for efficiently managing datasets in environments like Netflix’s streaming services, where high performance and reliability are paramount.</p><h1 id="1a6a" class="pe pf io bf pg ph pi pj gl pk pl pm gn pn po pp pq pr ps pt pu pv pw px py pz bk">Revised architecture</h1><p id="e493" class="pw-post-body-paragraph nv nw io nx b ny qb oa ob oc qc oe of go qd oh oi gr qe ok ol gu qf on oo op hp bk">Tudum was a perfect fit to battle-test RAW Hollow while it was pre-GA internally. Hollow’s high-density near cache significantly reduces I/O. Having our primary dataset in memory enables Tudum’s various microservices (page construction, search, personalization) to access data synchronously in O(1) time, simplifying architecture, reducing code complexity, and increasing fault tolerance.</p></div></div><div class="oy"><div class="ac cb"><div class="my qp mz qq na qr cf qs cg qt ci bh"><figure class="ot ou ov ow ox oy qv qw paragraph-image"><div role="button" tabindex="0" class="oz pa fl pb bh pc"><div class="oq or rc"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*XpvbAvfxMmfUq4oBC_E_BA.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*XpvbAvfxMmfUq4oBC_E_BA.png" /><img alt="" class="bh fw pd c" width="1000" height="900" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="qx ff qy oq or qz ra bf b bg ab du">Updated Tudum data architecture</figcaption></figure></div></div></div><div class="ac cb"><div class="ci bh hv hw hx hy"><p id="24fa" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">In our simplified architecture, we eliminated the Page Data Service, Key Value store, and Kafka infrastructure, in favor of RAW Hollow. By embedding the in-memory client directly into our read-path services, we avoid per-request I/O and reduce roundtrip time.</p><h1 id="6980" class="pe pf io bf pg ph pi pj gl pk pl pm gn pn po pp pq pr ps pt pu pv pw px py pz bk">Migration results</h1><p id="9b6a" class="pw-post-body-paragraph nv nw io nx b ny qb oa ob oc qc oe of go qd oh oi gr qe ok ol gu qf on oo op hp bk">The updated architecture yielded a monumental reduction in data propagation times, and the reduced I/O led to faster request times as an added bonus. Hollow’s compression alleviated our concerns about our data being “too big” to fit in memory. Storing three years’ of unhydrated data requires only a 130MB memory footprint — 25% of its uncompressed size in an Iceberg table!</p><p id="adc5" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">Writers and editors can preview changes in seconds instead of minutes, while still maintaining high-availability and in-memory caching for Tudum visitors — the best of both worlds.</p><p id="b810" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">But what about the faster request times? The diagram below illustrates the before &amp; after timing to fulfil a request for Tudum’s home page. All of Tudum’s read-path services leverage Hollow in-memory state, leading to a significant increase in page construction speed and personalization algorithms. Controlling for factors like TLS, authentication, request logging, and WAF filtering, homepage construction time decreased from ~1.4 seconds to ~0.4 seconds!</p></div></div><div class="oy"><div class="ac cb"><div class="my qp mz qq na qr cf qs cg qt ci bh"><figure class="ot ou ov ow ox oy qv qw paragraph-image"><div role="button" tabindex="0" class="oz pa fl pb bh pc"><div class="oq or rd"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*o7LgLonS37PvGPY04iVMCg.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*o7LgLonS37PvGPY04iVMCg.png" /><img alt="" class="bh fw pd c" width="1000" height="793" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="qx ff qy oq or qz ra bf b bg ab du">Home page construction time</figcaption></figure></div></div></div><div class="ac cb"><div class="ci bh hv hw hx hy"><p id="f4b8" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">An attentive reader might notice that we have now tightly-coupled our Page Construction Service with the Hollow In-Memory State. This tight-coupling is used only in Tudum-specific applications. However, caution is needed if sharing the Hollow In-Memory Client with other engineering teams, as it could limit your ability to make schema changes or deprecations.</p><h1 id="a4d4" class="pe pf io bf pg ph pi pj gl pk pl pm gn pn po pp pq pr ps pt pu pv pw px py pz bk">Key Learnings</h1><ol class=""><li id="a421" class="nv nw io nx b ny qb oa ob oc qc oe of go qd oh oi gr qe ok ol gu qf on oo op rb qi qj bk">CQRS is a powerful design paradigm for scale, if you can tolerate some eventual consistency.</li><li id="7771" class="nv nw io nx b ny qk oa ob oc ql oe of go qm oh oi gr qn ok ol gu qo on oo op rb qi qj bk">Minimizing the number of sequential operations can significantly reduce response times. I/O is often the main enemy of performance.</li><li id="c7fc" class="nv nw io nx b ny qk oa ob oc ql oe of go qm oh oi gr qn ok ol gu qo on oo op rb qi qj bk">Caching is complicated. Cache invalidation is a hard problem. By holding an entire dataset in memory, you can eliminate an entire class of problems.</li></ol><p id="2f6b" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">In the next episode, we’ll share how <a class="ag hb" href="http://tudum.com" rel="noopener ugc nofollow" target="_blank">Tudum.com</a> leverages Server Driven UI to rapidly build and deploy new experiences for Netflix fans. Stay tuned!</p><h1 id="5218" class="pe pf io bf pg ph pi pj gl pk pl pm gn pn po pp pq pr ps pt pu pv pw px py pz bk">Credits</h1><p id="5972" class="pw-post-body-paragraph nv nw io nx b ny qb oa ob oc qc oe of go qd oh oi gr qe ok ol gu qf on oo op hp bk">Thanks to <a class="ag hb" href="https://www.linkedin.com/in/koszewnik" rel="noopener ugc nofollow" target="_blank">Drew Koszewnik</a>, <a class="ag hb" href="https://www.linkedin.com/in/govindvenkatramankrishnan" rel="noopener ugc nofollow" target="_blank">Govind Venkatraman Krishnan</a>, <a class="ag hb" href="https://www.linkedin.com/in/nick-mooney-193849/" rel="noopener ugc nofollow" target="_blank">Nick Mooney</a></p></div></div></div>]]></description>
      <link>https://netflixtechblog.com/netflix-tudum-architecture-from-cqrs-with-kafka-to-cqrs-with-raw-hollow-86d141b72e52</link>
      <guid>https://netflixtechblog.com/netflix-tudum-architecture-from-cqrs-with-kafka-to-cqrs-with-raw-hollow-86d141b72e52</guid>
      <pubDate>Thu, 10 Jul 2025 21:31:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Driving Content Delivery Efficiency Through Classifying Cache Misses]]></title>
      <description><![CDATA[<div><div></div><p id="7f7c" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">By <a class="ag hb" href="https://www.linkedin.com/in/mvipulbharat/" rel="noopener ugc nofollow" target="_blank">Vipul Marlecha</a>, <a class="ag hb" href="https://www.linkedin.com/in/lara-deek-79773966" rel="noopener ugc nofollow" target="_blank">Lara Deek</a>, <a class="ag hb" href="https://www.linkedin.com/in/thiaraortiz" rel="noopener ugc nofollow" target="_blank">Thiara Ortiz</a></p><p id="b29b" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk"><em class="oq">The mission of </em><a class="ag hb" href="https://openconnect.netflix.com/en/#what-is-open-connect" rel="noopener ugc nofollow" target="_blank"><em class="oq">Open Connect</em></a><em class="oq">, our dedicated content delivery network (CDN), is to deliver the best quality of experience (QoE) to our members. By localizing our Open Connect Appliances (OCAs), we bring Netflix content closer to the end user. This is achieved through close partnerships with internet service providers (ISPs) worldwide. Our ability to efficiently localize traffic, known as Content Delivery Efficiency, is a critical component of Open Connect’s service.</em></p><p id="8d23" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk"><em class="oq">In this post, we discuss one of the frameworks we use to evaluate our efficiency and identify sources of inefficiencies. Specifically, we classify the causes of traffic not being served from local servers, a phenomenon that we refer to as cache misses.</em></p><h2 id="fb9e" class="or os io bf ot gk ou dy gl gm ov ea gn go ow gp gq gr ox gs gt gu oy gv gw oz bk">Why does Netflix have the Open Connect Program?</h2><p id="f869" class="pw-post-body-paragraph nv nw io nx b ny pa oa ob oc pb oe of go pc oh oi gr pd ok ol gu pe on oo op hp bk">The Open Connect Program is a cornerstone of Netflix’s commitment to delivering unparalleled QoE for our customers. By localizing traffic delivery from Open Connect servers at IX or ISP sites, we significantly enhance the speed and reliability of content delivery. The inherent latencies of data traveling across physical links, compounded by Internet infrastructure components like routers and network stacks, can disrupt a seamless viewing experience. Delays in video start times, reduced initial video quality, and the frustrating occurrence of buffering lead to an overall reduction in customer QoE. Open Connect empowers Netflix to maintain hyper-efficiency, ensuring a flawless client experience for new, latency-sensitive, on-demand content such as live streams and ads.</p><p id="827e" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">Our custom-built servers, known as Open Connect Appliances (OCAs), are designed for both efficiency and cost-effectiveness. By logging detailed historical streaming behavior and using it to model and forecast future trends, we hyper-optimize our OCAs for long-term caching efficiency. We build methods to efficiently and reliably store, stream, and move our content.</p><p id="dae4" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">The mission of Open Connect hinges on our ability to effectively localize content on our OCAs globally, despite limited storage space, and also by design with specific storage sizes. This ensures that our cost and power efficiency metrics continue to improve, enhancing client QoE and reducing costs for our ISP partners. A critical question we continuously ask is: How do we evaluate and monitor which bytes should have been served from local OCAs but resulted in a cache miss?</p><p id="361a" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk"><strong class="nx ip">The Anatomy of a Playback Request</strong></p><p id="03de" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">Let us start by introducing the logic that directs or “steers” a specific Netflix client device to its dedicated OCA. The lifecycle from when a client device presses play until the video starts being streamed to that device is referred to as “playback.” Figure 1 illustrates the logical components involved in playback.</p><figure class="pi pj pk pl pm pn pf pg paragraph-image"><div role="button" tabindex="0" class="po pp fl pq bh pr"><div class="pf pg ph"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*VvQOBlOQLLAkBFOw%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*VvQOBlOQLLAkBFOw%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*VvQOBlOQLLAkBFOw%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*VvQOBlOQLLAkBFOw%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*VvQOBlOQLLAkBFOw%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*VvQOBlOQLLAkBFOw%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*VvQOBlOQLLAkBFOw%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*VvQOBlOQLLAkBFOw 640w, https://miro.medium.com/v2/resize:fit:720/0*VvQOBlOQLLAkBFOw 720w, https://miro.medium.com/v2/resize:fit:750/0*VvQOBlOQLLAkBFOw 750w, https://miro.medium.com/v2/resize:fit:786/0*VvQOBlOQLLAkBFOw 786w, https://miro.medium.com/v2/resize:fit:828/0*VvQOBlOQLLAkBFOw 828w, https://miro.medium.com/v2/resize:fit:1100/0*VvQOBlOQLLAkBFOw 1100w, https://miro.medium.com/v2/resize:fit:1400/0*VvQOBlOQLLAkBFOw 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh fw ps c" width="700" height="405" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="ec11" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk"><strong class="nx ip">Figure 1:</strong> Components for Playback</p><p id="eb5d" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">The components involved in playback are important to understand as we elaborate on the concept of how we determine a cache miss versus hit. Independent of client requests, every OCA in our CDN periodically reports its capacity and health, learned BGP routes, and current list of stored files. All of this data is reported to the Cache Control Service (CCS). When a member hits the play button, this request is sent to our AWS services, specifically the Playback Apps service. After Playback Apps determines which files correspond to a specific movie request, it issues a request to “steer” the client’s playback request to OCAs via the Steering Service. The Steering Service in turn, using the data reported from OCAs to CCS as well as other client information such as geo location, identifies the set of OCAs that can satisfy that client’s request. This set of OCAs is then returned in the form of rank-ordered URLs to the client device, the client connects to the top-ranked OCA and requests the files it needs to begin the video stream.</p><h2 id="11f7" class="or os io bf ot gk ou dy gl gm ov ea gn go ow gp gq gr ox gs gt gu oy gv gw oz bk">What is a Cache Miss?</h2><p id="0d83" class="pw-post-body-paragraph nv nw io nx b ny pa oa ob oc pb oe of go pc oh oi gr pd ok ol gu pe on oo op hp bk">A cache miss occurs when bytes are not served from the best available OCA for a given Netflix client, independent of OCA state. For each playback request, the Steering Service computes a ranked list of local sites for the client, ordered by network proximity alone. This ranked list of sites is known as the “proximity rank.” Network proximity is determined based on the IP ranges (BGP routes) that are advertised by our ISP partners. Any OCA from the first “most proximal” site on this list is the most preferred and closest, having advertised the longest, most specific matching prefix to the client’s IP address. A cache miss is logged when bytes are not streamed from any OCA at this first local site, and we log when and why that happens.</p><p id="9760" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">It is important to note that our concept of cache misses is viewed from the client’s perspective, focusing on the optimal delivery source for the end user and prepositioning content accordingly, rather than relying on traditional CDN proxy caching mechanisms. Our “prepositioning” differentiator allows us to prioritize client QoE by ensuring content is served from the most optimal OCA.</p><p id="4340" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">We attribute cache misses to three logical categories. The intuition behind the delineated categories is that each category informs parallel strategies to achieve content delivery efficiency.</p><ul class=""><li id="22ba" class="nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op pt pu pv bk"><strong class="nx ip">Content Miss:</strong> This happens when the files were not found on OCAs in the local site. In previous articles like “<a class="ag hb" rel="noopener ugc nofollow" href="https://netflixtechblog.com/content-popularity-for-open-connect-b86d56f613b" target="_blank" data-discover="true">Content Popularity for Open Connect</a>” and “<a class="ag hb" rel="noopener ugc nofollow" href="https://netflixtechblog.com/distributing-content-to-open-connect-3e3e391d4dc9" target="_blank" data-discover="true">Distributing Content to Open Connect</a>,” we discuss how we decide what content to prioritize populating first onto our OCAs. A sample of efforts this insights informs include: (1) how accurately we predict the popularity of content, (2) how rapidly we pre-position that content, (3) how well we design our OCA hardware, and (4) how well we provision storage capacity at our locations of presence.</li><li id="6c6c" class="nv nw io nx b ny pw oa ob oc px oe of go py oh oi gr pz ok ol gu qa on oo op pt pu pv bk"><strong class="nx ip">Health Miss:</strong> This happens when the local site’s OCA hardware resources are becoming saturated, and one or more OCA can not handle more traffic. As a result, we direct clients to other OCAs with capacity to serve that content. Each OCA has a control loop that monitors its bottleneck metrics (such as CPU, disk usage, etc.) and assesses its ability to serve additional traffic. This is referred to as “OCA health.” Insight into health misses informs efforts such as: (1) how well we load balance traffic across OCAs with heterogeneous hardware resources, (2) how well we provision enough copies of highly popular content to distribute massive traffic, which is also tied to how accurately we predict the popularity of content, and (3) how well we preposition content to specific hardware components with varying traffic serve capabilities and bottlenecks.</li></ul><p id="88df" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">Next we will dig into the framework we built to log and compute these metrics in real-time, with some extra attention to technical detail.</p><h2 id="354c" class="or os io bf ot gk ou dy gl gm ov ea gn go ow gp gq gr ox gs gt gu oy gv gw oz bk">Cache Miss Computation Framework</h2><h2 id="7414" class="or os io bf ot gk ou dy gl gm ov ea gn go ow gp gq gr ox gs gt gu oy gv gw oz bk">Logging Components</h2><p id="06fa" class="pw-post-body-paragraph nv nw io nx b ny pa oa ob oc pb oe of go pc oh oi gr pd ok ol gu pe on oo op hp bk">There are two critical data components that we log, gather, and analyze to compute cache misses:</p><ul class=""><li id="fef8" class="nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op pt pu pv bk"><strong class="nx ip">Steering Playback Manifest Logs:</strong> Within the Steering Service, we compute and log the ranked list of sites for each client request, i.e. the “proximity rank” introduced earlier. We also enrich that list with information that reflects the logical decisions and filters our algorithms applied across all proximity ranks given that point-in-time state of our systems. This information allows us to replay/simulate any hypothetical scenario easily, such as to evaluate whether an outage across all sites in the first proximity rank would overwhelm sites in the second proximity rank, and many more such scenarios!</li><li id="d1c3" class="nv nw io nx b ny pw oa ob oc px oe of go py oh oi gr pz ok ol gu qa on oo op pt pu pv bk"><strong class="nx ip">OCA Server Logs:</strong> Once a Netflix client connects with an OCA to begin video streaming, the OCAs log any data regarding that streaming session, such as the files streamed and total bytes. All OCA logs are consolidated to identify which OCA(s) each client actually watched its video stream from, and the amount of content streamed.</li></ul><p id="0d1e" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">The above logs are joined for every Netflix client’s playback request to compute detailed cache miss metrics (in bytes and hours streamed) at different aggregation levels (such as per OCA, movie, file, encode type, country, and so on).</p><h2 id="508a" class="or os io bf ot gk ou dy gl gm ov ea gn go ow gp gq gr ox gs gt gu oy gv gw oz bk">System Architecture</h2><p id="63f2" class="pw-post-body-paragraph nv nw io nx b ny pa oa ob oc pb oe of go pc oh oi gr pd ok ol gu pe on oo op hp bk">Figure 2 outlines how the logging components fit into the general engineering architecture that allows us to compute content miss metrics at low-latency and almost real-time.</p><figure class="pi pj pk pl pm pn pf pg paragraph-image"><div role="button" tabindex="0" class="po pp fl pq bh pr"><div class="pf pg qb"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*PlQ4xv4Si8iWGnW1%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*PlQ4xv4Si8iWGnW1%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*PlQ4xv4Si8iWGnW1%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*PlQ4xv4Si8iWGnW1%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*PlQ4xv4Si8iWGnW1%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*PlQ4xv4Si8iWGnW1%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*PlQ4xv4Si8iWGnW1%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*PlQ4xv4Si8iWGnW1 640w, https://miro.medium.com/v2/resize:fit:720/0*PlQ4xv4Si8iWGnW1 720w, https://miro.medium.com/v2/resize:fit:750/0*PlQ4xv4Si8iWGnW1 750w, https://miro.medium.com/v2/resize:fit:786/0*PlQ4xv4Si8iWGnW1 786w, https://miro.medium.com/v2/resize:fit:828/0*PlQ4xv4Si8iWGnW1 828w, https://miro.medium.com/v2/resize:fit:1100/0*PlQ4xv4Si8iWGnW1 1100w, https://miro.medium.com/v2/resize:fit:1400/0*PlQ4xv4Si8iWGnW1 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh fw ps c" width="700" height="354" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="8a58" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk"><strong class="nx ip">Figure 2:</strong> Components of the cache miss computation framework.</p><p id="b723" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">We will now describe the system requirements of each component.</p><ol class=""><li id="7a9e" class="nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op qc pu pv bk"><strong class="nx ip">Log Emission</strong>: The logs for computing cache miss are emitted to Kafka clusters in each of our evaluated AWS regions, enabling us to send logs with the lowest possible latency. After a client device makes a playback request, the Steering Service generates a <em class="oq">steering playback manifest</em>, logs it, and sends the data to a Kafka cluster. Kafka is used for event streaming at Netflix because of its high-throughput event processing, low latency, and reliability. After the client device starts the video stream from an OCA, the OCA stores information about the bytes served for each file requested by each unique client playback stream. This data is what we refer to as <em class="oq">OCA server logs</em>.</li><li id="4c2c" class="nv nw io nx b ny pw oa ob oc px oe of go py oh oi gr pz ok ol gu qa on oo op qc pu pv bk"><strong class="nx ip">Log Consolidation</strong>: The logs emitted by the Steering Service and the OCAs can result in data for a single playback request being distributed across different AWS regions, because logs are recorded in geographically distributed Kafka clusters. <em class="oq">OCA server logs</em> might be stored in one region’s Kafka cluster while <em class="oq">steering playback manifest logs</em> are stored in another. One approach to consolidate data for a single playback is to build complex many-to-many joins. In streaming pipelines, performing these joins requires replicating logs across all regions, which leads to data duplication and increased complexity. This setup complicates downstream data processing and inflates operational costs due to multiple redundant cross-region data transfers. To overcome these challenges, we perform a cross-region transfer only once, consolidating all logs into a single region.</li><li id="89ce" class="nv nw io nx b ny pw oa ob oc px oe of go py oh oi gr pz ok ol gu qa on oo op qc pu pv bk"><strong class="nx ip">Log Enrichment</strong>: We enrich the logs during streaming joins with metadata using various slow-changing dimension tables and services so that we have the necessary information about the OCA and the played content.</li><li id="4bd0" class="nv nw io nx b ny pw oa ob oc px oe of go py oh oi gr pz ok ol gu qa on oo op qc pu pv bk"><strong class="nx ip">Streaming Window-Based Join</strong>: We perform a streaming window-based join to merge the <em class="oq">steering playback manifest logs</em> with the <em class="oq">OCA server logs</em>. Performing enrichment and log consolidation upstream allows for more seamless and un-interrupted joining of our log data sources.</li><li id="b8ac" class="nv nw io nx b ny pw oa ob oc px oe of go py oh oi gr pz ok ol gu qa on oo op qc pu pv bk"><strong class="nx ip">Cache Miss Calculations</strong>: After joining the logs, we compute the cache miss metrics. The computation checks whether the client played content from an OCA in the first site listed in the <em class="oq">steering playback manifest</em>’s proximity rank or from another site. When a video stream occurs at a higher proximity rank, this indicates that a cache miss occurred.</li></ol><h1 id="f012" class="qd os io bf ot qe qf qg gl qh qi qj gn qk ql qm qn qo qp qq qr qs qt qu qv qw bk">Data Model to Evaluate Cache Misses</h1><p id="c8ef" class="pw-post-body-paragraph nv nw io nx b ny pa oa ob oc pb oe of go pc oh oi gr pd ok ol gu pe on oo op hp bk">One of the most exciting opportunities we have enabled through these logs (in these authors’ opinions) is the ability to replay our logic offline and in simulations with variable parameters, to reproduce impact in production under different conditions. This allows us to test new conditions, features, and hypothetical scenarios without impacting production Netflix traffic.</p><p id="c132" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">To achieve the above, our data should satisfy two main conditions. First, the data should be comprehensive in representing the state of each distinct logical step involved in steering, including the decisions and their reasons. In order to achieve this, the underlying logic, here the Steering Service, needs to be built in a modularized fashion, where each logical component overlays data from the prior component, resulting in a rich blurb representing the system’s full state, which is finally logged. This all needs to be achieved without adding perceivable latency to client playback requests! Second, the data should be in a format that allows near-real-time aggregate metrics for monitoring purposes.</p><p id="dae9" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">Some components of our final, joined data model that enables us to collect rich insights in a scalable and timely manner are listed in Table 1.</p><p id="4cfa" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk"><strong class="nx ip">Table 1: Unified Data Model after joining <em class="oq">steering playback manifest</em> and <em class="oq">OCA server logs</em>.</strong></p><figure class="pi pj pk pl pm pn pf pg paragraph-image"><div role="button" tabindex="0" class="po pp fl pq bh pr"><div class="pf pg qx"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*7kbXh8GcB8P75TPWfsjTig.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*7kbXh8GcB8P75TPWfsjTig.png" /><img alt="" class="bh fw ps c" width="700" height="619" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><h2 id="65ca" class="or os io bf ot gk ou dy gl gm ov ea gn go ow gp gq gr ox gs gt gu oy gv gw oz bk">Cache Miss Computation Sample</h2><p id="ba38" class="pw-post-body-paragraph nv nw io nx b ny pa oa ob oc pb oe of go pc oh oi gr pd ok ol gu pe on oo op hp bk">Let us share an example of how we compute cache miss metrics. For a given unique client play request, we know we had a cache miss when the client streams from an OCA that is not in the client’s first proximity rank. As you can see from Table 1, each file needed for a client’s video streaming session is linked to routable OCAs and their corresponding sites with a proximity rank. These are 0 based indexes with proximity rank zero indicating the most optimal OCA for the client. “Proximity Rank Zero” indicates that the client connected to an OCA in the most preferred site(s), thus no misses occurred. Higher proximity ranks indicate a miss has occurred. The aggregation of all bytes and hours streamed from non-preferred sites constitutes a missed opportunity for Netflix and are reported in our cache miss metrics.</p><p id="ef7e" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk"><strong class="nx ip">Decision Labels and Bytes Sent</strong></p><p id="bd09" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">Sourced from the <em class="oq">steering playback manifest logs</em>, we record why we did not select an OCA for playback. These are denoted by:</p><ul class=""><li id="7f1f" class="nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op pt pu pv bk">“H”: Health miss.</li><li id="ae1f" class="nv nw io nx b ny pw oa ob oc px oe of go py oh oi gr pz ok ol gu qa on oo op pt pu pv bk">“C”: Content miss.</li></ul><p id="5c56" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk"><strong class="nx ip">Metrics Calculation and Categorization</strong></p><p id="891c" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">For each file needed for a client’s video streaming session, we can categorize the bytes streamed by the client into different types of misses:</p><ul class=""><li id="f864" class="nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op pt pu pv bk">No Miss: If proximity rank is zero, bytes were streamed from the optimal OCA.</li><li id="e378" class="nv nw io nx b ny pw oa ob oc px oe of go py oh oi gr pz ok ol gu qa on oo op pt pu pv bk">Health Miss (“H”): Miss due to the OCA reporting high utilization.</li><li id="213c" class="nv nw io nx b ny pw oa ob oc px oe of go py oh oi gr pz ok ol gu qa on oo op pt pu pv bk">Content Miss (“C”): Miss due to the OCA not having the content available locally.</li></ul><h2 id="c8cd" class="or os io bf ot gk ou dy gl gm ov ea gn go ow gp gq gr ox gs gt gu oy gv gw oz bk">How are miss metrics used to monitor our efficiency?</h2><p id="5903" class="pw-post-body-paragraph nv nw io nx b ny pa oa ob oc pb oe of go pc oh oi gr pd ok ol gu pe on oo op hp bk">Open Connect uses cache miss metrics to manage our Open Connect infrastructure. One of the team’s goals is to reduce the frequency of these cache misses, as they indicate that our members are being served by less proximal OCAs. By maintaining a detailed set of metrics that reveal the reasons behind cache misses, we can set up alerts to quickly identify when members are streaming from suboptimal locations. This is crucial because we operate a global CDN with millions of members worldwide and tens of thousands of servers.</p><p id="fca9" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">The figure below illustrates how we track the volume of total streaming traffic alongside the proportion of traffic streamed from less preferred locations due to content shedding. By calculating the ratio of content shed traffic to total streamed traffic, we derive a content shed ratio:</p><p id="1ea9" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">content shed ratio = content shed traffic total streamed traffic</p><figure class="pi pj pk pl pm pn pf pg paragraph-image"><div role="button" tabindex="0" class="po pp fl pq bh pr"><div class="pf pg qy"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*qoDQqy_9y6mffW9C%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*qoDQqy_9y6mffW9C%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*qoDQqy_9y6mffW9C%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*qoDQqy_9y6mffW9C%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*qoDQqy_9y6mffW9C%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*qoDQqy_9y6mffW9C%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*qoDQqy_9y6mffW9C%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*qoDQqy_9y6mffW9C 640w, https://miro.medium.com/v2/resize:fit:720/0*qoDQqy_9y6mffW9C 720w, https://miro.medium.com/v2/resize:fit:750/0*qoDQqy_9y6mffW9C 750w, https://miro.medium.com/v2/resize:fit:786/0*qoDQqy_9y6mffW9C 786w, https://miro.medium.com/v2/resize:fit:828/0*qoDQqy_9y6mffW9C 828w, https://miro.medium.com/v2/resize:fit:1100/0*qoDQqy_9y6mffW9C 1100w, https://miro.medium.com/v2/resize:fit:1400/0*qoDQqy_9y6mffW9C 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh fw ps c" width="700" height="376" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="687d" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">This active monitoring of content shedding allows us to maintain a tight feedback loop to ensure the efficacy of our deployment and prediction algorithms, streaming traffic, and the QoE of our members. Given that content shedding can occur for multiple reasons, it is essential to have clear signals indicating when it happens, along with known and automated remediation strategies, such as mechanisms to quickly deploy mispredicted content onto OCAs. When special intervention is necessary to minimize shedding, we use it as an opportunity to enhance our systems as well as to ensure they are comprehensive in considering all known failure cases.</p><h2 id="20b3" class="or os io bf ot gk ou dy gl gm ov ea gn go ow gp gq gr ox gs gt gu oy gv gw oz bk">Conclusion</h2><p id="fde1" class="pw-post-body-paragraph nv nw io nx b ny pa oa ob oc pb oe of go pc oh oi gr pd ok ol gu pe on oo op hp bk">Open Connect’s unique strategy requires us to be incredibly efficient in delivering content from our OCAs. We closely track miss metrics to ensure we are maximizing the traffic our members stream from most proximal locations. This ensures we are delivering the best quality of experience to our members globally.</p><p id="e265" class="pw-post-body-paragraph nv nw io nx b ny nz oa ob oc od oe of go og oh oi gr oj ok ol gu om on oo op hp bk">Our methods for managing cache misses are evolving, especially with the introduction of new streaming types like Live and Ads, which have different streaming behaviors and access patterns compared to traditional video. We remain committed to identifying and seizing opportunities for improvement as we face new challenges.</p></div>]]></description>
      <link>https://netflixtechblog.com/driving-content-delivery-efficiency-through-classifying-cache-misses-ffcf08026b6c</link>
      <guid>https://netflixtechblog.com/driving-content-delivery-efficiency-through-classifying-cache-misses-ffcf08026b6c</guid>
      <pubDate>Wed, 02 Jul 2025 17:20:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[AV1 @ Scale: Film Grain Synthesis, The Awakening]]></title>
      <description><![CDATA[<div class="ac cb"><div class="ci bh hv hw hx hy"><div><div><h2 id="57a9" class="pw-subtitle-paragraph jm in io bf b jn jo jp jq jr js jt ju jv jw jx jy jz ka kb cq du"><em class="jl">Unleashing Film Grain Synthesis on Netflix and Enhancing Visuals for Millions</em></h2><div></div><p id="aa6f" class="pw-post-body-paragraph oi oj io ok b jn ol om on jq oo op oq go or os ot gr ou ov ow gu ox oy oz pa hp bk"><a class="ag hb" href="https://www.linkedin.com/in/li-heng-chen-a75458a2/" rel="noopener ugc nofollow" target="_blank">Li-Heng Chen</a>, <a class="ag hb" href="https://www.linkedin.com/in/andreynorkin/" rel="noopener ugc nofollow" target="_blank">Andrey Norkin</a>, <a class="ag hb" href="https://www.linkedin.com/in/liwei-guo/" rel="noopener ugc nofollow" target="_blank">Liwei Guo</a>, <a class="ag hb" href="https://www.linkedin.com/in/henryzhili/" rel="noopener ugc nofollow" target="_blank">Zhi Li</a>, <a class="ag hb" href="https://www.linkedin.com/in/agataopalach/" rel="noopener ugc nofollow" target="_blank">Agata Opalach</a> and <a class="ag hb" href="https://www.linkedin.com/in/anush-moorthy-b8451142/" rel="noopener ugc nofollow" target="_blank">Anush Moorthy</a></p><figure class="pe pf pg ph pi pj pb pc paragraph-image"><div role="button" tabindex="0" class="pk pl fl pm bh pn"><div class="pb pc pd"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*7WHYl75ij_-W7YDLj2rz2w.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*7WHYl75ij_-W7YDLj2rz2w.png" /><img alt="" class="bh fw po c" width="700" height="371" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="358d" class="pw-post-body-paragraph oi oj io ok b jn ol om on jq oo op oq go or os ot gr ou ov ow gu ox oy oz pa hp bk"><strong class="ok ip">Picture this: you’re watching a classic film, and the subtle dance of film grain adds a layer of authenticity and nostalgia to every scene.</strong> This grain, formed from tiny particles during the film’s development, is more than just a visual effect. It plays a key role in storytelling by enhancing the film’s depth and contributing to its realism. However, film grain is as elusive as it is beautiful. Its random nature makes it notoriously difficult to compress. Traditional compression algorithms struggle to manage it, often forcing a choice between preserving the grain and reducing file size.</p><p id="b381" class="pw-post-body-paragraph oi oj io ok b jn ol om on jq oo op oq go or os ot gr ou ov ow gu ox oy oz pa hp bk">In the digital age, noise remains a ubiquitous element in video content. Camera sensor noise introduces its own characteristics, while filmmakers often add intentional grain during post-production to evoke mood or a vintage feel. These elements create a visually rich experience that tests conventional compression methods.</p><p id="177f" class="pw-post-body-paragraph oi oj io ok b jn ol om on jq oo op oq go or os ot gr ou ov ow gu ox oy oz pa hp bk">We’re giving members globally a transformed streaming experience with the recent rollout of AV1 Film Grain Synthesis (FGS) streams. While FGS has been part of the AV1 standard since its inception, we only enabled it for a limited number of titles during <a class="ag hb" rel="noopener ugc nofollow" href="https://netflixtechblog.com/bringing-av1-streaming-to-netflix-members-tvs-b7fc88e42320" target="_blank" data-discover="true">our initial launch of the AV1 codec in 2021</a>. Now, we’re enabling this innovative technology at scale, leveraging it to preserve the artistic integrity of film grain while optimizing data efficiency. In this blog post, we’ll explore how FGS revolutionizes video streaming and enhances your viewing experience.</p><h1 id="34c5" class="pp pq io bf pr ps pt jp gl pu pv js gn pw px py pz qa qb qc qd qe qf qg qh qi bk">Understanding Film Grain Synthesis in AV1</h1><p id="3ec2" class="pw-post-body-paragraph oi oj io ok b jn qj om on jq qk op oq go ql os ot gr qm ov ow gu qn oy oz pa hp bk">The AV1 Film Grain Synthesis tool models film grain through two key components, with model parameters estimated before the encoding of the denoised video:</p><p id="fefa" class="pw-post-body-paragraph oi oj io ok b jn ol om on jq oo op oq go or os ot gr ou ov ow gu ox oy oz pa hp bk"><strong class="ok ip">Film Grain Pattern</strong>: an <em class="qo">auto-regressive (AR) model</em> is used to replicate the pattern of film grain. The key parameters are the AR coefficients, which can be estimated from the residual between the source video and the denoised video, essentially capturing the noise. This model captures the spatial correlation between the grain samples, ensuring that the noise characteristics of the original content are accurately preserved. By adjusting the auto-regressive coefficients {ai}, the model can control the grain’s shape, making it appear coarser or finer. With these coefficients, a 64x64 noise template is generated, as illustrated in the animation below. To construct the noise layer during playback, random 32x32 patches are extracted from the 64x64 noise template and added to the decoded video.</p><figure class="pe pf pg ph pi pj pb pc paragraph-image"><div role="button" tabindex="0" class="pk pl fl pm bh pn"><div class="pb pc qp"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*uNAjI0TwiFlfclpC%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*uNAjI0TwiFlfclpC%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*uNAjI0TwiFlfclpC%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*uNAjI0TwiFlfclpC%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*uNAjI0TwiFlfclpC%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*uNAjI0TwiFlfclpC%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*uNAjI0TwiFlfclpC%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*uNAjI0TwiFlfclpC 640w, https://miro.medium.com/v2/resize:fit:720/0*uNAjI0TwiFlfclpC 720w, https://miro.medium.com/v2/resize:fit:750/0*uNAjI0TwiFlfclpC 750w, https://miro.medium.com/v2/resize:fit:786/0*uNAjI0TwiFlfclpC 786w, https://miro.medium.com/v2/resize:fit:828/0*uNAjI0TwiFlfclpC 828w, https://miro.medium.com/v2/resize:fit:1100/0*uNAjI0TwiFlfclpC 1100w, https://miro.medium.com/v2/resize:fit:1400/0*uNAjI0TwiFlfclpC 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh fw po c" width="700" height="291" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="qq ff qr pb pc qs qt bf b bg ab du">Fig. 1 The synthesis process of the 64x64 noise template using the simplest AR kernel with a lag parameter L=1. Each noise value is calculated as a linear combination of previously synthesized noise sample values, with AR coefficients a0, a1, a2, a3 and a white Gaussian noise (wgn) component.</figcaption></figure><p id="0ab0" class="pw-post-body-paragraph oi oj io ok b jn ol om on jq oo op oq go or os ot gr ou ov ow gu ox oy oz pa hp bk"><strong class="ok ip">Film Grain Intensity</strong>: a <em class="qo">scaling function</em> is employed to control the grain’s appearance under varying lighting conditions. This function, estimated during the encoding process, models the relationship between pixel value and noise intensity using a piecewise linear function. This allows for precise adjustments to the grain strength based on video brightness and color. Consequently, the film grain strength is adapted to the areas of the picture, closely recreating the look of the original video. The animation below demonstrates how the grain intensity is adjusted by the scaling function:</p><figure class="pe pf pg ph pi pj pb pc paragraph-image"><div role="button" tabindex="0" class="pk pl fl pm bh pn"><div class="pb pc qu"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*juBpnCJo31CO0imqoAyTpA.gif" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*juBpnCJo31CO0imqoAyTpA.gif" /><img alt="" class="bh fw po c" width="700" height="289" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="qq ff qr pb pc qs qt bf b bg ab du">Fig. 2 Illustration of the scaling function’s impact on film grain intensity. Left: The scaling function graph showing the relationship between pixel value and scaling intensity. Right: A grayscale SMPTE bars frame with film grain applied according to the scaling function.</figcaption></figure><p id="c03f" class="pw-post-body-paragraph oi oj io ok b jn ol om on jq oo op oq go or os ot gr ou ov ow gu ox oy oz pa hp bk">With these models specified by AV1 standard, the encoding process first removes the film grain from the video. The standard does not mandate a specific method for this step, allowing users to choose their preferred denoiser. Following the denoising, the video is compressed, and the grain’s pattern and intensity are estimated and transmitted alongside the compressed video data. During playback, the film grain is recreated and reintegrated into the video using a block-based method. This approach is optimized for consumer devices, ensuring smooth playback and high-quality visuals. For a more detailed explanation, please refer to the <a class="ag hb" href="https://norkin.org/pdf/DCC_2018_AV1_film_grain.pdf" rel="noopener ugc nofollow" target="_blank">original paper</a>.</p><p id="e365" class="pw-post-body-paragraph oi oj io ok b jn ol om on jq oo op oq go or os ot gr ou ov ow gu ox oy oz pa hp bk">By combining these components, the AV1 Film Grain Synthesis tool preserves the artistic integrity of film grain while making the content “easier to compress” by denoising the source video prior to encoding. This process enables high-quality video streaming, even in content with heavy grain, resulting in significant bitrate savings and improved visual quality.</p><h1 id="3b4b" class="pp pq io bf pr ps pt jp gl pu pv js gn pw px py pz qa qb qc qd qe qf qg qh qi bk">Visual Quality Improvement, Bitrate Reduction, and Member Benefits</h1><p id="d1b0" class="pw-post-body-paragraph oi oj io ok b jn qj om on jq qk op oq go ql os ot gr qm ov ow gu qn oy oz pa hp bk">In our pursuit of premium streaming quality, enabling AV1 Film Grain Synthesis has led to significant bitrate reduction, allowing us to deliver high-quality video with less data while preserving the artistic integrity of film grain. Below, we showcase visual examples highlighting the improved quality and reduced bitrate, using a frame from the Netflix title <a class="ag hb" href="https://www.netflix.com/title/80996324" rel="noopener ugc nofollow" target="_blank"><em class="qo">They Cloned Tyrone</em></a>:</p></div></div><div class="pj"><div class="ac cb"><div class="nl qv nm qw nn qx cf qy cg qz ci bh"><figure class="pe pf pg ph pi pj rb rc paragraph-image"><div role="button" tabindex="0" class="pk pl fl pm bh pn"><div class="pb pc ra"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*6mrNIZrV9yJhsw3yn-ZfJQ.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*6mrNIZrV9yJhsw3yn-ZfJQ.png" /><img alt="" class="bh fw po c" width="1000" height="563" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="qq ff qr pb pc qs qt bf b bg ab du">A source video frame from <em class="jl">They Cloned Tyrone</em></figcaption></figure><figure class="nh pj rb rc paragraph-image"><div role="button" tabindex="0" class="pk pl fl pm bh pn"><div class="pb pc ra"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*1GnNYWlnl343CHFdQAHEHw.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*1GnNYWlnl343CHFdQAHEHw.png" /><img alt="" class="bh fw po c" width="1000" height="563" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="qq ff qr pb pc qs qt bf b bg ab du">Regular AV1 (without FGS) @ 8274 kbps</figcaption></figure><figure class="nh pj rb rc paragraph-image"><div role="button" tabindex="0" class="pk pl fl pm bh pn"><div class="pb pc ra"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*STC86FuqtMqO_hbz-H3CBQ.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*STC86FuqtMqO_hbz-H3CBQ.png" /><img alt="" class="bh fw po c" width="1000" height="563" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="qq ff qr pb pc qs qt bf b bg ab du">AV1 with FGS @ 2804 kbps</figcaption></figure></div></div></div><div class="ac cb"><div class="ci bh hv hw hx hy"><p id="2e2e" class="pw-post-body-paragraph oi oj io ok b jn ol om on jq oo op oq go or os ot gr ou ov ow gu ox oy oz pa hp bk">The visual comparison highlights a significant bitrate reduction of approximately 66%, with regular AV1 encoding at 8274 kbps compared to AV1 with FGS at 2804 kbps. In this example, which features strong film grain, it may be observed that the regular version exhibits distorted noise with a discrete cosine transform (DCT)-like pattern. In contrast, the FGS version preserves the integrity of the film grain at a lower bitrate.</p><p id="0d0a" class="pw-post-body-paragraph oi oj io ok b jn ol om on jq oo op oq go or os ot gr ou ov ow gu ox oy oz pa hp bk">Additionally, synthesized noise effectively masks compression artifacts, resulting in a more visually appealing experience. In this comparison below, both the regular AV1 stream and the AV1 FGS stream without synthesized noise (equivalent to compressing the denoised video) show compression artifacts. In contrast, the AV1 FGS stream with grain synthesis (rightmost figure) improves visual quality through contrast masking in human visual systems. The added film grain, a form of mask, effectively conceals some compression artifacts.</p></div></div><div class="pj"><div class="ac cb"><div class="nl qv nm qw nn qx cf qy cg qz ci bh"><div class="pe pf pg ph pi ac mh"><figure class="nh pj rd re rb rc rf paragraph-image"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*ms9UEY7w_LyQFv14kZJHWQ.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*ms9UEY7w_LyQFv14kZJHWQ.png" /><img alt="" class="bh fw po c" width="220" height="320" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></figure><figure class="nh pj rd re rb rc rf paragraph-image"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*um6VAX8lSeGHPTT--94rZA.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*um6VAX8lSeGHPTT--94rZA.png" /><img alt="" class="bh fw po c" width="220" height="320" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></figure><figure class="nh pj rd re rb rc rf paragraph-image"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*mB9eWSgP9Fete-g-dqYSig.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*mB9eWSgP9Fete-g-dqYSig.png" /><img alt="" class="bh fw po c" width="220" height="320" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture><figcaption class="qq ff qr pb pc qs qt bf b bg ab du rg fl rh ri">Cropped frame comparison: Regular AV1 stream (Left), AV1 FGS stream <strong class="bf pr">without</strong> grain synthesis during decoding (Middle), and AV1 FGS stream with grain synthesis (Right).</figcaption></figure></div></div></div></div><div class="ac cb"><div class="ci bh hv hw hx hy"><p id="5092" class="pw-post-body-paragraph oi oj io ok b jn ol om on jq oo op oq go or os ot gr ou ov ow gu ox oy oz pa hp bk">Currently, we lack a dedicated quality model for film grain synthesis. The noise appearing at different pixel locations between the source and decoded video poses challenges for pixelwise comparison methods like <a class="ag hb" href="https://en.wikipedia.org/wiki/Peak_signal-to-noise_ratio" rel="noopener ugc nofollow" target="_blank">PSNR</a> or <a class="ag hb" rel="noopener ugc nofollow" href="https://netflixtechblog.com/toward-a-practical-perceptual-video-quality-metric-653f208b9652" target="_blank" data-discover="true">VMAF</a>, leading to penalized quality scores. Despite this, our internal assessment highlights the improvements in visual quality and the value of these advancements.</p><p id="2281" class="pw-post-body-paragraph oi oj io ok b jn ol om on jq oo op oq go or os ot gr ou ov ow gu ox oy oz pa hp bk">To evaluate the impact of AV1 Film Grain Synthesis, we selected approximately 300 titles from the Netflix catalog, each with varying levels of graininess. The bar chart below illustrates a 36% reduction in average bitrate for resolutions of 1080p and above when AV1 film grain synthesis is enabled, highlighting its efficacy in optimizing data usage. For resolutions below 1080p, the reduction in bitrate is relatively small, reaching only a 10% decrease, likely because noise is filtered out during the downscaling process. Furthermore, enabling the film grain synthesis coding tool consistently introduces syntax overhead to the bitstream.</p><figure class="pe pf pg ph pi pj pb pc paragraph-image"><div role="button" tabindex="0" class="pk pl fl pm bh pn"><div class="pb pc qp"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*fVfNc9DIh878py8G%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*fVfNc9DIh878py8G%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*fVfNc9DIh878py8G%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*fVfNc9DIh878py8G%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*fVfNc9DIh878py8G%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*fVfNc9DIh878py8G%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*fVfNc9DIh878py8G%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*fVfNc9DIh878py8G 640w, https://miro.medium.com/v2/resize:fit:720/0*fVfNc9DIh878py8G 720w, https://miro.medium.com/v2/resize:fit:750/0*fVfNc9DIh878py8G 750w, https://miro.medium.com/v2/resize:fit:786/0*fVfNc9DIh878py8G 786w, https://miro.medium.com/v2/resize:fit:828/0*fVfNc9DIh878py8G 828w, https://miro.medium.com/v2/resize:fit:1100/0*fVfNc9DIh878py8G 1100w, https://miro.medium.com/v2/resize:fit:1400/0*fVfNc9DIh878py8G 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh fw po c" width="700" height="296" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="qq ff qr pb pc qs qt bf b bg ab du">Fig. 3: Comparison of average values across resolution categories between regular AV1 streams (without film grain synthesis) and AV1 streams with film grain synthesis enabled.</figcaption></figure><p id="1e09" class="pw-post-body-paragraph oi oj io ok b jn ol om on jq oo op oq go or os ot gr ou ov ow gu ox oy oz pa hp bk">Finally, we conducted A/B testing prior to rollout to understand the overall streaming impact of enabling AV1 Film Grain Synthesis. This testing showcased a smoother and more reliable Quality of Experience (QoE) for our members. The improvements include:</p><ul class=""><li id="242b" class="oi oj io ok b jn ol om on jq oo op oq go or os ot gr ou ov ow gu ox oy oz pa rj rk rl bk"><strong class="ok ip">Lower Initial and Average Bitrate</strong>: Bitrate at the start of the playback reduced by 24% and average bitrate by 31.6%, lower network bandwidth requirements and reduced storage needs for downloaded streams.</li><li id="9d6e" class="oi oj io ok b jn rm om on jq rn op oq go ro os ot gr rp ov ow gu rq oy oz pa rj rk rl bk"><strong class="ok ip">Decreased Playback Errors</strong>: Playback error rate reduced by approximately 3%.</li><li id="dbc9" class="oi oj io ok b jn rm om on jq rn op oq go ro os ot gr rp ov ow gu rq oy oz pa rj rk rl bk"><strong class="ok ip">Reduced Rebuffering</strong>: 10% fewer rebuffers and a 5% reduction in rebuffer duration resulting from the lower bitrate.</li><li id="9907" class="oi oj io ok b jn rm om on jq rn op oq go ro os ot gr rp ov ow gu rq oy oz pa rj rk rl bk"><strong class="ok ip">Faster Start Play</strong>: Start play delay reduced by 10%, potentially due to the lower bitrate, which may help devices reach the target buffer level more quickly.</li><li id="66c9" class="oi oj io ok b jn rm om on jq rn op oq go ro os ot gr rp ov ow gu rq oy oz pa rj rk rl bk"><strong class="ok ip">Improved Playback Stability</strong>: Observed 10% fewer noticeable bitrate drops and a 10% reduction in the time users spend adjusting their playback position during video playback, likely influenced by reduced bitrate and rebuffering.</li><li id="81d7" class="oi oj io ok b jn rm om on jq rn op oq go ro os ot gr rp ov ow gu rq oy oz pa rj rk rl bk"><strong class="ok ip">Higher Resolution Streaming</strong>: About 0.7% of viewing hours shifted from lower resolutions (≤ 1080p) to 2160p on 4K-capable devices. This shift is attributed to reduced bitrates at switching points, which make it easier to achieve the highest resolution during a session.</li></ul><h1 id="cc76" class="pp pq io bf pr ps pt jp gl pu pv js gn pw px py pz qa qb qc qd qe qf qg qh qi bk">Behind the Scenes: Our Film Grain Adventure Continues</h1><p id="08cc" class="pw-post-body-paragraph oi oj io ok b jn qj om on jq qk op oq go ql os ot gr qm ov ow gu qn oy oz pa hp bk">We’re always excited to share our progress with the community. This blog provides an overview of our journey: from the initial launch of the AV1 codec to the recent addition of Film Grain Synthesis (FGS) streams, highlighting the impact these innovations have on Netflix’s streaming quality. Since March, we’ve been rolling out FGS across scale, and many users can now enjoy the FGS-enabled streams, provided their device supports this feature. We encourage you to watch some of the author’s favorite titles <a class="ag hb" href="https://www.netflix.com/title/81985186" rel="noopener ugc nofollow" target="_blank">The Hot Spot</a>, <a class="ag hb" href="https://www.netflix.com/title/70018511" rel="noopener ugc nofollow" target="_blank">Kung Fu Cult Master</a>, <a class="ag hb" href="https://www.netflix.com/title/70043379" rel="noopener ugc nofollow" target="_blank">Initial D</a>, <a class="ag hb" href="https://www.netflix.com/title/70005331" rel="noopener ugc nofollow" target="_blank">God of Gamblers II</a>, <a class="ag hb" href="https://www.netflix.com/title/80203996" rel="noopener ugc nofollow" target="_blank">Baahubali 2: The Conclusion</a>, or <a class="ag hb" href="https://www.netflix.com/title/81487660" rel="noopener ugc nofollow" target="_blank">Dept. Q</a> (you may need to toggle off HDR from the settings menu) on Netflix to experience the new FGS streams firsthand.</p><p id="2872" class="pw-post-body-paragraph oi oj io ok b jn ol om on jq oo op oq go or os ot gr ou ov ow gu ox oy oz pa hp bk">In the next post, we will share how we did this in <a class="ag hb" rel="noopener ugc nofollow" href="https://netflixtechblog.com/rebuilding-netflix-video-processing-pipeline-with-microservices-4e5e6310e359" target="_blank" data-discover="true">our video encoding pipeline</a>, detailing the process and insights we’ve gained. Stay tuned to the <a class="ag hb" href="https://netflixtechblog.com/" rel="noopener ugc nofollow" target="_blank">Netflix Tech Blog</a> for the latest updates.</p><h1 id="4c65" class="pp pq io bf pr ps pt jp gl pu pv js gn pw px py pz qa qb qc qd qe qf qg qh qi bk">Acknowledgments</h1><p id="950b" class="pw-post-body-paragraph oi oj io ok b jn qj om on jq qk op oq go ql os ot gr qm ov ow gu qn oy oz pa hp bk">This achievement is the result of a collaborative effort among several Open Connect teams at Netflix, including Video Algorithms, Media Encoding Pipeline, Media Foundations, Infrastructure Capacity Planning, and Open Connect Control Plane. We also received invaluable support from Client &amp; Partner Technologies, Streaming &amp; Discovery Experiences, Media Compute &amp; Storage Infrastructure, Data Science &amp; Engineering, and the Global Production Technology team. We would like to express our sincere gratitude to the following individuals for their contributions to the project’s success:</p><ul class=""><li id="5f2e" class="oi oj io ok b jn ol om on jq oo op oq go or os ot gr ou ov ow gu ox oy oz pa rj rk rl bk">Prudhvi Kumar Chaganti and Ken Thomas for the discussion and assistance on rollout strategy</li><li id="cb28" class="oi oj io ok b jn rm om on jq rn op oq go ro os ot gr rp ov ow gu rq oy oz pa rj rk rl bk">Poojarani Chennai Natarajan, Lara Deek , Ivan Ivanov, and Ishaan Shastri for their essential support in planning and operations for Open Connect.</li><li id="0a02" class="oi oj io ok b jn rm om on jq rn op oq go ro os ot gr rp ov ow gu rq oy oz pa rj rk rl bk">Alex Chang for his support in everything related to data analysis, and Jessica Tweneboah and Amelia Taylor for their assistance with AB testing.</li><li id="2acf" class="oi oj io ok b jn rm om on jq rn op oq go ro os ot gr rp ov ow gu rq oy oz pa rj rk rl bk">David Zheng, Janet Xue, Scott Bolter, Brian Li, Allan Zhou, Vivian Li, Sarah Kurdoghlian, Artem Danylenko, Greg Freedman, and many other dedicated team members played a crucial role in device certification and collaboration with device partners. Their efforts significantly improved compatibility across platforms. (Spoiler alert: this was one of the biggest challenges we faced for productizing AV1 FGS!)</li><li id="9eec" class="oi oj io ok b jn rm om on jq rn op oq go ro os ot gr rp ov ow gu rq oy oz pa rj rk rl bk">Javier Fernandez-Ivern and Ritesh Makharia expertly managed the playback logic</li><li id="834e" class="oi oj io ok b jn rm om on jq rn op oq go ro os ot gr rp ov ow gu rq oy oz pa rj rk rl bk">Joseph McCormick and JD Vandenberg for providing valuable insights from a content production point of view, and Alex ‘Ally’ Michaelson for assisting in monitoring customer service.</li><li id="06c6" class="oi oj io ok b jn rm om on jq rn op oq go ro os ot gr rp ov ow gu rq oy oz pa rj rk rl bk">A special thanks to Roger Quero, who played a key role in supporting various aspects of the project and contributed significantly to its overall success while he was at Netflix.</li></ul></div></div></div></div>]]></description>
      <link>https://netflixtechblog.com/av1-scale-film-grain-synthesis-the-awakening-ee09cfdff40b</link>
      <guid>https://netflixtechblog.com/av1-scale-film-grain-synthesis-the-awakening-ee09cfdff40b</guid>
      <pubDate>Wed, 02 Jul 2025 16:21:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Model Once, Represent Everywhere: UDA (Unified Data Architecture) at Netflix]]></title>
      <description><![CDATA[<div class="ac cb"><div class="ci bh ht hu hv hw"><div><div></div><p id="0456" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk">By <a class="ag gz" href="https://www.linkedin.com/in/ahutter/" rel="noopener ugc nofollow" target="_blank">Alex Hutter</a>, <a class="ag gz" href="https://www.linkedin.com/in/bertails/" rel="noopener ugc nofollow" target="_blank">Alexandre Bertails</a>, <a class="ag gz" href="https://www.linkedin.com/in/clairezwang0612/" rel="noopener ugc nofollow" target="_blank">Claire Wang</a>, <a class="ag gz" href="https://www.linkedin.com/in/haoyuan-h-98b587134/" rel="noopener ugc nofollow" target="_blank">Haoyuan He</a>, <a class="ag gz" href="https://www.linkedin.com/in/kishore-banala/" rel="noopener ugc nofollow" target="_blank">Kishore Banala</a>, <a class="ag gz" href="https://www.linkedin.com/in/peterroyal/" rel="noopener ugc nofollow" target="_blank">Peter Royal</a>, <a class="ag gz" href="https://www.linkedin.com/in/shervinafshar/" rel="noopener ugc nofollow" target="_blank">Shervin Afshar</a></p><p id="e595" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk">As Netflix’s offerings grow — across films, series, games, live events, and ads — so does the complexity of the systems that support it. Core business concepts like ‘actor’ or ‘movie’ are modeled in many places: in our Enterprise GraphQL Gateway powering internal apps, in our asset management platform storing media assets, in our media computing platform that powers encoding pipelines, to name a few. Each system models these concepts differently and in isolation, with little coordination or shared understanding. While they often operate on the same concepts, these systems remain largely unaware of that fact, and of each other.</p><figure class="oy oz pa pb pc pd ov ow paragraph-image"><div role="button" tabindex="0" class="pe pf fi pg bh ph"><div class="ov ow ox"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*wNYAhebbErEdYROL%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*wNYAhebbErEdYROL%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*wNYAhebbErEdYROL%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*wNYAhebbErEdYROL%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*wNYAhebbErEdYROL%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*wNYAhebbErEdYROL%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*wNYAhebbErEdYROL%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*wNYAhebbErEdYROL 640w, https://miro.medium.com/v2/resize:fit:720/0*wNYAhebbErEdYROL 720w, https://miro.medium.com/v2/resize:fit:750/0*wNYAhebbErEdYROL 750w, https://miro.medium.com/v2/resize:fit:786/0*wNYAhebbErEdYROL 786w, https://miro.medium.com/v2/resize:fit:828/0*wNYAhebbErEdYROL 828w, https://miro.medium.com/v2/resize:fit:1100/0*wNYAhebbErEdYROL 1100w, https://miro.medium.com/v2/resize:fit:1400/0*wNYAhebbErEdYROL 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="Spider-Man Pointing meme with each Spider-Man labelled as: “it’s a movie”, “it’s a tv show”, “it’s a game”." class="bh fu pi c" width="700" height="467" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="9f61" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk">As a result, several challenges emerge:</p><ul class=""><li id="c101" class="oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou pj pk pl bk"><strong class="oc in">Duplicated and Inconsistent Models</strong> — Teams re-model the same business entities in different systems, leading to conflicting definitions that are hard to reconcile.</li><li id="e1ed" class="oa ob im oc b od pm of og oh pn oj ok gm po om on gp pp op oq gs pq os ot ou pj pk pl bk"><strong class="oc in">Inconsistent Terminology</strong> — Even within a single system, teams may use different terms for the same concept, or the same term for different concepts, making collaboration harder.</li><li id="b67d" class="oa ob im oc b od pm of og oh pn oj ok gm po om on gp pp op oq gs pq os ot ou pj pk pl bk"><strong class="oc in">Data Quality Issues</strong> — Discrepancies and broken references are hard to detect across our many microservices. While identifiers and foreign keys exist, they are inconsistently modeled and poorly documented, requiring manual work from domain experts to find and fix any data issues.</li><li id="86d8" class="oa ob im oc b od pm of og oh pn oj ok gm po om on gp pp op oq gs pq os ot ou pj pk pl bk"><strong class="oc in">Limited Connectivity</strong> — Within systems, relationships between data are constrained by what each system supports. Across systems, they are effectively non-existent.</li></ul><p id="c853" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk">To address these challenges, we need new foundations that allow us to define a model once, at the conceptual level, and reuse those definitions everywhere. But it isn’t enough to just document concepts; we need to connect them to real systems and data. And more than just connect, we have to project those definitions outward, generating schemas and enforcing consistency across systems. The conceptual model must become part of the control plane.</p><p id="e618" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk">These were the core ideas that led us to build UDA.</p><h1 id="03da" class="pr ps im bf pt pu pv pw gj px py pz gl qa qb qc qd qe qf qg qh qi qj qk ql qm bk">Introducing UDA</h1><p id="0faf" class="pw-post-body-paragraph oa ob im oc b od qn of og oh qo oj ok gm qp om on gp qq op oq gs qr os ot ou hn bk"><strong class="oc in">UDA (Unified Data Architecture)</strong> is the foundation for connected data in <a class="ag gz" rel="noopener ugc nofollow" href="https://netflixtechblog.com/netflix-studio-engineering-overview-ed60afcfa0ce" target="_blank" data-discover="true">Content Engineering</a>. It enables teams to model domains once and represent them consistently across systems — powering automation, discoverability, and <a class="ag gz" href="https://en.wikipedia.org/wiki/Semantic_interoperability" rel="noopener ugc nofollow" target="_blank">semantic interoperability</a>.</p><p id="5baf" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk"><strong class="oc in">Using UDA, users and systems can:</strong></p><p id="105b" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk"><strong class="oc in">Register and connect domain models </strong>— formal conceptualizations of federated business domains expressed as data.</p><ul class=""><li id="27d4" class="oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou pj pk pl bk"><strong class="oc in">Why? </strong>So everyone uses the same official definitions for business concepts, which avoids confusion and stops different teams from rebuilding similar models in conflicting ways.</li></ul><p id="6a56" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk"><strong class="oc in">Catalog and map domain models to data containers</strong>, such as GraphQL type resolvers served by a <a class="ag gz" rel="noopener ugc nofollow" href="https://netflixtechblog.com/open-sourcing-the-netflix-domain-graph-service-framework-graphql-for-spring-boot-92b9dcecda18" target="_blank" data-discover="true">Domain Graph Service</a>, <a class="ag gz" rel="noopener ugc nofollow" href="https://netflixtechblog.com/data-mesh-a-data-movement-and-processing-platform-netflix-1288bcab2873" target="_blank" data-discover="true">Data Mesh sources</a>, or Iceberg tables, through their representation as a graph.</p><ul class=""><li id="36c0" class="oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou pj pk pl bk"><strong class="oc in">Why?</strong> To make it easy to find where the actual data for these business concepts lives (e.g., in which specific database, table, or service) and understand how it’s structured there.</li></ul><p id="b0a9" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk"><strong class="oc in">Transpile domain models into schema definition languages</strong> like GraphQL, Avro, SQL, RDF, and Java, while preserving semantics.</p><ul class=""><li id="5c58" class="oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou pj pk pl bk"><strong class="oc in">Why? </strong>To automatically create consistent technical data structures (schemas) for various systems directly from the domain models, saving developers manual effort and reducing errors caused by out-of-sync definitions.</li></ul><p id="d376" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk"><strong class="oc in">Move data faithfully between data containers</strong>, such as from federated GraphQL entities to <a class="ag gz" rel="noopener ugc nofollow" href="https://netflixtechblog.com/data-mesh-a-data-movement-and-processing-platform-netflix-1288bcab2873" target="_blank" data-discover="true">Data Mesh</a> (a general purpose data movement and processing platform for moving data between Netflix systems at scale), Change Data Capture (CDC) sources to joinable Iceberg Data Products.</p><ul class=""><li id="f165" class="oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou pj pk pl bk"><strong class="oc in">Why? </strong>To save developer time by automatically handling how data is moved and correctly transformed between different systems. This means less manual work to configure data movement, ensuring data shows up consistently and accurately wherever it’s needed.</li></ul><p id="1067" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk"><strong class="oc in">Discover and explore domain concepts </strong>via search and graph traversal.</p><ul class=""><li id="2c8b" class="oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou pj pk pl bk"><strong class="oc in">Why? </strong>So anyone can more easily find the specific business information they’re looking for, understand how different concepts and data are related, and be confident they are accessing the correct information.</li></ul><p id="c87b" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk"><strong class="oc in">Programmatically introspect the knowledge graph</strong> using Java, GraphQL, or SPARQL.</p><ul class=""><li id="a8f7" class="oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou pj pk pl bk"><strong class="oc in">Why?</strong> So developers can build smarter applications that leverage this connected business information, automate more complex data-dependent workflows, and help uncover new insights from the relationships in the data.</li></ul><p id="cb20" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk"><strong class="oc in">This post introduces the foundations of UDA</strong> as a knowledge graph, connecting domain models to data containers through mappings, and grounded in an in-house <a class="ag gz" href="https://en.wikipedia.org/wiki/Metamodeling#:~:text=A%20metamodel%2F%20surrogate%20model%20is,representing%20input%20and%20output%20relations" rel="noopener ugc nofollow" target="_blank">metamodel</a>, or model of models, called Upper. Upper defines the language for domain modeling in UDA and enables projections that automatically generate schemas and pipelines across systems.</p><figure class="oy oz pa pb pc pd ov ow paragraph-image"><div role="button" tabindex="0" class="pe pf fi pg bh ph"><div class="ov ow qs"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*j1I2cLD0vtfE9IQfNiUwVQ.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*j1I2cLD0vtfE9IQfNiUwVQ.png" /><img alt="Image of the UDA knowledge graph. A central node representing a domain model is connected to other nodes representing Data Mesh, GraphQL, and Iceberg data containers." class="bh fu pi c" width="700" height="723" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="qt fc qu ov ow qv qw bf b bg ab ed">The same domain model can be connected to semantically equivalent data containers in the UDA knowledge graph.</figcaption></figure><p id="b7d0" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk"><strong class="oc in">This post also highlights two systems</strong> that leverage UDA in production:</p><p id="995a" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk"><strong class="oc in">Primary Data Management (PDM)</strong> is our platform for managing authoritative reference data and taxonomies. PDM turns domain models into flat or hierarchical taxonomies that drive a generated UI for business users. These taxonomy models are projected into Avro and GraphQL schemas, automatically provisioning data products in the Warehouse and GraphQL APIs in the <a class="ag gz" rel="noopener ugc nofollow" href="https://netflixtechblog.com/how-netflix-scales-its-api-with-graphql-federation-part-1-ae3557c187e2" target="_blank" data-discover="true">Enterprise Gateway</a>.</p><p id="8584" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk"><strong class="oc in">Sphere</strong> is our self-service operational reporting tool for business users. Sphere uses UDA to catalog and relate business concepts across systems, enabling discovery through familiar terms like ‘actor’ or ‘movie.’ Once concepts are selected, Sphere walks the knowledge graph and generates SQL queries to retrieve data from the warehouse, no manual joins or technical mediation required.</p><h2 id="ba47" class="qx ps im bf pt gi qy du gj gk qz dw gl gm ra gn go gp rb gq gr gs rc gt gu rd bk">UDA is a Knowledge Graph</h2><p id="b725" class="pw-post-body-paragraph oa ob im oc b od qn of og oh qo oj ok gm qp om on gp qq op oq gs qr os ot ou hn bk"><strong class="oc in">UDA needs to solve the </strong><a class="ag gz" href="https://en.wikipedia.org/wiki/Data_integration" rel="noopener ugc nofollow" target="_blank"><strong class="oc in">data integration</strong></a><strong class="oc in"> problem. </strong>We needed a data catalog unified with a schema registry, but with a hard requirement for <a class="ag gz" href="https://en.wikipedia.org/wiki/Semantic_integration#:~:text=Semantic%20integration%20is%20the%20process,from%20diverse%20sources" rel="noopener ugc nofollow" target="_blank">semantic integration</a>. Connecting business concepts to schemas and data containers in a graph-like structure, grounded in strong semantic foundations, naturally led us to consider a <a class="ag gz" href="https://en.wikipedia.org/wiki/Knowledge_graph" rel="noopener ugc nofollow" target="_blank">knowledge graph</a> approach.</p><p id="0e88" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk"><strong class="oc in">We chose RDF and SHACL as the foundation for UDA’s knowledge graph</strong>. But operationalizing them at enterprise scale surfaced several challenges:</p><ul class=""><li id="2065" class="oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou pj pk pl bk"><strong class="oc in">RDF lacked a usable information model.</strong> While RDF offers a flexible graph structure, it provides little guidance on how to organize data into <a class="ag gz" href="https://www.w3.org/TR/rdf12-concepts/#dfn-named-graph" rel="noopener ugc nofollow" target="_blank">named graphs</a>, manage ontology ownership, or define governance boundaries. Standard <a class="ag gz" href="https://www.w3.org/2001/sw/wiki/Linking_patterns" rel="noopener ugc nofollow" target="_blank">follow-your-nose mechanisms</a> like owl:imports apply only to ontologies and don’t extend to named graphs; we needed a generalized mechanism to express and resolve dependencies between them.</li><li id="fd63" class="oa ob im oc b od pm of og oh pn oj ok gm po om on gp pp op oq gs pq os ot ou pj pk pl bk"><strong class="oc in">SHACL is not a modeling language for enterprise data.</strong> Designed to validate native RDF, SHACL assumes globally unique URIs and a single data graph. But enterprise data is structured around local schemas and typed keys, as in GraphQL, Avro, or SQL. SHACL could not express these patterns, making it difficult to model and validate real-world data across heterogeneous systems.</li><li id="bcdb" class="oa ob im oc b od pm of og oh pn oj ok gm po om on gp pp op oq gs pq os ot ou pj pk pl bk"><strong class="oc in">Teams lacked shared authoring practices.</strong> Without strong guidelines, teams modeled their ontologies inconsistently breaking semantic interoperability. Even subtle differences in style, structure, or naming led to divergent interpretations and made transpilation harder to define consistently across schemas.</li><li id="f055" class="oa ob im oc b od pm of og oh pn oj ok gm po om on gp pp op oq gs pq os ot ou pj pk pl bk"><strong class="oc in">Ontology tooling lacked support for collaborative modeling.</strong> Unlike GraphQL Federation, ontology frameworks had no built-in support for modular contributions, team ownership, or safe federation. Most engineers found the tools and concepts unfamiliar, and available authoring environments lacked the structure needed for coordinated contributions.</li></ul><p id="4a57" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk"><strong class="oc in">To address these challenges, UDA adopts a named-graph-first information model.</strong> Each named graph conforms to a governing model, itself a named graph in the knowledge graph. This systematic approach ensures resolution, modularity, and enables governance across the entire graph. While a full description of UDA’s information infrastructure is beyond the scope of this post, the next sections explain how UDA bootstraps the knowledge graph with its metamodel and uses it to model data container representations and mappings.</p><h2 id="5d0e" class="qx ps im bf pt gi qy du gj gk qz dw gl gm ra gn go gp rb gq gr gs rc gt gu rd bk">Upper is Domain Modeling</h2><p id="440f" class="pw-post-body-paragraph oa ob im oc b od qn of og oh qo oj ok gm qp om on gp qq op oq gs qr os ot ou hn bk"><strong class="oc in">Upper is a language for formally describing domains — business or system — and their concepts</strong>. <a class="ag gz" href="https://en.wikipedia.org/wiki/Conceptualization_(information_science)" rel="noopener ugc nofollow" target="_blank">These concepts are organized into domain models</a>: controlled vocabularies that define classes of keyed entities, their attributes, and their relationships to other entities, which may be keyed or nested, within the same domain or across domains. Keyed concepts within a domain model can be organized in taxonomies of types, which can be as complex as the business or the data system needs them to be. Keyed concepts can also be extended from other domain models — that is, new attributes and relationships can be <a class="ag gz" href="https://tomgruber.org/writing/onto-design.pdf#page=4" rel="noopener ugc nofollow" target="_blank">contributed monotonically</a>. Finally, Upper ships with a rich set of datatypes for attribute values, which can also be customized per domain.</p><figure class="oy oz pa pb pc pd ov ow paragraph-image"><div role="button" tabindex="0" class="pe pf fi pg bh ph"><div class="ov ow ox"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*A_-GpZLvqbxuVdkH%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*A_-GpZLvqbxuVdkH%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*A_-GpZLvqbxuVdkH%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*A_-GpZLvqbxuVdkH%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*A_-GpZLvqbxuVdkH%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*A_-GpZLvqbxuVdkH%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*A_-GpZLvqbxuVdkH%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*A_-GpZLvqbxuVdkH 640w, https://miro.medium.com/v2/resize:fit:720/0*A_-GpZLvqbxuVdkH 720w, https://miro.medium.com/v2/resize:fit:750/0*A_-GpZLvqbxuVdkH 750w, https://miro.medium.com/v2/resize:fit:786/0*A_-GpZLvqbxuVdkH 786w, https://miro.medium.com/v2/resize:fit:828/0*A_-GpZLvqbxuVdkH 828w, https://miro.medium.com/v2/resize:fit:1100/0*A_-GpZLvqbxuVdkH 1100w, https://miro.medium.com/v2/resize:fit:1400/0*A_-GpZLvqbxuVdkH 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="Visualization of the UDA graph representation of a One Piece character. The Character node in the graph is connected to a Devil Fruit node. The Devil Fruit node is connected to a Devil Fruit Type node." class="bh fu pi c" width="700" height="397" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="qt fc qu ov ow qv qw bf b bg ab ed"><em class="re">The graph representation of the onepiece: domain model from our UI. Depicted here you can see how Characters are related to Devil Fruit, and that each Devil Fruit has a type.</em></figcaption></figure><p id="82dc" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk"><strong class="oc in">Upper domain models are data</strong>. They are expressed as <a class="ag gz" href="https://www.w3.org/TR/rdf12-concepts/" rel="noopener ugc nofollow" target="_blank">conceptual RDF</a> and organized into named graphs, making them introspectable, queryable, and versionable within the UDA knowledge graph. This graph unifies not just the domain models themselves, but also the schemas they transpile to — GraphQL, Avro, Iceberg, Java — and the mappings that connect domain concepts to concrete data containers, such as GraphQL type resolvers served by a <a class="ag gz" rel="noopener ugc nofollow" href="https://netflixtechblog.com/open-sourcing-the-netflix-domain-graph-service-framework-graphql-for-spring-boot-92b9dcecda18" target="_blank" data-discover="true">Domain Graph Service</a>, <a class="ag gz" rel="noopener ugc nofollow" href="https://netflixtechblog.com/data-mesh-a-data-movement-and-processing-platform-netflix-1288bcab2873" target="_blank" data-discover="true">Data Mesh sources</a>, or Iceberg tables, through their representations. Upper raises the level of abstraction above traditional ontology languages: it defines a strict subset of <a class="ag gz" href="https://www.w3.org/2001/sw/wiki/Main_Page" rel="noopener ugc nofollow" target="_blank">semantic technologies</a> from the W3C tailored and generalized for domain modeling. It builds on ontology frameworks like RDFS, OWL, and SHACL so domain authors can model effectively without even needing to learn what an ontology is.</p><figure class="oy oz pa pb pc pd ov ow paragraph-image"><div role="button" tabindex="0" class="pe pf fi pg bh ph"><div class="ov ow rf"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*SGMUpJucEWhdlZsd4blz3A.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*SGMUpJucEWhdlZsd4blz3A.png" /><img alt="Screenshot of UDA UI showing domain model for One Piece serialized as Turtle." class="bh fu pi c" width="700" height="767" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="qt fc qu ov ow qv qw bf b bg ab ed">UDA domain model for One Piece. <a class="ag gz" href="https://github.com/Netflix-Skunkworks/uda/blob/9627a97fcd972a41ec910be3f928ea7692d38714/uda-intro-blog/onepiece.ttl" rel="noopener ugc nofollow" target="_blank">Link to full definition</a>.</figcaption></figure><p id="eed1" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk"><strong class="oc in">Upper is the metamodel for Connected Data in UDA — the model for all models</strong>. It is designed as a bootstrapping <a class="ag gz" href="https://en.wikipedia.org/wiki/Upper_ontology" rel="noopener ugc nofollow" target="_blank">upper ontology</a>, which means that Upper is <em class="rg">self-referencing</em>, because it models itself as a domain model; <em class="rg">self-describing</em>, because it defines the very concept of a domain model; and <em class="rg">self-validating</em>, because it conforms to its own model. This approach enables UDA to bootstrap its own infrastructure: Upper itself is projected into a generated Jena-based Java API and GraphQL schema used in GraphQL service federated into Netflix’s Enterprise GraphQL gateway. These same generated APIs are then used by the projections and the UI. Because all domain models are <a class="ag gz" href="https://en.wikipedia.org/wiki/Conservative_extension" rel="noopener ugc nofollow" target="_blank">conservative extensions</a> of Upper, other system domain models — including those for GraphQL, Avro, Data Mesh, and Mappings — integrate seamlessly into the same runtime, enabling consistent data semantics and interoperability across schemas.</p><figure class="oy oz pa pb pc pd ov ow paragraph-image"><div role="button" tabindex="0" class="pe pf fi pg bh ph"><div class="ov ow rh"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*5tJcW2A6lLrNi257%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*5tJcW2A6lLrNi257%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*5tJcW2A6lLrNi257%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*5tJcW2A6lLrNi257%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*5tJcW2A6lLrNi257%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*5tJcW2A6lLrNi257%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*5tJcW2A6lLrNi257%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*5tJcW2A6lLrNi257 640w, https://miro.medium.com/v2/resize:fit:720/0*5tJcW2A6lLrNi257 720w, https://miro.medium.com/v2/resize:fit:750/0*5tJcW2A6lLrNi257 750w, https://miro.medium.com/v2/resize:fit:786/0*5tJcW2A6lLrNi257 786w, https://miro.medium.com/v2/resize:fit:828/0*5tJcW2A6lLrNi257 828w, https://miro.medium.com/v2/resize:fit:1100/0*5tJcW2A6lLrNi257 1100w, https://miro.medium.com/v2/resize:fit:1400/0*5tJcW2A6lLrNi257 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="Screenshot of an IDE. It shows Java code using the generated API from the Upper metamodel to traverse and print terms from a domain domain in the top while the bottom contains the output of an execution." class="bh fu pi c" width="700" height="801" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="qt fc qu ov ow qv qw bf b bg ab ed">Traversing a domain model programmatically using the Java API generated from the Upper metamodel.</figcaption></figure><h2 id="67b9" class="qx ps im bf pt gi qy du gj gk qz dw gl gm ra gn go gp rb gq gr gs rc gt gu rd bk">Data Container Representations</h2><p id="168d" class="pw-post-body-paragraph oa ob im oc b od qn of og oh qo oj ok gm qp om on gp qq op oq gs qr os ot ou hn bk"><strong class="oc in">Data containers are repositories of information. </strong>They contain instance data that conform to their own schema languages or type systems: federated entities from GraphQL services, Avro records from Data Mesh sources, rows from Iceberg tables, or objects from Java APIs. Each container operates within the context of a system that imposes its own structural and operational constraints.</p><figure class="oy oz pa pb pc pd ov ow paragraph-image"><div role="button" tabindex="0" class="pe pf fi pg bh ph"><div class="ov ow ri"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*qUzAb6-TC2HL8qAWAW1Xlw.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*qUzAb6-TC2HL8qAWAW1Xlw.png" /><img alt="Screenshot of a UI showing details for a Data Mesh Source containing One Piece Characters." class="bh fu pi c" width="700" height="759" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="qt fc qu ov ow qv qw bf b bg ab ed">A Data Mesh source is a data container.</figcaption></figure><p id="707f" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk"><strong class="oc in">Data container </strong><a class="ag gz" href="https://en.wikipedia.org/wiki/Knowledge_representation_and_reasoning" rel="noopener ugc nofollow" target="_blank"><strong class="oc in">representations</strong></a><strong class="oc in"> are data.</strong> They are faithful interpretations of the members of data systems as graph data. UDA captures the definition of these systems as their own domain models, the system domains. These models encode both the information architecture of the systems and the schemas of the data containers within. They provide a blueprint for translating the systems into graph representations.</p></div></div><div class="pd"><div class="ac cb"><div class="nd rj ne rk nf rl cf rm cg rn ci bh"><figure class="oy oz pa pb pc pd ro rp paragraph-image"><div role="button" tabindex="0" class="pe pf fi pg bh ph"><div class="ov ow ox"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*6QzelmSRrIj1G881%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*6QzelmSRrIj1G881%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*6QzelmSRrIj1G881%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*6QzelmSRrIj1G881%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*6QzelmSRrIj1G881%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*6QzelmSRrIj1G881%201100w,%20https://miro.medium.com/v2/resize:fit:2000/format:webp/0*6QzelmSRrIj1G881%202000w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 1000px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*6QzelmSRrIj1G881 640w, https://miro.medium.com/v2/resize:fit:720/0*6QzelmSRrIj1G881 720w, https://miro.medium.com/v2/resize:fit:750/0*6QzelmSRrIj1G881 750w, https://miro.medium.com/v2/resize:fit:786/0*6QzelmSRrIj1G881 786w, https://miro.medium.com/v2/resize:fit:828/0*6QzelmSRrIj1G881 828w, https://miro.medium.com/v2/resize:fit:1100/0*6QzelmSRrIj1G881 1100w, https://miro.medium.com/v2/resize:fit:2000/0*6QzelmSRrIj1G881 2000w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 1000px" /><img alt="Screenshot of an IDE showing two files open side by side. On the left is a system domain model for Data Mesh. On the right is a representation of a Data Mesh source containing One Piece Character data." class="bh fu pi c" width="1000" height="604" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="qt fc qu ov ow qv qw bf b bg ab ed"><em class="re">Side by side/super imposed image of data container schema and representation. </em><a class="ag gz" href="https://github.com/Netflix-Skunkworks/uda/blob/9627a97fcd972a41ec910be3f928ea7692d38714/uda-intro-blog/onepiece_character_data_container.ttl" rel="noopener ugc nofollow" target="_blank"><em class="re">Link to full data container representation</em></a><em class="re">.</em></figcaption></figure></div></div></div><div class="ac cb"><div class="ci bh ht hu hv hw"><p id="c76d" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk"><strong class="oc in">UDA catalogs the data container representations into the knowledge graph.</strong> It records the coordinates and metadata of the underlying data assets, but unlike a traditional catalog, it only tracks assets that are semantically connected to domain models. This enables users and systems to connect concepts from domain models to the concrete locations where corresponding instance data can be accessed. Those connections are called <em class="rg">Mappings</em>.</p><h2 id="31d6" class="qx ps im bf pt gi qy du gj gk qz dw gl gm ra gn go gp rb gq gr gs rc gt gu rd bk">Mappings</h2><p id="886e" class="pw-post-body-paragraph oa ob im oc b od qn of og oh qo oj ok gm qp om on gp qq op oq gs qr os ot ou hn bk"><strong class="oc in">Mappings are data that connect domain models to data containers.</strong> Every element in a domain model is addressable, from the domain model itself down to specific attributes and relationships. Likewise, data container representations make all components addressable, from an Iceberg table to an individual column, or from a GraphQL type to a specific field. A Mapping connects nodes in a subgraph of the domain model to nodes in a subgraph of a container representation. Visually, the Mapping is the set of arcs that link those two graphs together.</p><figure class="oy oz pa pb pc pd ov ow paragraph-image"><div role="button" tabindex="0" class="pe pf fi pg bh ph"><div class="ov ow rq"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*it3X5Vu8plWX5QvN_AJkgw.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*it3X5Vu8plWX5QvN_AJkgw.png" /><img alt="Screenshot of UDA UI showing a mapping between a concept in UDA and a Data Mesh Source." class="bh fu pi c" width="700" height="695" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="qt fc qu ov ow qv qw bf b bg ab ed"><em class="re">A mapping between a domain model and a Data Mesh Source from the UDA UI. </em><a class="ag gz" href="https://github.com/Netflix-Skunkworks/uda/blob/9627a97fcd972a41ec910be3f928ea7692d38714/uda-intro-blog/onepiece_character_mappings.ttl" rel="noopener ugc nofollow" target="_blank"><em class="re">Link to full mapping</em></a><em class="re">.</em></figcaption></figure><p id="900c" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk"><strong class="oc in">Mappings enable discovery.</strong> Starting from a domain concept, users and systems can walk the knowledge graph to find where that concept is materialized — in which data system, in which container, and even how a specific attribute or relationship is physically accessed. The inverse is also supported: given a data container, one can trace back to the domain concepts it participates in.</p><p id="d391" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk"><strong class="oc in">Mappings shape UDA’s approach to semantic data integration.</strong> Most existing schema languages are not expressive enough in capturing richer semantics of a domain to address requirements for data integration (<a class="ag gz" href="https://doi.org/10.1007/978-3-319-49340-4_8" rel="noopener ugc nofollow" target="_blank">for example</a>, “accessibility of data, providing semantic context to support its interpretation, and establishing meaningful links between data”). A trivial example of this could be seen in the lack of built-in facilities in Avro to represent foreign keys, making it very hard to express how entities relate across Data Mesh sources. Mappings, together with the corresponding system domain models, allow for such relationships, and many other constraints, to be defined in the domain models and used programmatically in actual data systems.</p><p id="936f" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk"><strong class="oc in">Mappings enable intent-based automation.</strong> Data is not always available in the systems where consumers need it. Because Mappings encode both meaning and location, UDA can reason about how data should move, preserving semantics, without requiring the consumer to specify how it should be done. Beyond the cataloging use case, connecting to existing containers, UDA automatically derives <em class="rg">canonical Mappings</em> from registered domain models as part of the projection process.</p><h2 id="4cec" class="qx ps im bf pt gi qy du gj gk qz dw gl gm ra gn go gp rb gq gr gs rc gt gu rd bk">Projections</h2><p id="b219" class="pw-post-body-paragraph oa ob im oc b od qn of og oh qo oj ok gm qp om on gp qq op oq gs qr os ot ou hn bk"><strong class="oc in">A projection produces a concrete data container.</strong> These containers, such as a GraphQL schema or a Data Mesh source, implement the characteristics derived from a registered domain model. Each projection is a concrete realization of Upper’s denotational semantics, ensuring <a class="ag gz" href="https://en.wikipedia.org/wiki/Semantic_interoperability" rel="noopener ugc nofollow" target="_blank">semantic interoperability</a> across all containers projected from the same domain model.</p><p id="c0d9" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk"><strong class="oc in">Projections produce consistent public contracts across systems.</strong> The data containers generated by projections encode data contracts in the form of schemas, derived by transpiling a domain model into the target container’s schema language. UDA currently supports transpilation to GraphQL and Avro schemas.</p><p id="10a4" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk">The GraphQL transpilation produces a schema that adheres to the <a class="ag gz" href="https://spec.graphql.org/October2021/#sec-Overview" rel="noopener ugc nofollow" target="_blank">official GraphQL spec</a> with the ability to generate all GraphQL types defined in the spec. Given that the UDA domain model can be federated, it also supports generating federated graphQL schemas. Below is an example of a transpiled GraphQL schema.</p></div></div><div class="pd"><div class="ac cb"><div class="nd rj ne rk nf rl cf rm cg rn ci bh"><figure class="oy oz pa pb pc pd ro rp paragraph-image"><div role="button" tabindex="0" class="pe pf fi pg bh ph"><div class="ov ow ox"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*NPXB3ujnUGSIklei%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*NPXB3ujnUGSIklei%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*NPXB3ujnUGSIklei%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*NPXB3ujnUGSIklei%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*NPXB3ujnUGSIklei%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*NPXB3ujnUGSIklei%201100w,%20https://miro.medium.com/v2/resize:fit:2000/format:webp/0*NPXB3ujnUGSIklei%202000w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 1000px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*NPXB3ujnUGSIklei 640w, https://miro.medium.com/v2/resize:fit:720/0*NPXB3ujnUGSIklei 720w, https://miro.medium.com/v2/resize:fit:750/0*NPXB3ujnUGSIklei 750w, https://miro.medium.com/v2/resize:fit:786/0*NPXB3ujnUGSIklei 786w, https://miro.medium.com/v2/resize:fit:828/0*NPXB3ujnUGSIklei 828w, https://miro.medium.com/v2/resize:fit:1100/0*NPXB3ujnUGSIklei 1100w, https://miro.medium.com/v2/resize:fit:2000/0*NPXB3ujnUGSIklei 2000w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 1000px" /><img alt="Screenshot of an IDE showing two files open side by side. On the left is the definition of a Character in UDA. On the right is transpiled GraphQL schema." class="bh fu pi c" width="1000" height="424" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="qt fc qu ov ow qv qw bf b bg ab ed"><em class="re">Domain model on the left, with transpiled GraphQL schema on the right. </em><a class="ag gz" href="https://github.com/Netflix-Skunkworks/uda/blob/9627a97fcd972a41ec910be3f928ea7692d38714/uda-intro-blog/onepiece.graphqls" rel="noopener ugc nofollow" target="_blank"><em class="re">Link to full transpiled GraphQL schema</em></a><em class="re">.</em></figcaption></figure></div></div></div><div class="ac cb"><div class="ci bh ht hu hv hw"><p id="252b" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk">The Avro transpilation produces a schema that is a Data Mesh flavor of Avro, which includes some customization on top of the <a class="ag gz" href="https://avro.apache.org/docs/1.12.0/specification/" rel="noopener ugc nofollow" target="_blank">official Avro spec</a>. This schema is used to automatically create a Data Mesh source container. Below is an example of a transpiled Avro schema.</p></div></div><div class="pd"><div class="ac cb"><div class="nd rj ne rk nf rl cf rm cg rn ci bh"><figure class="oy oz pa pb pc pd ro rp paragraph-image"><div role="button" tabindex="0" class="pe pf fi pg bh ph"><div class="ov ow ox"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*uVInkj5S3PYTqNA-%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*uVInkj5S3PYTqNA-%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*uVInkj5S3PYTqNA-%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*uVInkj5S3PYTqNA-%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*uVInkj5S3PYTqNA-%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*uVInkj5S3PYTqNA-%201100w,%20https://miro.medium.com/v2/resize:fit:2000/format:webp/0*uVInkj5S3PYTqNA-%202000w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 1000px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*uVInkj5S3PYTqNA- 640w, https://miro.medium.com/v2/resize:fit:720/0*uVInkj5S3PYTqNA- 720w, https://miro.medium.com/v2/resize:fit:750/0*uVInkj5S3PYTqNA- 750w, https://miro.medium.com/v2/resize:fit:786/0*uVInkj5S3PYTqNA- 786w, https://miro.medium.com/v2/resize:fit:828/0*uVInkj5S3PYTqNA- 828w, https://miro.medium.com/v2/resize:fit:1100/0*uVInkj5S3PYTqNA- 1100w, https://miro.medium.com/v2/resize:fit:2000/0*uVInkj5S3PYTqNA- 2000w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 1000px" /><img alt="Screenshot of an IDE showing two files open side by side. On the left is the definition of a Devil Fruit in UDA. On the right is transpiled Avro schema." class="bh fu pi c" width="1000" height="476" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="qt fc qu ov ow qv qw bf b bg ab ed"><em class="re">Domain model on the left, with transpiled Avro schema on the right. </em><a class="ag gz" href="https://github.com/Netflix-Skunkworks/uda/blob/9627a97fcd972a41ec910be3f928ea7692d38714/uda-intro-blog/onepiece.avro" rel="noopener ugc nofollow" target="_blank"><em class="re">Link to full transpiled Avro schema</em></a><em class="re">.</em></figcaption></figure></div></div></div><div class="ac cb"><div class="ci bh ht hu hv hw"><p id="6ead" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk"><strong class="oc in">Projections can automatically populate data containers. </strong>Some projections, such as those to GraphQL schemas or Data Mesh sources produce empty containers that require developers to populate the data. This might be creating GraphQL APIs or pushing events onto Data Mesh sources. Conversely, other containers, like Iceberg Tables, are automatically created and populated by UDA. For Iceberg Tables, UDA leverages the Data Mesh platform to automatically create data streams to move data into tables. This process utilizes much of the same infrastructure detailed in this blog post <a class="ag gz" rel="noopener ugc nofollow" href="https://netflixtechblog.com/data-movement-in-netflix-studio-via-data-mesh-3fddcceb1059" target="_blank" data-discover="true">here</a>.</p><p id="bf20" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk"><strong class="oc in">Projections have mappings. </strong>UDA automatically generates and manages mappings between the newly created data containers and the projected domain model.</p><h1 id="240f" class="pr ps im bf pt pu pv pw gj px py pz gl qa qb qc qd qe qf qg qh qi qj qk ql qm bk">Early Adopters</h1><h2 id="76f7" class="qx ps im bf pt gi qy du gj gk qz dw gl gm ra gn go gp rb gq gr gs rc gt gu rd bk">Controlled Vocabularies (PDM)</h2><p id="da3d" class="pw-post-body-paragraph oa ob im oc b od qn of og oh qo oj ok gm qp om on gp qq op oq gs qr os ot ou hn bk">The full range of Netflix’s business activities relies on a sprawling data model that captures the details of our many business processes. Teams need to be able to coordinate operational activities to ensure that content production is complete, advertising campaigns are in place, and promotional assets are ready to deploy. We implicitly depend upon a singular definition of shared concepts, such as content production is complete. Multiple definitions create coordination challenges. Software (and humans) don’t know that the definitions mean the same thing.</p><p id="800e" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk">We started the Primary Data Management (PDM) initiative to create unified and consistent definitions for the core concepts in our data model. These definitions form <strong class="oc in">controlled vocabularies</strong>, standardized and governed lists for what values are permitted within certain fields in our data model.</p><p id="be83" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk"><strong class="oc in">Primary Data Management (PDM) is a single place where business users can manage controlled vocabularies. </strong>Our data model governance has been scattered across different tools and teams creating coordination challenges. This is an information management problem relating to the definition, maintenance and consistent use of reference data and taxonomies. This problem is not unique to Netflix, so we looked outward for existing solutions to this problem.</p><figure class="oy oz pa pb pc pd ov ow paragraph-image"><div role="button" tabindex="0" class="pe pf fi pg bh ph"><div class="ov ow rr"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*GJMad4GU29YxPONf%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*GJMad4GU29YxPONf%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*GJMad4GU29YxPONf%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*GJMad4GU29YxPONf%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*GJMad4GU29YxPONf%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*GJMad4GU29YxPONf%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*GJMad4GU29YxPONf%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*GJMad4GU29YxPONf 640w, https://miro.medium.com/v2/resize:fit:720/0*GJMad4GU29YxPONf 720w, https://miro.medium.com/v2/resize:fit:750/0*GJMad4GU29YxPONf 750w, https://miro.medium.com/v2/resize:fit:786/0*GJMad4GU29YxPONf 786w, https://miro.medium.com/v2/resize:fit:828/0*GJMad4GU29YxPONf 828w, https://miro.medium.com/v2/resize:fit:1100/0*GJMad4GU29YxPONf 1100w, https://miro.medium.com/v2/resize:fit:1400/0*GJMad4GU29YxPONf 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="Screenshot of PDM UI" class="bh fu pi c" width="700" height="420" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="qt fc qu ov ow qv qw bf b bg ab ed">Managing the taxonomy of One Piece characters in PDM.</figcaption></figure><p id="3a7b" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk"><strong class="oc in">PDM uses the Simple Knowledge Organization System (</strong><a class="ag gz" href="https://www.w3.org/TR/skos-primer" rel="noopener ugc nofollow" target="_blank"><strong class="oc in">SKOS</strong></a><strong class="oc in">)</strong> <strong class="oc in">model</strong>. It is a W3C data standard designed for modeling knowledge. Its terminology is abstract, with Concepts that can be organized into ConceptSchemes and properties to describe various types of relationships. Every system is hardcoded against <em class="rg">something</em>, that’s how software knows how to manipulate data. We want a system that can work with a data model as its input, so we still need <em class="rg">something</em> concrete to build the software against. This is what SKOS provides, a generic basis for modeling knowledge that our system can understand.</p><p id="cd41" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk"><strong class="oc in">PDM uses Domain Models to integrate SKOS into the rest of Content Engineering’s ecosystem. </strong>A core premise of the system is that it takes a domain model as input, and everything that <em class="rg">can</em> be derived <em class="rg">is</em> derived from that model. PDM builds a user interface based upon the model definition and leverages UDA to project this model into type-safe interfaces for other systems to use. The system will provision a Domain Graph Service (DGS) within our federated GraphQL API environment using a GraphQL schema that UDA projects from the domain model. UDA is also used to provision data movement pipelines which are able to feed our <a class="ag gz" rel="noopener ugc nofollow" href="https://netflixtechblog.com/how-netflix-content-engineering-makes-a-federated-graph-searchable-5c0c1c7d7eaf" target="_blank" data-discover="true">GraphSearch</a> infrastructure as well as move data into the warehouse. The data movement systems use Avro schemas, and UDA creates a projection from the domain model to Avro.</p><p id="2309" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk"><strong class="oc in">Consumers of controlled vocabularies never know they’re using SKOS. </strong>Domain models use terms that fit in with the domain. SKOS’s generic notion of <em class="rg">broader</em> and <em class="rg">narrower</em> to define a hierarchy are hidden from consumers as super-properties within the model. This allows consumers to work with language that is familiar to them while enabling PDM to work with any model. The best of both worlds.</p><h2 id="47a4" class="qx ps im bf pt gi qy du gj gk qz dw gl gm ra gn go gp rb gq gr gs rc gt gu rd bk">Operational Reporting (Sphere)</h2><p id="cb2d" class="pw-post-body-paragraph oa ob im oc b od qn of og oh qo oj ok gm qp om on gp qq op oq gs qr os ot ou hn bk"><strong class="oc in">Operational reporting serves the detailed day-to-day activities and processes of a business domain.</strong> It is a reporting paradigm specialized in covering high-resolution, low-latency data sets.</p><p id="1d94" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk"><strong class="oc in">Operational reporting systems should generate reports without relying on technical intermediaries. </strong>Operational reporting systems need to address the persistent challenge of empowering business users to explore and obtain the data they need, when they need it. Without such self-service systems, requests for new reports or data extracts often result in back-and-forth exchanges, where the initial query may not exactly meet business users’ expectations, requiring further clarification and refinement.</p><p id="905e" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk"><strong class="oc in">Data discovery and query generation are two relevant aspects of data integration. </strong>Supplying end-users with an accurate, contextual, and user-friendly data discovery experience provides a basis for query generation mechanism which produces syntactically correct and semantically reliable queries.</p><p id="0968" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk"><strong class="oc in">Operational reports are predominantly run on data hydrated from GraphQL services into the Data Warehouse. </strong>You can read about our journey from conventional data movement to streaming data pipelines based on CDC and GraphQL hydration in <a class="ag gz" rel="noopener ugc nofollow" href="https://netflixtechblog.com/data-movement-in-netflix-studio-via-data-mesh-3fddcceb1059" target="_blank" data-discover="true">this blog post</a>. Among the challenging byproducts of this approach was that a single, distinct data concept is now present in two places (GraphQL and data warehouse), with some disparity in semantic context to guide and support the interpretations and connectivity of that data. To address this, we formulate a mechanism to use the syntax and semantics captured in the federated schema from <a class="ag gz" rel="noopener ugc nofollow" href="https://netflixtechblog.com/how-netflix-scales-its-api-with-graphql-federation-part-1-ae3557c187e2" target="_blank" data-discover="true">Netflix’s Enterprise GraphQL</a> and populate <em class="rg">representational domain models</em> in UDA to preserve those details and add more.</p><p id="28bb" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk"><strong class="oc in">Domain models enable the data discovery experience. </strong>Metadata aggregated from various data-producing systems is captured in UDA domain models using a unified vocabulary. This metadata is surfaced for the users’ search and discovery needs; instead of specifying exact tables and join keys, users simply can search for familiar business concepts such as ‘actors’ or ‘movies’. We use UDA models to disambiguate and resolve the intended concepts and their related data entities.</p><p id="7eee" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk"><strong class="oc in">UDA knowledge graph is the data landscape for query generation. </strong>Once concepts are discovered and their mappings to corresponding data containers are identified and located in the knowledge graph, we use them to establish join strategies. Through graph traversal, we identify <em class="rg">boundaries</em> and <em class="rg">islands</em> within the data landscape. This ensures only feasible, joinable combinations are selected while weeding out semantically incorrect and non-executable query candidates.</p><figure class="oy oz pa pb pc pd ov ow paragraph-image"><div role="button" tabindex="0" class="pe pf fi pg bh ph"><div class="ov ow ox"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*EFEfzwY-3Tb6521O%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*EFEfzwY-3Tb6521O%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*EFEfzwY-3Tb6521O%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*EFEfzwY-3Tb6521O%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*EFEfzwY-3Tb6521O%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*EFEfzwY-3Tb6521O%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*EFEfzwY-3Tb6521O%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*EFEfzwY-3Tb6521O 640w, https://miro.medium.com/v2/resize:fit:720/0*EFEfzwY-3Tb6521O 720w, https://miro.medium.com/v2/resize:fit:750/0*EFEfzwY-3Tb6521O 750w, https://miro.medium.com/v2/resize:fit:786/0*EFEfzwY-3Tb6521O 786w, https://miro.medium.com/v2/resize:fit:828/0*EFEfzwY-3Tb6521O 828w, https://miro.medium.com/v2/resize:fit:1100/0*EFEfzwY-3Tb6521O 1100w, https://miro.medium.com/v2/resize:fit:1400/0*EFEfzwY-3Tb6521O 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="Screenshot of Sphere’s UI" class="bh fu pi c" width="700" height="353" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="qt fc qu ov ow qv qw bf b bg ab ed">Generating a report in Sphere.</figcaption></figure><p id="107a" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk"><strong class="oc in">Sphere is a UDA-powered self-service operational reporting system. </strong>The solution based on knowledge graphs described above is called Sphere. Seeing self-service operational reporting through this lens, we can improve business users’ agency in access to operational data. They are empowered to explore, assemble, and refine reports at the conceptual level, while technical complexities are managed by the system.</p><h1 id="76ef" class="pr ps im bf pt pu pv pw gj px py pz gl qa qb qc qd qe qf qg qh qi qj qk ql qm bk">Stay Tuned</h1><p id="844c" class="pw-post-body-paragraph oa ob im oc b od qn of og oh qo oj ok gm qp om on gp qq op oq gs qr os ot ou hn bk">UDA marks a fundamental shift in how we approach data modeling within Content Engineering. By providing a unified knowledge graph composed of what we know about our various data systems and the business concepts within them, we’ve made information more consistent, connected, and discoverable across our organization. We’re excited about future applications of these ideas such as:</p><ul class=""><li id="8106" class="oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou pj pk pl bk">Supporting additional projections like Protobuf/gRPC</li><li id="9af0" class="oa ob im oc b od pm of og oh pn oj ok gm po om on gp pp op oq gs pq os ot ou pj pk pl bk">Materializing the knowledge graph of instance data for querying, profiling, and management</li><li id="ec0a" class="oa ob im oc b od pm of og oh pn oj ok gm po om on gp pp op oq gs pq os ot ou pj pk pl bk">Finally solving some of the initial <a class="ag gz" rel="noopener ugc nofollow" href="https://netflixtechblog.com/how-netflix-content-engineering-makes-a-federated-graph-searchable-5c0c1c7d7eaf" target="_blank" data-discover="true">challenges</a> posed by Graph Search (that actually inspired some of this work)</li></ul><p id="2012" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk">If you’re interested in this space, we’d love to connect — whether you’re exploring new roles down the road or just want to swap ideas.</p><p id="844f" class="pw-post-body-paragraph oa ob im oc b od oe of og oh oi oj ok gm ol om on gp oo op oq gs or os ot ou hn bk">Expect to see future blog posts exploring PDM and Sphere in more detail soon!</p><h2 id="e536" class="qx ps im bf pt gi qy du gj gk qz dw gl gm ra gn go gp rb gq gr gs rc gt gu rd bk">Credits</h2><p id="b15e" class="pw-post-body-paragraph oa ob im oc b od qn of og oh qo oj ok gm qp om on gp qq op oq gs qr os ot ou hn bk">Thanks to <a class="ag gz" href="https://www.linkedin.com/in/andreaslegenbauer/" rel="noopener ugc nofollow" target="_blank">Andreas Legenbauer</a>, <a class="ag gz" href="https://www.linkedin.com/in/bernardo-g-4414b41/" rel="noopener ugc nofollow" target="_blank">Bernardo Gomez Palacio Valdes</a>, <a class="ag gz" href="https://www.linkedin.com/in/czhao/" rel="noopener ugc nofollow" target="_blank">Charles Zhao</a>, <a class="ag gz" href="https://www.linkedin.com/in/christopherchonguw/" rel="noopener ugc nofollow" target="_blank">Christopher Chong</a>, <a class="ag gz" href="https://www.linkedin.com/in/deepa-krishnan-593b60/" rel="noopener ugc nofollow" target="_blank">Deepa Krishnan</a>, <a class="ag gz" href="https://www.linkedin.com/in/gpesma/" rel="noopener ugc nofollow" target="_blank">George Pesmazoglou</a>, <a class="ag gz" href="https://www.linkedin.com/in/jsilvax/" rel="noopener ugc nofollow" target="_blank">Jessica Silva</a>, <a class="ag gz" href="https://www.linkedin.com/in/katherine-anderson-77074159/" rel="noopener ugc nofollow" target="_blank">Katherine Anderson</a>, <a class="ag gz" href="https://www.linkedin.com/in/malikday/" rel="noopener ugc nofollow" target="_blank">Malik Day</a>, <a class="ag gz" href="https://www.linkedin.com/in/ritabogdanovashapkina/" rel="noopener ugc nofollow" target="_blank">Rita Bogdanova</a>, <a class="ag gz" href="https://www.linkedin.com/in/ruoyunzheng/" rel="noopener ugc nofollow" target="_blank">Ruoyun Zheng</a>, <a class="ag gz" href="https://www.linkedin.com/in/shawn-s-b80821b0/" rel="noopener ugc nofollow" target="_blank">Shawn Stedman</a>, <a class="ag gz" href="https://www.linkedin.com/in/suchitagoyal/" rel="noopener ugc nofollow" target="_blank">Suchita Goyal</a>, <a class="ag gz" href="http://www.linkedin.com/in/utkarshshrivastava/" rel="noopener ugc nofollow" target="_blank">Utkarsh Shrivastava</a>, <a class="ag gz" href="https://www.linkedin.com/in/yoomikoh/" rel="noopener ugc nofollow" target="_blank">Yoomi Koh</a>, <a class="ag gz" href="https://www.linkedin.com/in/yuliashmeleva/" rel="noopener ugc nofollow" target="_blank">Yulia Shmeleva</a></p></div></div></div>]]></description>
      <link>https://netflixtechblog.com/uda-unified-data-architecture-6a6aee261d8d</link>
      <guid>https://netflixtechblog.com/uda-unified-data-architecture-6a6aee261d8d</guid>
      <pubDate>Thu, 12 Jun 2025 16:56:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[FM-Intent: Predicting User Session Intent with Hierarchical Multi-Task Learning]]></title>
      <description><![CDATA[<div><div></div><p id="e0a5" class="pw-post-body-paragraph ob oc io od b oe of og oh oi oj ok ol go om on oo gr op oq or gu os ot ou ov hp bk">Authors: <a class="ag hb" href="https://www.linkedin.com/in/sejoon-oh/" rel="noopener ugc nofollow" target="_blank">Sejoon Oh</a>, <a class="ag hb" href="https://www.linkedin.com/in/moumitab/" rel="noopener ugc nofollow" target="_blank">Moumita Bhattacharya</a>, <a class="ag hb" href="https://www.linkedin.com/in/yesufeng/" rel="noopener ugc nofollow" target="_blank">Yesu Feng</a>, <a class="ag hb" href="https://www.linkedin.com/in/sudarshanlamkhede/" rel="noopener ugc nofollow" target="_blank">Sudarshan Lamkhede</a>, <a class="ag hb" href="https://www.linkedin.com/in/markhsiao/" rel="noopener ugc nofollow" target="_blank">Ko-Jen Hsiao</a>, and <a class="ag hb" href="https://www.linkedin.com/in/jbasilico/" rel="noopener ugc nofollow" target="_blank">Justin Basilico</a></p><h1 id="115a" class="ow ox io bf oy oz pa pb gl pc pd pe gn pf pg ph pi pj pk pl pm pn po pp pq pr bk">Motivation</h1><p id="cb4a" class="pw-post-body-paragraph ob oc io od b oe ps og oh oi pt ok ol go pu on oo gr pv oq or gu pw ot ou ov hp bk">Recommender systems have become essential components of digital services across e-commerce, streaming media, and social networks [1, 2]. At Netflix, these systems drive significant product and business impact by connecting members with relevant content at the right time [3, 4]. While our recommendation <strong class="od ip">foundation model (FM)</strong> has made substantial progress in understanding user preferences through large-scale learning from interaction histories (please refer to this <a class="ag hb" href="https://netflixtechblog.medium.com/foundation-model-for-personalized-recommendation-1a0bd8e02d39" rel="noopener"><strong class="od ip"><em class="px">article</em></strong></a> about FM @ Netflix), there is an opportunity to further enhance its capabilities. By extending FM to incorporate the prediction of underlying user intents, we aim to enrich its understanding of user sessions beyond next-item prediction, thereby offering a more comprehensive and nuanced recommendation experience.</p><p id="1371" class="pw-post-body-paragraph ob oc io od b oe of og oh oi oj ok ol go om on oo gr op oq or gu os ot ou ov hp bk">Recent research has highlighted the importance of understanding user intent in online platforms [5, 6, 7, 8]. As Xia et al. [8] demonstrated at Pinterest, predicting a user’s future intent can lead to more accurate and personalized recommendations. However, existing intent prediction approaches typically employ simple multi-task learning that adds intent prediction heads to next-item prediction models without establishing a hierarchical relationship between these tasks.</p><p id="62f1" class="pw-post-body-paragraph ob oc io od b oe of og oh oi oj ok ol go om on oo gr op oq or gu os ot ou ov hp bk">To address these limitations, we introduce <strong class="od ip"><em class="px">FM-Intent</em></strong>, a novel recommendation model that enhances our foundation model through hierarchical multi-task learning. FM-Intent captures a user’s latent session intent using both short-term and long-term implicit signals as proxies, then leverages this intent prediction to improve next-item recommendations. Unlike conventional approaches, FM-Intent establishes a clear hierarchy where intent predictions directly inform item recommendations, creating a more coherent and effective recommendation pipeline.</p><p id="16e8" class="pw-post-body-paragraph ob oc io od b oe of og oh oi oj ok ol go om on oo gr op oq or gu os ot ou ov hp bk">FM-Intent makes three key contributions:</p><ol class=""><li id="9b16" class="ob oc io od b oe of og oh oi oj ok ol go om on oo gr op oq or gu os ot ou ov py pz qa bk">A novel recommendation model that captures user intent on the Netflix platform and enhances next-item prediction using this intent information.</li><li id="538d" class="ob oc io od b oe qb og oh oi qc ok ol go qd on oo gr qe oq or gu qf ot ou ov py pz qa bk">A hierarchical multi-task learning approach that effectively models both short-term and long-term user interests.</li><li id="c154" class="ob oc io od b oe qb og oh oi qc ok ol go qd on oo gr qe oq or gu qf ot ou ov py pz qa bk">Comprehensive experimental validation showing significant performance improvements over state-of-the-art models, including our foundation model.</li></ol><h1 id="04a3" class="ow ox io bf oy oz pa pb gl pc pd pe gn pf pg ph pi pj pk pl pm pn po pp pq pr bk">Understanding User Intent in Netflix</h1><p id="d624" class="pw-post-body-paragraph ob oc io od b oe ps og oh oi pt ok ol go pu on oo gr pv oq or gu pw ot ou ov hp bk">In the Netflix ecosystem, user intent manifests through various interaction metadata, as illustrated in Figure 1. FM-Intent leverages these implicit signals to predict both user intent and next-item recommendations.</p><figure class="qj qk ql qm qn qo qg qh paragraph-image"><div role="button" tabindex="0" class="qp qq fl qr bh qs"><div class="qg qh qi"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*3pMgS5u3TepefPLB%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*3pMgS5u3TepefPLB%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*3pMgS5u3TepefPLB%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*3pMgS5u3TepefPLB%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*3pMgS5u3TepefPLB%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*3pMgS5u3TepefPLB%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*3pMgS5u3TepefPLB%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*3pMgS5u3TepefPLB 640w, https://miro.medium.com/v2/resize:fit:720/0*3pMgS5u3TepefPLB 720w, https://miro.medium.com/v2/resize:fit:750/0*3pMgS5u3TepefPLB 750w, https://miro.medium.com/v2/resize:fit:786/0*3pMgS5u3TepefPLB 786w, https://miro.medium.com/v2/resize:fit:828/0*3pMgS5u3TepefPLB 828w, https://miro.medium.com/v2/resize:fit:1100/0*3pMgS5u3TepefPLB 1100w, https://miro.medium.com/v2/resize:fit:1400/0*3pMgS5u3TepefPLB 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh fw qt c" width="700" height="329" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="cd73" class="pw-post-body-paragraph ob oc io od b oe of og oh oi oj ok ol go om on oo gr op oq or gu os ot ou ov hp bk"><em class="px">Figure 1: Overview of user engagement data in Netflix. User intent can be associated with several interaction metadata. We leverage various implicit signals to predict user intent and next-item.</em></p><p id="798c" class="pw-post-body-paragraph ob oc io od b oe of og oh oi oj ok ol go om on oo gr op oq or gu os ot ou ov hp bk">In Netflix, there can be multiple types of user intents. For instance,</p><blockquote class="qu qv qw"><p id="78cd" class="ob oc px od b oe of og oh oi oj ok ol go om on oo gr op oq or gu os ot ou ov hp bk"><strong class="od ip"><em class="io">Action Type</em></strong>: Categories reflecting what users intend to do on Netflix, such as discovering new content versus continuing previously started content. For example, when a member plays a follow-up episode of something they were already watching, this can be categorized as “continue watching” intent.</p><p id="5e27" class="ob oc px od b oe of og oh oi oj ok ol go om on oo gr op oq or gu os ot ou ov hp bk"><strong class="od ip"><em class="io">Genre Preference</em></strong>: The pre-defined genre labels (e.g., Action, Thriller, Comedy) that indicate a user’s content preferences during a session. These preferences can shift significantly between sessions, even for the same user.</p><p id="39f6" class="ob oc px od b oe of og oh oi oj ok ol go om on oo gr op oq or gu os ot ou ov hp bk"><strong class="od ip"><em class="io">Movie/Show Type</em></strong>: Whether a user is looking for a movie (typically a single, longer viewing experience) or a TV show (potentially multiple episodes of shorter duration).</p><p id="ebf9" class="ob oc px od b oe of og oh oi oj ok ol go om on oo gr op oq or gu os ot ou ov hp bk"><strong class="od ip"><em class="io">Time-since-release</em></strong>: Whether the user prefers newly released content, recent content (e.g., between a week and a month), or evergreen catalog titles.</p></blockquote><p id="cff1" class="pw-post-body-paragraph ob oc io od b oe of og oh oi oj ok ol go om on oo gr op oq or gu os ot ou ov hp bk">These dimensions serve as proxies for the latent user intent, which is often not directly observable but crucial for providing relevant recommendations.</p><h1 id="9ea0" class="ow ox io bf oy oz pa pb gl pc pd pe gn pf pg ph pi pj pk pl pm pn po pp pq pr bk">FM-Intent Model Architecture</h1><p id="30e5" class="pw-post-body-paragraph ob oc io od b oe ps og oh oi pt ok ol go pu on oo gr pv oq or gu pw ot ou ov hp bk">FM-Intent employs a hierarchical multi-task learning approach with three major components, as illustrated in Figure 2.</p><figure class="qj qk ql qm qn qo qg qh paragraph-image"><div class="qg qh qx"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*TO3z9jPiu2QZR-xnL7erMQ.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*TO3z9jPiu2QZR-xnL7erMQ.png" /><img alt="" class="bh fw qt c" width="624" height="830" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></figure><p id="0f7c" class="pw-post-body-paragraph ob oc io od b oe of og oh oi oj ok ol go om on oo gr op oq or gu os ot ou ov hp bk"><em class="px">Figure 2: An architectural illustration of our hierarchical multi-task learning model FM-Intent for user intent and item predictions. We use ground-truth intent and item-ID labels to optimize predictions.</em></p><h2 id="2f73" class="qy ox io bf oy gk qz dy gl gm ra ea gn go rb gp gq gr rc gs gt gu rd gv gw re bk">1. Input Feature Sequence Formation</h2><p id="852b" class="pw-post-body-paragraph ob oc io od b oe ps og oh oi pt ok ol go pu on oo gr pv oq or gu pw ot ou ov hp bk">The first component constructs rich input features by combining interaction metadata. The input feature for each interaction combines categorical embeddings and numerical features, creating a comprehensive representation of user behavior.</p><h2 id="be50" class="qy ox io bf oy gk qz dy gl gm ra ea gn go rb gp gq gr rc gs gt gu rd gv gw re bk">2. User Intent Prediction</h2><p id="754b" class="pw-post-body-paragraph ob oc io od b oe ps og oh oi pt ok ol go pu on oo gr pv oq or gu pw ot ou ov hp bk">The intent prediction component processes the input feature sequence through a Transformer encoder and generates predictions for multiple intent signals.</p><p id="7364" class="pw-post-body-paragraph ob oc io od b oe of og oh oi oj ok ol go om on oo gr op oq or gu os ot ou ov hp bk">The Transformer encoder effectively models the long-term interest of users through multi-head attention mechanisms. For each prediction task, the intent encoding is transformed into prediction scores via fully-connected layers.</p><p id="eefb" class="pw-post-body-paragraph ob oc io od b oe of og oh oi oj ok ol go om on oo gr op oq or gu os ot ou ov hp bk">A key innovation in FM-Intent is the attention-based aggregation of individual intent predictions. This approach generates a comprehensive intent embedding that captures the relative importance of different intent signals for each user, providing valuable insights for personalization and explanation.</p><h2 id="19c9" class="qy ox io bf oy gk qz dy gl gm ra ea gn go rb gp gq gr rc gs gt gu rd gv gw re bk">3. Next-Item Prediction with Hierarchical Multi-Task Learning</h2><p id="6faa" class="pw-post-body-paragraph ob oc io od b oe ps og oh oi pt ok ol go pu on oo gr pv oq or gu pw ot ou ov hp bk">The final component combines the input features with the user intent embedding to make more accurate next-item recommendations.</p><p id="32a2" class="pw-post-body-paragraph ob oc io od b oe of og oh oi oj ok ol go om on oo gr op oq or gu os ot ou ov hp bk">FM-Intent employs hierarchical multi-task learning where intent predictions are conducted first, and their results are used as input features for the next-item prediction task. This hierarchical relationship ensures that the next-item recommendations are informed by the predicted user intent, creating a more coherent and effective recommendation model.</p><h1 id="6df1" class="ow ox io bf oy oz pa pb gl pc pd pe gn pf pg ph pi pj pk pl pm pn po pp pq pr bk">Offline Results</h1><p id="f9d0" class="pw-post-body-paragraph ob oc io od b oe ps og oh oi pt ok ol go pu on oo gr pv oq or gu pw ot ou ov hp bk">We conducted comprehensive offline experiments on sampled Netflix user engagement data to evaluate FM-Intent’s performance. Note that FM-Intent uses a much smaller dataset for training compared to the FM production model due to its complex hierarchical prediction architecture.</p><h2 id="eb63" class="qy ox io bf oy gk qz dy gl gm ra ea gn go rb gp gq gr rc gs gt gu rd gv gw re bk">Next-Item and Next-Intent Prediction Accuracy</h2><p id="47fb" class="pw-post-body-paragraph ob oc io od b oe ps og oh oi pt ok ol go pu on oo gr pv oq or gu pw ot ou ov hp bk">Table 1 compares FM-Intent with several state-of-the-art sequential recommendation models, including our production model (FM-Intent-V0).</p><figure class="qj qk ql qm qn qo qg qh paragraph-image"><div role="button" tabindex="0" class="qp qq fl qr bh qs"><div class="qg qh qi"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*7h7aSQhq7U_heAUu%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*7h7aSQhq7U_heAUu%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*7h7aSQhq7U_heAUu%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*7h7aSQhq7U_heAUu%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*7h7aSQhq7U_heAUu%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*7h7aSQhq7U_heAUu%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*7h7aSQhq7U_heAUu%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*7h7aSQhq7U_heAUu 640w, https://miro.medium.com/v2/resize:fit:720/0*7h7aSQhq7U_heAUu 720w, https://miro.medium.com/v2/resize:fit:750/0*7h7aSQhq7U_heAUu 750w, https://miro.medium.com/v2/resize:fit:786/0*7h7aSQhq7U_heAUu 786w, https://miro.medium.com/v2/resize:fit:828/0*7h7aSQhq7U_heAUu 828w, https://miro.medium.com/v2/resize:fit:1100/0*7h7aSQhq7U_heAUu 1100w, https://miro.medium.com/v2/resize:fit:1400/0*7h7aSQhq7U_heAUu 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh fw qt c" width="700" height="203" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="e95c" class="pw-post-body-paragraph ob oc io od b oe of og oh oi oj ok ol go om on oo gr op oq or gu os ot ou ov hp bk"><em class="px">Table 1: Next-item and next-intent prediction results of baselines and our proposed method FM-Intent on the Netflix user engagement dataset.</em></p><p id="c67b" class="pw-post-body-paragraph ob oc io od b oe of og oh oi oj ok ol go om on oo gr op oq or gu os ot ou ov hp bk">All metrics are represented as relative % improvements compared to the SOTA baseline: TransAct. N/A indicates that a model is not capable of predicting a certain intent. Note that we added additional fully-connected layers to LSTM, GRU, and Transformer baselines in order to predict user intent, while we used original implementations for other baselines. FM-Intent demonstrates statistically significant improvement of 7.4% in next-item prediction accuracy compared to the best baseline (TransAct).</p><p id="01c4" class="pw-post-body-paragraph ob oc io od b oe of og oh oi oj ok ol go om on oo gr op oq or gu os ot ou ov hp bk">Most baseline models show limited performance as they either cannot predict user intent or cannot incorporate intent predictions into next-item recommendations. Our production model (FM-Intent-V0) performs well but lacks the ability to predict and leverage user intent. Note that FM-Intent-V0 is trained with a smaller dataset for a fair comparison with other models; the actual production model is trained with a much larger dataset.</p><h1 id="c46f" class="ow ox io bf oy oz pa pb gl pc pd pe gn pf pg ph pi pj pk pl pm pn po pp pq pr bk">Qualitative Analysis: User Clustering</h1><figure class="qj qk ql qm qn qo qg qh paragraph-image"><div role="button" tabindex="0" class="qp qq fl qr bh qs"><div class="qg qh rf"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*nBqH-ZHRXRfR3-e0%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*nBqH-ZHRXRfR3-e0%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*nBqH-ZHRXRfR3-e0%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*nBqH-ZHRXRfR3-e0%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*nBqH-ZHRXRfR3-e0%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*nBqH-ZHRXRfR3-e0%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*nBqH-ZHRXRfR3-e0%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*nBqH-ZHRXRfR3-e0 640w, https://miro.medium.com/v2/resize:fit:720/0*nBqH-ZHRXRfR3-e0 720w, https://miro.medium.com/v2/resize:fit:750/0*nBqH-ZHRXRfR3-e0 750w, https://miro.medium.com/v2/resize:fit:786/0*nBqH-ZHRXRfR3-e0 786w, https://miro.medium.com/v2/resize:fit:828/0*nBqH-ZHRXRfR3-e0 828w, https://miro.medium.com/v2/resize:fit:1100/0*nBqH-ZHRXRfR3-e0 1100w, https://miro.medium.com/v2/resize:fit:1400/0*nBqH-ZHRXRfR3-e0 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh fw qt c" width="700" height="753" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="c3dc" class="pw-post-body-paragraph ob oc io od b oe of og oh oi oj ok ol go om on oo gr op oq or gu os ot ou ov hp bk"><em class="px">Figure 3: K-means++ (K=10) clustering of user intent embeddings found by FM-Intent; FM-Intent finds unique clusters of users that share the similar intent.</em></p><p id="b54d" class="pw-post-body-paragraph ob oc io od b oe of og oh oi oj ok ol go om on oo gr op oq or gu os ot ou ov hp bk">FM-Intent generates meaningful user intent embeddings that can be used for clustering users with similar intents. Figure 3 visualizes 10 distinct clusters identified through K-means++ clustering.These clusters reveal meaningful user segments with distinct viewing patterns:</p><ul class=""><li id="de73" class="ob oc io od b oe of og oh oi oj ok ol go om on oo gr op oq or gu os ot ou ov rg pz qa bk">Users who primarily discover new content versus those who continue watching recent/favorite content.</li><li id="67ed" class="ob oc io od b oe qb og oh oi qc ok ol go qd on oo gr qe oq or gu qf ot ou ov rg pz qa bk">Genre enthusiasts (e.g., <em class="px">anime/kids content viewers</em>).</li><li id="d511" class="ob oc io od b oe qb og oh oi qc ok ol go qd on oo gr qe oq or gu qf ot ou ov rg pz qa bk">Users with specific viewing patterns (e.g., <em class="px">Rewatchers</em> versus <em class="px">casual viewers</em>).</li></ul><h1 id="2e06" class="ow ox io bf oy oz pa pb gl pc pd pe gn pf pg ph pi pj pk pl pm pn po pp pq pr bk">Potential Applications of FM-Intent</h1><p id="5515" class="pw-post-body-paragraph ob oc io od b oe ps og oh oi pt ok ol go pu on oo gr pv oq or gu pw ot ou ov hp bk">FM-Intent has been successfully integrated into Netflix’s recommendation ecosystem, can be leveraged for several downstream applications:</p><blockquote class="qu qv qw"><p id="39d4" class="ob oc px od b oe of og oh oi oj ok ol go om on oo gr op oq or gu os ot ou ov hp bk"><strong class="od ip">Personalized UI Optimization</strong>: The predicted user intent could inform the layout and content selection on the Netflix homepage, emphasizing different rows based on whether users are in discovery mode, continue-watching mode, or exploring specific genres.</p><p id="710a" class="ob oc px od b oe of og oh oi oj ok ol go om on oo gr op oq or gu os ot ou ov hp bk"><strong class="od ip">Analytics and User Understanding</strong>: Intent embeddings and clusters provide valuable insights into viewing patterns and preferences, informing content acquisition and production decisions.</p><p id="c3c1" class="ob oc px od b oe of og oh oi oj ok ol go om on oo gr op oq or gu os ot ou ov hp bk"><strong class="od ip">Enhanced Recommendation Signals</strong>: Intent predictions serve as features for other recommendation models, improving their accuracy and relevance.</p><p id="3b9a" class="ob oc px od b oe of og oh oi oj ok ol go om on oo gr op oq or gu os ot ou ov hp bk"><strong class="od ip">Search Optimization</strong>: Real-time intent predictions help prioritize search results based on the user’s current session intent.</p></blockquote><h1 id="cc1a" class="ow ox io bf oy oz pa pb gl pc pd pe gn pf pg ph pi pj pk pl pm pn po pp pq pr bk">Conclusion</h1><p id="6e9a" class="pw-post-body-paragraph ob oc io od b oe ps og oh oi pt ok ol go pu on oo gr pv oq or gu pw ot ou ov hp bk">FM-Intent represents an advancement in Netflix’s recommendation capabilities by enhancing them with hierarchical multi-task learning for user intent prediction. Our comprehensive experiments demonstrate that FM-Intent significantly outperforms state-of-the-art models, including our prior foundation model that focused solely on next-item prediction. By understanding not just what users might watch next but what underlying intents users have, we can provide more personalized, relevant, and satisfying recommendations.</p><h1 id="c957" class="ow ox io bf oy oz pa pb gl pc pd pe gn pf pg ph pi pj pk pl pm pn po pp pq pr bk">Acknowledgements</h1><p id="e32c" class="pw-post-body-paragraph ob oc io od b oe ps og oh oi pt ok ol go pu on oo gr pv oq or gu pw ot ou ov hp bk">We thank our stunning colleagues in the Foundation Model team &amp; AIMS org. for their valuable feedback and discussions. We also thank our partner teams for getting this up and running in production.</p><h1 id="2e36" class="ow ox io bf oy oz pa pb gl pc pd pe gn pf pg ph pi pj pk pl pm pn po pp pq pr bk">References</h1><p id="8218" class="pw-post-body-paragraph ob oc io od b oe ps og oh oi pt ok ol go pu on oo gr pv oq or gu pw ot ou ov hp bk">[1] Amatriain, X., &amp; Basilico, J. (2015). Recommender systems in industry: A netflix case study. In Recommender systems handbook (pp. 385–419). Springer.</p><p id="c60d" class="pw-post-body-paragraph ob oc io od b oe of og oh oi oj ok ol go om on oo gr op oq or gu os ot ou ov hp bk">[2] Gomez-Uribe, C. A., &amp; Hunt, N. (2015). The netflix recommender system: Algorithms, business value, and innovation. ACM Transactions on Management Information Systems (TMIS), 6(4), 1–19.</p><p id="84d0" class="pw-post-body-paragraph ob oc io od b oe of og oh oi oj ok ol go om on oo gr op oq or gu os ot ou ov hp bk">[3] Jannach, D., &amp; Jugovac, M. (2019). Measuring the business value of recommender systems. ACM Transactions on Management Information Systems (TMIS), 10(4), 1–23.</p><p id="b1fc" class="pw-post-body-paragraph ob oc io od b oe of og oh oi oj ok ol go om on oo gr op oq or gu os ot ou ov hp bk">[4] Bhattacharya, M., &amp; Lamkhede, S. (2022). Augmenting Netflix Search with In-Session Adapted Recommendations. In Proceedings of the 16th ACM Conference on Recommender Systems (pp. 542–545).</p><p id="3d01" class="pw-post-body-paragraph ob oc io od b oe of og oh oi oj ok ol go om on oo gr op oq or gu os ot ou ov hp bk">[5] Chen, Y., Liu, Z., Li, J., McAuley, J., &amp; Xiong, C. (2022). Intent contrastive learning for sequential recommendation. In Proceedings of the ACM Web Conference 2022 (pp. 2172–2182).</p><p id="cabd" class="pw-post-body-paragraph ob oc io od b oe of og oh oi oj ok ol go om on oo gr op oq or gu os ot ou ov hp bk">[6] Ding, Y., Ma, Y., Wong, W. K., &amp; Chua, T. S. (2021). Modeling instant user intent and content-level transition for sequential fashion recommendation. IEEE Transactions on Multimedia, 24, 2687–2700.</p><p id="f88f" class="pw-post-body-paragraph ob oc io od b oe of og oh oi oj ok ol go om on oo gr op oq or gu os ot ou ov hp bk">[7] Liu, Z., Chen, H., Sun, F., Xie, X., Gao, J., Ding, B., &amp; Shen, Y. (2021). Intent preference decoupling for user representation on online recommender system. In Proceedings of the Twenty-Ninth International Conference on International Joint Conferences on Artificial Intelligence (pp. 2575–2582).</p><p id="1c91" class="pw-post-body-paragraph ob oc io od b oe of og oh oi oj ok ol go om on oo gr op oq or gu os ot ou ov hp bk">[8] Xia, X., Eksombatchai, P., Pancha, N., Badani, D. D., Wang, P. W., Gu, N., Joshi, S. V., Farahpour, N., Zhang, Z., &amp; Zhai, A. (2023). TransAct: Transformer-based Realtime User Action Model for Recommendation at Pinterest. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (pp. 5249–5259).</p></div>]]></description>
      <link>https://netflixtechblog.com/fm-intent-predicting-user-session-intent-with-hierarchical-multi-task-learning-94c75e18f4b8</link>
      <guid>https://netflixtechblog.com/fm-intent-predicting-user-session-intent-with-hierarchical-multi-task-learning-94c75e18f4b8</guid>
      <pubDate>Wed, 21 May 2025 18:28:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Behind the Scenes: Building a Robust Ads Event Processing Pipeline]]></title>
      <description><![CDATA[<div class="ab ca"><div class="ch bg hu hv hw hx"><div><div></div><p id="3140" class="pw-post-body-paragraph oc od in oe b of og oh oi oj ok ol om gn on oo op gq oq or os gt ot ou ov ow ho bj"><a class="af ha" href="https://www.linkedin.com/in/kineshsatiya/" rel="noopener ugc nofollow" target="_blank">Kinesh Satiya</a></p><h2 id="c0fb" class="ox oy in be oz gj pa dx gk gl pb dz gm gn pc go gp gq pd gr gs gt pe gu gv pf bj">Introduction</h2><p id="0bc0" class="pw-post-body-paragraph oc od in oe b of pg oh oi oj ph ol om gn pi oo op gq pj or os gt pk ou ov ow ho bj">In a digital advertising platform, a robust feedback system is essential for the lifecycle and success of an ad campaign. This system comprises of diverse sub-systems designed to monitor, measure, and optimize ad campaigns. At Netflix, we embarked on a journey to build a robust event processing platform that not only meets the current demands but also scales for future needs. This blog post delves into the architectural evolution and technical decisions that underpin our Ads event processing pipeline.</p><p id="5676" class="pw-post-body-paragraph oc od in oe b of og oh oi oj ok ol om gn on oo op gq oq or os gt ot ou ov ow ho bj">Ad serving acts like the “brain” — making decisions, optimizing delivery and ensuring right Ad is shown to the right member at the right time. Meanwhile, ad events, after an Ad is rendered, function like “heartbeats”, continuously providing real-time feedback (oxygen/nutrients) that fuels better decision-making, optimizations, reporting, measurement, and billing. Expanding on this analogy:</p><ul class=""><li id="ea42" class="oc od in oe b of og oh oi oj ok ol om gn on oo op gq oq or os gt ot ou ov ow pl pm pn bj">Just as the brain relies on continuous blood flow, ad serving depends on a steady stream of ad events to adjust next ad serving decision, frequency capping, pacing, and personalization.</li><li id="6f87" class="oc od in oe b of po oh oi oj pp ol om gn pq oo op gq pr or os gt ps ou ov ow pl pm pn bj">If the nervous system stops sending signals (ad events stop flowing), the brain (ad serving) lacks critical insights and starts making poor decisions or even fails.</li><li id="e937" class="oc od in oe b of po oh oi oj pp ol om gn pq oo op gq pr or os gt ps ou ov ow pl pm pn bj">The healthier and more accurate the event stream (just like strong heart function), the better the ad serving system can adapt, optimize, and drive business outcomes.</li></ul><p id="2ec6" class="pw-post-body-paragraph oc od in oe b of og oh oi oj ok ol om gn on oo op gq oq or os gt ot ou ov ow ho bj">Let’s dive into the journey of building this pipeline.</p><h2 id="35f9" class="ox oy in be oz gj pa dx gk gl pb dz gm gn pc go gp gq pd gr gs gt pe gu gv pf bj">The Pilot</h2><p id="1001" class="pw-post-body-paragraph oc od in oe b of pg oh oi oj ph ol om gn pi oo op gq pj or os gt pk ou ov ow ho bj">In November 2022, we launched a brand <a class="af ha" href="https://about.netflix.com/en/news/announcing-basic-with-ads-us" rel="noopener ugc nofollow" target="_blank">new basic ads plan</a>, in partnership with Microsoft. The software systems extended the existing Netflix playback systems to play ads. Initially, the system was designed to be simple, secure, and efficient, with an underlying ethos of device-originated and server-proxied operations. The system consisted of three main components: the Microsoft Ad Server, Netflix Ads Manager, and Ad Event Handler. Each ad served required tracking to ensure the feedback loop functioned effectively, providing the external ad server with insights on impressions, frequency capping (advertiser policy that limits the number of times a user sees a specific ad), and monetization processes.</p><p id="4b23" class="pw-post-body-paragraph oc od in oe b of og oh oi oj ok ol om gn on oo op gq oq or os gt ot ou ov ow ho bj">Key features of this system include:</p><ol class=""><li id="568d" class="oc od in oe b of og oh oi oj ok ol om gn on oo op gq oq or os gt ot ou ov ow pt pm pn bj"><strong class="oe io">Client Request: </strong>Client devices request for ads during an ad break from Netflix playback systems, which is then decorated with information by ads manager to request ads from the ad server.</li><li id="405c" class="oc od in oe b of po oh oi oj pp ol om gn pq oo op gq pr or os gt ps ou ov ow pt pm pn bj"><strong class="oe io">Server-Side Ad Insertion:</strong> The Ad Server sends ad responses using the VAST (Video Ad Serving Template) format.</li><li id="8552" class="oc od in oe b of po oh oi oj pp ol om gn pq oo op gq pr or os gt ps ou ov ow pt pm pn bj"><strong class="oe io">Netflix Ads Manager:</strong> This service parses VAST documents, extracts tracking event information, and creates a simplified response structure for Netflix playback systems and client devices. <br /> — The tracking information is packed into a structured protobuf data model.<br /> — This structure is encrypted to create an opaque token.<br /> — The final response, informs the client devices, when to send an event and the corresponding token.</li><li id="d205" class="oc od in oe b of po oh oi oj pp ol om gn pq oo op gq pr or os gt ps ou ov ow pt pm pn bj"><strong class="oe io">Client Device:</strong> During ad playback, client devices send events accompanied by a token. The Netflix telemetry system then enqueues all these events in Kafka for asynchronous processing.</li><li id="e73a" class="oc od in oe b of po oh oi oj pp ol om gn pq oo op gq pr or os gt ps ou ov ow pt pm pn bj"><strong class="oe io">Ads Event Handler:</strong> This component is a Kafka consumer, that reads/decrypts the event payload and forwards the tracking information encoded back to the ad server and other vendors.</li></ol></div></div><div class="pu"><div class="ab ca"><div class="nf pv ng pw nh px ce py cf pz ch bg"><figure class="qd qe qf qg qh pu qi qj paragraph-image"><div role="button" tabindex="0" class="qk ql fk qm bg qn"><div class="qa qb qc"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*S_6xv6LxoRyL8KtUXqWueQ.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*S_6xv6LxoRyL8KtUXqWueQ.png" /></picture></div></div><figcaption class="qp fe qq qa qb qr qs be b bf z dt">Fig 1: Basic Ad Event Handling System</figcaption></figure></div></div></div><div class="ab ca"><div class="ch bg hu hv hw hx"><p id="ee5a" class="pw-post-body-paragraph oc od in oe b of og oh oi oj ok ol om gn on oo op gq oq or os gt ot ou ov ow ho bj">There is an <a class="af ha" rel="noopener ugc nofollow" href="https://netflixtechblog.com/ensuring-the-successful-launch-of-ads-on-netflix-f99490fdf1ba" target="_blank">excellent prior blog</a> post that explains how this systems was tested end-to-end at scale. This system design allowed us to quickly add new integrations for verification with vendors like DV, IAS and Nielsen for measurement.</p><h2 id="1fc6" class="ox oy in be oz gj pa dx gk gl pb dz gm gn pc go gp gq pd gr gs gt pe gu gv pf bj">The Expansion</h2><p id="06ee" class="pw-post-body-paragraph oc od in oe b of pg oh oi oj ph ol om gn pi oo op gq pj or os gt pk ou ov ow ho bj">As we continued to expand our third-party (3P) advertising vendors for measurement, tracking and verification, we identified a critical trend: growth in the volume of data encapsulated within opaque tokens. These tokens, which are cached on client devices, present a risk of elevated memory usage, potentially impacting device performance. We also anticipated increase in third-party tracking URLs, metadata needs, and more event types as our business added new capabilities.</p><p id="b65f" class="pw-post-body-paragraph oc od in oe b of og oh oi oj ok ol om gn on oo op gq oq or os gt ot ou ov ow ho bj">To strategically address these challenges, we introduced a new persistence layer using <a class="af ha" rel="noopener ugc nofollow" href="https://netflixtechblog.com/introducing-netflixs-key-value-data-abstraction-layer-1ea8a0a11b30" target="_blank">Key-Value abstraction</a>, between ad serving and event handling system: Ads Metadata Registry. This transient storage service stores metadata for each Ad served, and upon callback, event handler would read the tracking information to relay information to the vendors. The contract between the client device and Ads systems continues to use the opaque token per event, but now, instead of tracking information, it contains reference identifiers — Ad ID, the corresponding metadata record ID in the registry and the event name. This approach future proofed our systems to handle any growth in data that needs to pass from ad serving to event handling systems.</p></div></div><div class="pu"><div class="ab ca"><div class="nf pv ng pw nh px ce py cf pz ch bg"><figure class="qd qe qf qg qh pu qi qj paragraph-image"><div role="button" tabindex="0" class="qk ql fk qm bg qn"><div class="qa qb qc"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*0wP5cyAj84Vju7ryabJE1A.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*0wP5cyAj84Vju7ryabJE1A.png" /></picture></div></div><figcaption class="qp fe qq qa qb qr qs be b bf z dt">Fig 2: Storage service between Ad Serving &amp; Reporting</figcaption></figure></div></div></div><div class="ab ca"><div class="ch bg hu hv hw hx"><h2 id="f553" class="ox oy in be oz gj pa dx gk gl pb dz gm gn pc go gp gq pd gr gs gt pe gu gv pf bj">The Evolution</h2><p id="b9ba" class="pw-post-body-paragraph oc od in oe b of pg oh oi oj ph ol om gn pi oo op gq pj or os gt pk ou ov ow ho bj">In January of 2024, we decided to invest in in-house advertising technology platform. This implied that the event processing pipeline had to evolve significantly — attain parity with existing offerings and continue to support new product launches with rapid iterations using in-house Netflix Ad Server. This required re-evaluation of the entire architecture across all of Ads engineering teams.</p><p id="83b3" class="pw-post-body-paragraph oc od in oe b of og oh oi oj ok ol om gn on oo op gq oq or os gt ot ou ov ow ho bj">First, we made an inventory of the use-cases that would need to be supported through ad events.</p><ol class=""><li id="5134" class="oc od in oe b of og oh oi oj ok ol om gn on oo op gq oq or os gt ot ou ov ow pt pm pn bj">We’d need to start supporting frequency capping in-house for all ads through Netflix Ad server.</li><li id="8187" class="oc od in oe b of po oh oi oj pp ol om gn pq oo op gq pr or os gt ps ou ov ow pt pm pn bj">Incorporate pricing information for impressions to set the stage for billing events, which are used to charge advertisers.</li><li id="2966" class="oc od in oe b of po oh oi oj pp ol om gn pq oo op gq pr or os gt ps ou ov ow pt pm pn bj">A robust reporting system to share campaign reports with advertisers, combined with metrics data collection, helps assess the delivery and effectiveness of the campaign.</li><li id="da68" class="oc od in oe b of po oh oi oj pp ol om gn pq oo op gq pr or os gt ps ou ov ow pt pm pn bj">Scale event handler to perform tracking information look-ups across different vendors.</li></ol><p id="c00a" class="pw-post-body-paragraph oc od in oe b of og oh oi oj ok ol om gn on oo op gq oq or os gt ot ou ov ow ho bj">Next, we examined upcoming launches, such as Pause/Display ads, to gain deeper insights into our strategic initiatives. We recognized that Display Ads would utilize a distinct logging framework, suggesting that different upstream pipelines might deliver ad telemetry. However, the downstream use-cases were expected to remain largely consistent. Additionally, by reviewing the goals of our telemetry teams, we saw large initiatives aimed at upgrading the platform, indicating potential future migrations.</p><p id="ac48" class="pw-post-body-paragraph oc od in oe b of og oh oi oj ok ol om gn on oo op gq oq or os gt ot ou ov ow ho bj">Keeping the above insights &amp; challenges in mind,</p><ul class=""><li id="3983" class="oc od in oe b of og oh oi oj ok ol om gn on oo op gq oq or os gt ot ou ov ow pl pm pn bj">We planned a centralized ad event collection system. This centralized service would consolidate common operations like decryption of tokens, enrichment, hashing identifiers into a single step execution and provide a single unified data contract to consumers that is highly extensible (like being agnostic to ad server &amp; ad media).</li><li id="4f2e" class="oc od in oe b of po oh oi oj pp ol om gn pq oo op gq pr or os gt ps ou ov ow pl pm pn bj">We proposed moving all consumers of ad telemetry downstream of the centralized service. This creates a clean separation between upstream systems and consumers in Ads Engineering.</li><li id="ee4f" class="oc od in oe b of po oh oi oj pp ol om gn pq oo op gq pr or os gt ps ou ov ow pl pm pn bj">In the initial development phase of our advertising system, a crucial component was the creation of ad sessions based on individual ad events. This system was constructed using ad playback telemetry, which allowed us to gather essential metrics from these ad sessions. A significant decision in this plan was to position the ad sessionization process downstream of the raw ad events.</li><li id="f367" class="oc od in oe b of po oh oi oj pp ol om gn pq oo op gq pr or os gt ps ou ov ow pl pm pn bj">The proposal also recommended moving all our Ads data processing pipelines for reporting/analytics/metrics for Ads using the data published by the centralized system.</li></ul><p id="ea56" class="pw-post-body-paragraph oc od in oe b of og oh oi oj ok ol om gn on oo op gq oq or os gt ot ou ov ow ho bj">Putting together all the components in our vision -</p></div></div><div class="pu"><div class="ab ca"><div class="nf pv ng pw nh px ce py cf pz ch bg"><figure class="qd qe qf qg qh pu qi qj paragraph-image"><div role="button" tabindex="0" class="qk ql fk qm bg qn"><div class="qa qb qt"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*aPL3RHeFEzlw_psaLddWKw.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*aPL3RHeFEzlw_psaLddWKw.png" /></picture></div></div><figcaption class="qp fe qq qa qb qr qs be b bf z dt">Fig 3: Ad Event processing pipeline</figcaption></figure></div></div></div><div class="ab ca"><div class="ch bg hu hv hw hx"><p id="da0e" class="pw-post-body-paragraph oc od in oe b of og oh oi oj ok ol om gn on oo op gq oq or os gt ot ou ov ow ho bj">Key components on event processing pipeline -</p><p id="66c8" class="pw-post-body-paragraph oc od in oe b of og oh oi oj ok ol om gn on oo op gq oq or os gt ot ou ov ow ho bj"><strong class="oe io">Ads Event Publisher:</strong> This centralized system is responsible for collecting ads telemetry and providing unified ad events to the ads engineering teams. It supports various functions such as measurement, finance/billing, reporting, frequency capping, and maintaining an essential feedback loop back to the ad server.</p><p id="6e2e" class="pw-post-body-paragraph oc od in oe b of og oh oi oj ok ol om gn on oo op gq oq or os gt ot ou ov ow ho bj"><strong class="oe io">Realtime Consumers</strong></p><ol class=""><li id="8bf9" class="oc od in oe b of og oh oi oj ok ol om gn on oo op gq oq or os gt ot ou ov ow pt pm pn bj"><strong class="oe io">Frequency Capping: </strong>This system tracks impressions for each campaign, profile, and any other frequency capping parameters set up for the campaign. It is utilized by the Ad Server during each ad decision to ensure ads are served with frequency limits.</li><li id="0475" class="oc od in oe b of po oh oi oj pp ol om gn pq oo op gq pr or os gt ps ou ov ow pt pm pn bj"><strong class="oe io">Ads Metrics: </strong>This component is a Flink job that transforms raw data to a set of dimensions and metrics, subsequently writing to Apache Druid OLAP database. The streaming data is further backed by an offline process that corrects any inaccuracy during streaming ingestion and providing accurate metrics. It provides real-time metrics to assess the delivery health of campaigns and applies budget capping functionality.</li><li id="efd4" class="oc od in oe b of po oh oi oj pp ol om gn pq oo op gq pr or os gt ps ou ov ow pt pm pn bj"><strong class="oe io">Ads Sessionizer: </strong>An Apache Flink job that consolidates all events related to a single ad into an Ad Session. This session provides real-time information about ad playback, offering essential business insights and reporting. It is a crucial job that supports all downstream analytical and reporting processes.</li><li id="f42c" class="oc od in oe b of po oh oi oj pp ol om gn pq oo op gq pr or os gt ps ou ov ow pt pm pn bj"><strong class="oe io">Ads Event Handler: </strong>This service continuously sends information to ad vendors by reading tracking information from ad events, ensuring accurate and timely data exchange.</li></ol><p id="ff42" class="pw-post-body-paragraph oc od in oe b of og oh oi oj ok ol om gn on oo op gq oq or os gt ot ou ov ow ho bj"><strong class="oe io">Billing/Revenue: </strong>These are offline workflows designed to curate impressions, supporting billing and revenue recognition processes.</p><p id="56fe" class="pw-post-body-paragraph oc od in oe b of og oh oi oj ok ol om gn on oo op gq oq or os gt ot ou ov ow ho bj"><strong class="oe io">Ads Reporting &amp; Metrics: </strong>This service powers reporting module for our account managers and provides a centralized metrics API that help assess the delivery of a campaign.</p><p id="1868" class="pw-post-body-paragraph oc od in oe b of og oh oi oj ok ol om gn on oo op gq oq or os gt ot ou ov ow ho bj">This was a massive multi-quarter effort across different engineering teams. With extensive planning (kudos to our TPM team!) and coordination, we were able to iterate fast, build several services and execute the vision above, to power our in-house ads technology platform.</p><h2 id="67f3" class="ox oy in be oz gj pa dx gk gl pb dz gm gn pc go gp gq pd gr gs gt pe gu gv pf bj">Conclusion</h2><p id="7e2d" class="pw-post-body-paragraph oc od in oe b of pg oh oi oj ph ol om gn pi oo op gq pj or os gt pk ou ov ow ho bj">These systems have significantly accelerated our ability to launch new capabilities for the business.</p><ul class=""><li id="6878" class="oc od in oe b of og oh oi oj ok ol om gn on oo op gq oq or os gt ot ou ov ow pl pm pn bj">Through our partnership with Microsoft, Display Ad events were integrated into the new pipeline for reusability and ensuring when launching through Netflix ads systems, all use-cases were covered.</li><li id="abd9" class="oc od in oe b of po oh oi oj pp ol om gn pq oo op gq pr or os gt ps ou ov ow pl pm pn bj">Programmatic buying capabilities now support the exchange of numerous trackers and dynamic bid prices on impression events.</li><li id="8aac" class="oc od in oe b of po oh oi oj pp ol om gn pq oo op gq pr or os gt ps ou ov ow pl pm pn bj">Sharing opt-out signals helps ensure privacy and compliance with GDPR regulations for Ads business in Europe, supporting accurate reporting and measurement.</li><li id="5a25" class="oc od in oe b of po oh oi oj pp ol om gn pq oo op gq pr or os gt ps ou ov ow pl pm pn bj">New event types like Ad clicks and scanning of QR codes events also flow through the pipeline, ensuring all metrics and reporting are tracked consistently.</li></ul><p id="d5ce" class="pw-post-body-paragraph oc od in oe b of og oh oi oj ok ol om gn on oo op gq oq or os gt ot ou ov ow ho bj"><strong class="oe io">Key Takeways</strong></p><ul class=""><li id="1e7a" class="oc od in oe b of og oh oi oj ok ol om gn on oo op gq oq or os gt ot ou ov ow pl pm pn bj"><strong class="oe io">Strategic, incremental evolution:</strong> The development of our ads event processing systems has been a carefully orchestrated journey. Each iteration was meticulously planned by addressing existing challenges, anticipating future needs, and showcasing teamwork, planning, and coordination across various teams. These pillars have been fundamental to the success of this journey.</li><li id="8b0f" class="oc od in oe b of po oh oi oj pp ol om gn pq oo op gq pr or os gt ps ou ov ow pl pm pn bj"><strong class="oe io">Data contract:</strong> A clear data contract has been pivotal in ensuring consistency in interpretation and interoperability across our systems. By standardizing the data models, and establishing a clear data exchange between ad serving, and centralized event collection, our teams have been able to iterate at exceptional speed and continue to deliver many launches on time.</li><li id="bb6b" class="oc od in oe b of po oh oi oj pp ol om gn pq oo op gq pr or os gt ps ou ov ow pl pm pn bj"><strong class="oe io">Separation of concerns: </strong>Consumers are relieved from the need to understand each source of ad telemetry or manage updates and migrations. Instead, a centralized system handles these tasks, allowing consumers to focus on their core business logic.</li></ul><p id="27da" class="pw-post-body-paragraph oc od in oe b of og oh oi oj ok ol om gn on oo op gq oq or os gt ot ou ov ow ho bj">We have an exciting list of projects on the horizon. These include managing ad events from ads on Netflix live streams, de-duplication processes, and enriching data signals to deliver enhanced reporting and insights. Additionally, we are advancing our Native Ads strategy, integrating Conversion API for improved conversion tracking, among many others.</p><p id="1bbb" class="pw-post-body-paragraph oc od in oe b of og oh oi oj ok ol om gn on oo op gq oq or os gt ot ou ov ow ho bj">This is definitely not a season finale; it’s just the beginning of our journey to create a best-in-class ads technology platform. We warmly invite you to share your thoughts and comments with us. If you’re interested in learning more or becoming a part of this innovative journey, <a class="af ha" href="https://jobs.netflix.com/" rel="noopener ugc nofollow" target="_blank">Ads Engineering is hiring</a>!</p><h2 id="f86f" class="ox oy in be oz gj pa dx gk gl pb dz gm gn pc go gp gq pd gr gs gt pe gu gv pf bj"><em class="qu">Acknowledgements</em></h2><p id="bf29" class="pw-post-body-paragraph oc od in oe b of pg oh oi oj ph ol om gn pi oo op gq pj or os gt pk ou ov ow ho bj"><em class="qv">A special thanks to our amazing colleagues and teams who helped build our foundational post-impression system: </em><a class="af ha" href="https://www.linkedin.com/in/simonspencer1/" rel="noopener ugc nofollow" target="_blank"><em class="qv">Simon Spencer</em></a><em class="qv">, </em><a class="af ha" href="https://www.linkedin.com/in/priyankaavj/" rel="noopener ugc nofollow" target="_blank"><em class="qv">Priyankaa Vijayakumar,</em></a><a class="af ha" href="https://www.linkedin.com/in/indrajit-roy-choudhury-5b011754/" rel="noopener ugc nofollow" target="_blank"><em class="qv">Indrajit Roy Choudhury</em></a><em class="qv">; Ads TPM team — </em><a class="af ha" href="https://www.linkedin.com/in/sonyabellamy/" rel="noopener ugc nofollow" target="_blank"><em class="qv">Sonya Bellamy</em></a><em class="qv">; the Ad Serving Team —</em><a class="af ha" href="https://www.linkedin.com/in/andrewjsweeney/" rel="noopener ugc nofollow" target="_blank"><em class="qv"> Andrew Sweeney</em></a><em class="qv">, </em><a class="af ha" href="https://www.linkedin.com/in/tim-z-b9112034/" rel="noopener ugc nofollow" target="_blank"><em class="qv">Tim Zheng</em></a>, <a class="af ha" href="https://www.linkedin.com/in/haidongt/" rel="noopener ugc nofollow" target="_blank"><em class="qv">Haidong Tang</em></a><em class="qv"> and </em><a class="af ha" href="https://www.linkedin.com/in/edhbarker/" rel="noopener ugc nofollow" target="_blank"><em class="qv">Ed Barker</em></a><em class="qv">; the Ads Data Engineering Team — </em><a class="af ha" href="https://www.linkedin.com/in/sonalisharma/" rel="noopener ugc nofollow" target="_blank"><em class="qv">Sonali Sharma</em></a><em class="qv">, Harsha Arepalli, and </em><a class="af ha" href="https://www.linkedin.com/in/winifredtran/" rel="noopener ugc nofollow" target="_blank"><em class="qv">Wini Tran</em></a><em class="qv">; Product Data Systems — </em><a class="af ha" href="https://www.linkedin.com/in/d3cay/" rel="noopener ugc nofollow" target="_blank"><em class="qv">David Klosowski;</em></a><em class="qv"> and the entire Ads Reporting and Measurement team!</em></p></div></div></div>]]></description>
      <link>https://netflixtechblog.com/behind-the-scenes-building-a-robust-ads-event-processing-pipeline-e4e86caf9249</link>
      <guid>https://netflixtechblog.com/behind-the-scenes-building-a-robust-ads-event-processing-pipeline-e4e86caf9249</guid>
      <pubDate>Fri, 09 May 2025 21:44:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Measuring Dialogue Intelligibility for Netflix Content]]></title>
      <description><![CDATA[<div><div class="ab ca"><div class="ch bg hu hv hw hx"><div class="ih l"><div class="cp ab ii"><div><div class="bl" aria-hidden="false"><div class="ij q ik hz il ii ao"><p class="be b bf z dt">Featured</p></div></div></div></div></div></div></div></div><div class="ho im in io ip"><div class="ab ca"><div class="ch bg hu hv hw hx"><div><div></div><p id="b9e5" class="pw-post-body-paragraph ob oc is od b oe of og oh oi oj ok ol gn om on oo gq op oq or gt os ot ou ov ho bj"><em class="ow">Enhancing Member Experience Through Strategic Collaboration</em></p><p id="1886" class="pw-post-body-paragraph ob oc is od b oe of og oh oi oj ok ol gn om on oo gq op oq or gt os ot ou ov ho bj"><a class="af ha" href="https://www.linkedin.com/in/ozziesutherland/" rel="noopener ugc nofollow" target="_blank">Ozzie Sutherland</a>, <a class="af ha" href="https://www.linkedin.com/in/iroroorife/" rel="noopener ugc nofollow" target="_blank">Iroro Orife</a>, <a class="af ha" href="https://www.linkedin.com/in/chih-wei-wu-73081689/" rel="noopener ugc nofollow" target="_blank">Chih-Wei Wu</a>, <a class="af ha" href="https://www.linkedin.com/in/bhanusrikanth/" rel="noopener ugc nofollow" target="_blank">Bhanu Srikanth</a></p><p id="7f09" class="pw-post-body-paragraph ob oc is od b oe of og oh oi oj ok ol gn om on oo gq op oq or gt os ot ou ov ho bj">At Netflix, delivering the best possible experience for our members is at the heart of everything we do, and we know we can’t do it alone. That’s why we work closely with a diverse ecosystem of technology partners, combining their deep expertise with our creative and operational insights. Together, we explore new ideas, develop practical tools, and push technical boundaries in service of storytelling. This collaboration not only empowers the talented creatives working on our shows with better tools to bring their vision to life, but also helps us innovate in service of our members. By building these partnerships on trust, transparency, and shared purpose, we’re able to move faster and more meaningfully, always with the goal of making our stories more immersive, accessible, and enjoyable for audiences everywhere. One area where this collaboration is making a meaningful impact is in improving dialogue intelligibility, from set to screen. We call this the Dialogue Integrity Pipeline.</p><h2 id="4d08" class="ox oy is be oz gj pa dx gk gl pb dz gm gn pc go gp gq pd gr gs gt pe gu gv pf bj">Dialogue Integrity Pipeline</h2><p id="693a" class="pw-post-body-paragraph ob oc is od b oe pg og oh oi ph ok ol gn pi on oo gq pj oq or gt pk ot ou ov ho bj">We’ve all been there, settling in for a night of entertainment, only to find ourselves straining to catch what was just said on screen. You’re wrapped up in the story, totally invested, when suddenly a key line of dialogue vanishes into thin air. “Wait, what did they say? I can’t understand the dialogue! What just happened?”</p><p id="1d72" class="pw-post-body-paragraph ob oc is od b oe of og oh oi oj ok ol gn om on oo gq op oq or gt os ot ou ov ho bj">You may pick up the remote and rewind, turn up the volume, or try to stay with it and hope this doesn’t happen again. Creating sophisticated, modern series and films requires an incredible artistic &amp; technical effort. At Netflix, we strive to ensure those great stories are easy for the audience to enjoy. Dialogue intelligibility can break down at multiple points in what we call the <strong class="od it">Dialogue Integrity Pipeline</strong>, the journey from on-set capture to final playback at home. Many facets of the process can contribute to dialogue that’s difficult to understand:</p><ul class=""><li id="6e0e" class="ob oc is od b oe of og oh oi oj ok ol gn om on oo gq op oq or gt os ot ou ov pl pm pn bj">Naturalistic acting styles, diverse speech patterns, and accents</li><li id="ac5c" class="ob oc is od b oe po og oh oi pp ok ol gn pq on oo gq pr oq or gt ps ot ou ov pl pm pn bj">Noisy locations, microphone placement problems on set</li><li id="7a02" class="ob oc is od b oe po og oh oi pp ok ol gn pq on oo gq pr oq or gt ps ot ou ov pl pm pn bj">Cinematic (high dynamic range) mixing styles, excessive dialogue processing, substandard equipment</li><li id="c369" class="ob oc is od b oe po og oh oi pp ok ol gn pq on oo gq pr oq or gt ps ot ou ov pl pm pn bj">Audio compromises through the distribution pipeline</li><li id="71bd" class="ob oc is od b oe po og oh oi pp ok ol gn pq on oo gq pr oq or gt ps ot ou ov pl pm pn bj">TVs with inadequate speakers, noisy home environments</li></ul><p id="746a" class="pw-post-body-paragraph ob oc is od b oe of og oh oi oj ok ol gn om on oo gq op oq or gt os ot ou ov ho bj">Addressing these issues is critical to maintaining the standard of excellence our content deserves.</p><h2 id="d3ae" class="ox oy is be oz gj pa dx gk gl pb dz gm gn pc go gp gq pd gr gs gt pe gu gv pf bj">Measurement at Scale</h2><p id="0617" class="pw-post-body-paragraph ob oc is od b oe pg og oh oi ph ok ol gn pi on oo gq pj oq or gt pk ot ou ov ho bj">Netflix utilizes industry-standard loudness meters to measure content and its adherence to our core loudness specifications. This tool also provides feedback on audio dynamic range (loud to soft) which impacts dialogue intelligibility. The Audio Algorithms team at Netflix wanted to take these measurements further and develop a holistic understanding of dialogue intelligibility throughout the runtime of a given title.</p><p id="d830" class="pw-post-body-paragraph ob oc is od b oe of og oh oi oj ok ol gn om on oo gq op oq or gt os ot ou ov ho bj">The team developed a Speech Intelligibility measurement system based on the Short-time Objective Intelligibility (STOI) metric [<a class="af ha" href="https://www.researchgate.net/profile/Cees-Taal/publication/224219052_An_Algorithm_for_Intelligibility_Prediction_of_Time-Frequency_Weighted_Noisy_Speech/links/0deec51da9fbbc5eea000000/An-Algorithm-for-Intelligibility-Prediction-of-Time-Frequency-Weighted-Noisy-Speech.pdf" rel="noopener ugc nofollow" target="_blank">Taal et al.</a> (IEEE <em class="ow">Transactions on Audio, Speech, and Language Processing</em>)]. Firstly, a speech activity detector analyses the dialogue stem to render speech utterances, which are then compared to non-speech sounds in the mix, typically Music and Effects. Then the system calculates the Signal-to-Noise ratio, in each speech frequency band, the results of which are summarized succinctly, per-utterance on the range [0, 1.0], to quantify the degree to which competing Music and Effects can distract the listener.</p><figure class="pw px py pz qa qb pt pu paragraph-image"><div role="button" tabindex="0" class="qc qd fk qe bg qf"><div class="pt pu pv"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*WSViFfuvT8pcZshi%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*WSViFfuvT8pcZshi%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*WSViFfuvT8pcZshi%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*WSViFfuvT8pcZshi%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*WSViFfuvT8pcZshi%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*WSViFfuvT8pcZshi%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*WSViFfuvT8pcZshi%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*WSViFfuvT8pcZshi 640w, https://miro.medium.com/v2/resize:fit:720/0*WSViFfuvT8pcZshi 720w, https://miro.medium.com/v2/resize:fit:750/0*WSViFfuvT8pcZshi 750w, https://miro.medium.com/v2/resize:fit:786/0*WSViFfuvT8pcZshi 786w, https://miro.medium.com/v2/resize:fit:828/0*WSViFfuvT8pcZshi 828w, https://miro.medium.com/v2/resize:fit:1100/0*WSViFfuvT8pcZshi 1100w, https://miro.medium.com/v2/resize:fit:1400/0*WSViFfuvT8pcZshi 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /></picture></div></div><figcaption class="qh fe qi pt pu qj qk be b bf z dt">This chart shows how eSTOI (extended Short-Time Objective Intelligibility) method measures dialogue (fg [foreground] stem in the graphic) against non-speech (bg [background] stem in the graphic) to judge intelligibility based on competing non-speech sound.</figcaption></figure><h2 id="a25d" class="ox oy is be oz gj pa dx gk gl pb dz gm gn pc go gp gq pd gr gs gt pe gu gv pf bj">Optimizing Dialogue Prior to Delivery</h2><p id="b344" class="pw-post-body-paragraph ob oc is od b oe pg og oh oi ph ok ol gn pi on oo gq pj oq or gt pk ot ou ov ho bj">Understanding dialogue intelligibility across Netflix titles is invaluable, but our mission goes beyond analysis — we strive to empower creators with the tools to craft mixes that resonate seamlessly with audiences at home.</p><p id="1547" class="pw-post-body-paragraph ob oc is od b oe of og oh oi oj ok ol gn om on oo gq op oq or gt os ot ou ov ho bj">Seeing the lack of dedicated Dialogue Intelligibility Meter plugins for Digital Audio Workstations, we teamed up with industry leaders, Fraunhofer Institute for Digital Media Technology IDMT (Fraunhofer IDMT) and Nugen Audio to pioneer a solution that enhances creative control and ensures crystal-clear dialogue from mix to final delivery.</p><p id="19ac" class="pw-post-body-paragraph ob oc is od b oe of og oh oi oj ok ol gn om on oo gq op oq or gt os ot ou ov ho bj">We collaborated with Fraunhofer IDMT to adapt their machine-learning-based speech intelligibility solution for cross-platform plugin standards and brought in Nugen Audio to develop DAW-compatible plugins.</p><h2 id="c373" class="ox oy is be oz gj pa dx gk gl pb dz gm gn pc go gp gq pd gr gs gt pe gu gv pf bj">Fraunhofer IDMT</h2><p id="b95b" class="pw-post-body-paragraph ob oc is od b oe pg og oh oi ph ok ol gn pi on oo gq pj oq or gt pk ot ou ov ho bj">The Fraunhofer Department of Hearing, Speech, and Audio Technology HSA has done significant research and development on media processing tools that measure speech intelligibility. In 2020, the machine learning-based method was integrated into Steinberg’s Nuendo Digital Audio Workstation. We approached the Fraunhofer engineering team with a collaboration proposal to make their technology accessible to other audio workstations through the cross-platform VST (Virtual Studio Technology) and AAX (Avid Audio Extension) plugin standards. The scientists were keen on the project and provided their dialogue intelligibility library.</p><figure class="pw px py pz qa qb pt pu paragraph-image"><div role="button" tabindex="0" class="qc qd fk qe bg qf"><div class="pt pu ql"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*wuapXe2lajcx3tTj%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*wuapXe2lajcx3tTj%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*wuapXe2lajcx3tTj%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*wuapXe2lajcx3tTj%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*wuapXe2lajcx3tTj%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*wuapXe2lajcx3tTj%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*wuapXe2lajcx3tTj%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*wuapXe2lajcx3tTj 640w, https://miro.medium.com/v2/resize:fit:720/0*wuapXe2lajcx3tTj 720w, https://miro.medium.com/v2/resize:fit:750/0*wuapXe2lajcx3tTj 750w, https://miro.medium.com/v2/resize:fit:786/0*wuapXe2lajcx3tTj 786w, https://miro.medium.com/v2/resize:fit:828/0*wuapXe2lajcx3tTj 828w, https://miro.medium.com/v2/resize:fit:1100/0*wuapXe2lajcx3tTj 1100w, https://miro.medium.com/v2/resize:fit:1400/0*wuapXe2lajcx3tTj 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /></picture></div></div><figcaption class="qh fe qi pt pu qj qk be b bf z dt">The Fraunhofer IDMT Dialogue Intelligibility Meter integrated into the Steinberg Nuendo Digital Audio Workstation.</figcaption></figure><h2 id="698d" class="ox oy is be oz gj pa dx gk gl pb dz gm gn pc go gp gq pd gr gs gt pe gu gv pf bj">Nugen Audio</h2><p id="77fb" class="pw-post-body-paragraph ob oc is od b oe pg og oh oi ph ok ol gn pi on oo gq pj oq or gt pk ot ou ov ho bj">Nugen Audio created the VisLM plugin to provide sound teams with an efficient and accurate way to measure mixes for conformance to traditional broadcast &amp; streaming specifications — Full Mix Loudness, Dialogue Loudness, and True Peak. Since then, VisLM has become a widely used tool throughout the global post-production industry. Nugen Audio partnered with Fraunhofer, integrating the Fraunhofer IDMT Dialogue Intelligibility libraries into a new industry-first tool — Nugen DialogCheck. This tool gives <strong class="od it">re-recording mixers</strong> real-time insights, helping them adjust dialogue clarity at the most crucial points in the mixing process, ensuring every word is clear and understood.</p><figure class="pw px py pz qa qb pt pu paragraph-image"><div role="button" tabindex="0" class="qc qd fk qe bg qf"><div class="pt pu qm"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*gGt-DpKR806J2jqT%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*gGt-DpKR806J2jqT%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*gGt-DpKR806J2jqT%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*gGt-DpKR806J2jqT%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*gGt-DpKR806J2jqT%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*gGt-DpKR806J2jqT%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*gGt-DpKR806J2jqT%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*gGt-DpKR806J2jqT 640w, https://miro.medium.com/v2/resize:fit:720/0*gGt-DpKR806J2jqT 720w, https://miro.medium.com/v2/resize:fit:750/0*gGt-DpKR806J2jqT 750w, https://miro.medium.com/v2/resize:fit:786/0*gGt-DpKR806J2jqT 786w, https://miro.medium.com/v2/resize:fit:828/0*gGt-DpKR806J2jqT 828w, https://miro.medium.com/v2/resize:fit:1100/0*gGt-DpKR806J2jqT 1100w, https://miro.medium.com/v2/resize:fit:1400/0*gGt-DpKR806J2jqT 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /></picture></div></div></figure><h2 id="5c1f" class="ox oy is be oz gj pa dx gk gl pb dz gm gn pc go gp gq pd gr gs gt pe gu gv pf bj">Clearer Dialogue Through Collaboration</h2><p id="3af9" class="pw-post-body-paragraph ob oc is od b oe pg og oh oi ph ok ol gn pi on oo gq pj oq or gt pk ot ou ov ho bj">Crafting crystal-clear dialogue isn’t just a technical challenge — it’s an art that requires continuous innovation and strong industry collaboration. To empower creators, Netflix and its partners are embedding advanced intelligibility measurement tools directly into DAWs, giving sound teams the ability to:</p><ul class=""><li id="9e1e" class="ob oc is od b oe of og oh oi oj ok ol gn om on oo gq op oq or gt os ot ou ov pl pm pn bj">Detect and resolve dialogue clarity issues early in the mix.</li><li id="f1d8" class="ob oc is od b oe po og oh oi pp ok ol gn pq on oo gq pr oq or gt ps ot ou ov pl pm pn bj">Fine-tune speech intelligibility without compromising artistic intent.</li><li id="69a9" class="ob oc is od b oe po og oh oi pp ok ol gn pq on oo gq pr oq or gt ps ot ou ov pl pm pn bj">Deliver immersive, accessible storytelling to every viewer, in any listening environment.</li></ul><p id="ed32" class="pw-post-body-paragraph ob oc is od b oe of og oh oi oj ok ol gn om on oo gq op oq or gt os ot ou ov ho bj">At Netflix, we’re committed to pushing the boundaries of audio excellence. From pioneering the eSTOI (extended short-term objective intelligibility) method to collaborating with Fraunhofer and Nugen Audio on cutting-edge tools like the DialogCheck Plugin, we’re setting a new standard for dialogue clarity — ensuring every word is heard exactly as creators intended. But innovation doesn’t happen in isolation. By working together with our partners, we can continue to push the limits of what’s possible, fueling creativity and driving the future of storytelling.</p><p id="f0f9" class="pw-post-body-paragraph ob oc is od b oe of og oh oi oj ok ol gn om on oo gq op oq or gt os ot ou ov ho bj">Finally, we’d like to extend a heartfelt thanks to Scott Kramer for his contributions to this initiative.</p></div></div></div></div>]]></description>
      <link>https://netflixtechblog.com/measuring-dialogue-intelligibility-for-netflix-content-58c13d2a6f6e</link>
      <guid>https://netflixtechblog.com/measuring-dialogue-intelligibility-for-netflix-content-58c13d2a6f6e</guid>
      <pubDate>Thu, 08 May 2025 02:40:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[How Netflix Accurately Attributes eBPF Flow Logs]]></title>
      <description><![CDATA[<div><div></div><p id="84ec" class="pw-post-body-paragraph oc od io oe b of og oh oi oj ok ol om gp on oo op gs oq or os gv ot ou ov ow hp bk">By <a class="ag hc" href="https://www.linkedin.com/in/chengxie90/" rel="noopener ugc nofollow" target="_blank">Cheng Xie</a>, <a class="ag hc" href="https://www.linkedin.com/in/bryan-shultz-85983829/" rel="noopener ugc nofollow" target="_blank">Bryan Shultz</a>, and <a class="ag hc" href="https://www.linkedin.com/in/christine-xu-1b77191b/" rel="noopener ugc nofollow" target="_blank">Christine Xu</a></p><p id="be87" class="pw-post-body-paragraph oc od io oe b of og oh oi oj ok ol om gp on oo op gs oq or os gv ot ou ov ow hp bk">In a previous <a class="ag hc" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/how-netflix-uses-ebpf-flow-logs-at-scale-for-network-insight-e3ea997dca96">blog post</a>, we described how Netflix uses eBPF to capture TCP flow logs at scale for enhanced network insights. In this post, we delve deeper into how Netflix solved a core problem: accurately attributing flow IP addresses to workload identities.</p><h1 id="3fde" class="ox oy io bf oz pa pb pc gm pd pe pf go pg ph pi pj pk pl pm pn po pp pq pr ps bk">A Brief Recap</h1><p id="9147" class="pw-post-body-paragraph oc od io oe b of pt oh oi oj pu ol om gp pv oo op gs pw or os gv px ou ov ow hp bk"><strong class="oe ip">FlowExporter</strong> is a sidecar that runs alongside all Netflix workloads. It uses eBPF and <a class="ag hc" href="https://www.brendangregg.com/blog/2018-03-22/tcp-tracepoints.html" rel="noopener ugc nofollow" target="_blank">TCP tracepoints</a> to monitor TCP socket state changes. When a TCP socket closes, FlowExporter generates a flow log record that includes the IP addresses, ports, timestamps, and additional socket statistics. On average, 5 million records are produced per second.</p><p id="4fce" class="pw-post-body-paragraph oc od io oe b of og oh oi oj ok ol om gp on oo op gs oq or os gv ot ou ov ow hp bk">In cloud environments, IP addresses are reassigned to different workloads as workload instances are created and terminated, so IP addresses alone cannot provide insights on which workloads are communicating. To make the flow logs useful, each IP address must be attributed to its corresponding workload identity. <strong class="oe ip">FlowCollector</strong>, a backend service, collects flow logs from FlowExporter instances across the fleet, attributes the IP addresses, and sends these attributed flows to Netflix’s <a class="ag hc" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/data-mesh-a-data-movement-and-processing-platform-netflix-1288bcab2873">Data Mesh</a> for subsequent stream and batch processing.</p><p id="6fbd" class="pw-post-body-paragraph oc od io oe b of og oh oi oj ok ol om gp on oo op gs oq or os gv ot ou ov ow hp bk">The eBPF flow logs provide a comprehensive view of service topology and network health across Netflix’s extensive microservices fleet, regardless of the programming language, RPC mechanism, or application-layer protocol used by individual workloads.</p><h1 id="a716" class="ox oy io bf oz pa pb pc gm pd pe pf go pg ph pi pj pk pl pm pn po pp pq pr ps bk">The Problem with Misattribution</h1><p id="21b9" class="pw-post-body-paragraph oc od io oe b of pt oh oi oj pu ol om gp pv oo op gs pw or os gv px ou ov ow hp bk">Accurately attributing flow IP addresses to workload identities has been a significant challenge since our eBPF flow logs were introduced.</p><p id="413b" class="pw-post-body-paragraph oc od io oe b of og oh oi oj ok ol om gp on oo op gs oq or os gv ot ou ov ow hp bk">As noted in our previous blog post, our initial attribution approach relied on <a class="ag hc" href="https://youtu.be/8C9xNVYbCVk?si=Mqic7typcyB-v3JR&amp;t=1687" rel="noopener ugc nofollow" target="_blank">Sonar</a>, an internal IP address tracking service that emits an event whenever an IP address in Netflix’s AWS VPCs is assigned or unassigned to a workload. FlowCollector consumes a stream of IP address change events from Sonar and uses this information to attribute flow IP addresses in real-time.</p><figure class="qb qc qd qe qf qg py pz paragraph-image"><div role="button" tabindex="0" class="qh qi fm qj bh qk"><div class="py pz qa"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*QIn-JibEFM2CLans%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*QIn-JibEFM2CLans%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*QIn-JibEFM2CLans%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*QIn-JibEFM2CLans%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*QIn-JibEFM2CLans%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*QIn-JibEFM2CLans%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*QIn-JibEFM2CLans%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*QIn-JibEFM2CLans 640w, https://miro.medium.com/v2/resize:fit:720/0*QIn-JibEFM2CLans 720w, https://miro.medium.com/v2/resize:fit:750/0*QIn-JibEFM2CLans 750w, https://miro.medium.com/v2/resize:fit:786/0*QIn-JibEFM2CLans 786w, https://miro.medium.com/v2/resize:fit:828/0*QIn-JibEFM2CLans 828w, https://miro.medium.com/v2/resize:fit:1100/0*QIn-JibEFM2CLans 1100w, https://miro.medium.com/v2/resize:fit:1400/0*QIn-JibEFM2CLans 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh fx ql c" width="700" height="217" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="fd88" class="pw-post-body-paragraph oc od io oe b of og oh oi oj ok ol om gp on oo op gs oq or os gv ot ou ov ow hp bk">The fundamental drawback of this method is that it can lead to misattribution. Delays and failures are inevitable in distributed systems, which may delay IP address change events from reaching FlowCollector. For instance, an IP address may initially be assigned to workload X but later reassigned to workload Y. However, if the change event for this reassignment is delayed, FlowCollector will continue to assume that the IP address belongs to workload X, resulting in misattributed flows. Additionally, event timestamps may be inaccurate depending on how they are captured.</p><p id="7a27" class="pw-post-body-paragraph oc od io oe b of og oh oi oj ok ol om gp on oo op gs oq or os gv ot ou ov ow hp bk">Misattribution rendered the flow data unreliable for decision-making. Users often depend on flow logs to validate workload dependencies, but misattribution creates confusion. Without expert knowledge of expected dependencies, users would struggle to identify or confirm misattribution. Moreover, misattribution occurred frequently for critical services with a large footprint due to frequent IP address changes. Overall, misattribution makes fleet-wide dependency analysis impractical.</p><p id="29df" class="pw-post-body-paragraph oc od io oe b of og oh oi oj ok ol om gp on oo op gs oq or os gv ot ou ov ow hp bk">As a workaround, we made FlowCollector hold received flows for 15 minutes before attribution, allowing time for delayed IP address change events. While this approach reduced misattribution, it did not eliminate it. Moreover, the waiting period made the data less fresh, reducing its utility for real-time analysis.</p><p id="200e" class="pw-post-body-paragraph oc od io oe b of og oh oi oj ok ol om gp on oo op gs oq or os gv ot ou ov ow hp bk">Fully eliminating misattribution is crucial because it only takes a single misattributed flow to produce an incorrect workload dependency. Solving this problem required a complete rethinking of our approach. Over the past year, Netflix developed a new attribution method that has finally eliminated misattribution, as detailed in the rest of this post.</p><h1 id="276f" class="ox oy io bf oz pa pb pc gm pd pe pf go pg ph pi pj pk pl pm pn po pp pq pr ps bk">Attributing Local IP Addresses</h1><p id="994e" class="pw-post-body-paragraph oc od io oe b of pt oh oi oj pu ol om gp pv oo op gs pw or os gv px ou ov ow hp bk">Each socket has two IP addresses: a local IP address and a remote IP address. Previously, we used the same method to attribute both. However, attributing the local IP address should be a simpler task since the local IP address belongs to the instance where FlowExporter captures the socket. Therefore, FlowExporter should determine the local workload identity from its environment and attribute the local IP address before sending the flow to FlowCollector.</p><p id="da55" class="pw-post-body-paragraph oc od io oe b of og oh oi oj ok ol om gp on oo op gs oq or os gv ot ou ov ow hp bk">This is straightforward for workloads running directly on EC2 instances, as Netflix’s <a class="ag hc" href="https://www.youtube.com/watch?v=-mmOT9I6JlY" rel="noopener ugc nofollow" target="_blank">Metatron</a> provisions workload identity certificates to each EC2 instance at boot time. FlowExporter can simply read these certificates from the local disk to determine the local workload identity.</p><p id="81e0" class="pw-post-body-paragraph oc od io oe b of og oh oi oj ok ol om gp on oo op gs oq or os gv ot ou ov ow hp bk">Attributing local IP addresses for container workloads running on Netflix’s container platform, <a class="ag hc" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/titus-the-netflix-container-management-platform-is-now-open-source-f868c9fb5436">Titus</a>, is more challenging. FlowExporter runs at the container host level, where each host manages multiple container workloads with different identities. When FlowExporter’s eBPF programs receive a socket event from TCP tracepoints in the kernel, the socket may have been created by one of the container workloads or by the host itself. Therefore, FlowExporter must determine which workload to attribute the socket’s local IP address to. To solve this problem, we leveraged <a class="ag hc" href="https://www.youtube.com/watch?v=fmUM9bMoCNE" rel="noopener ugc nofollow" target="_blank">IPMan</a>, Netflix’s container IP address assignment service. IPManAgent, a daemon running on every container host, is responsible for assigning and unassigning IP addresses. As container workloads are launched, IPManAgent writes an IP-address-to-workload-ID mapping to an eBPF map, which FlowExporter’s eBPF programs can then use to look up the workload ID associated with a socket local IP address.</p><figure class="qb qc qd qe qf qg py pz paragraph-image"><div role="button" tabindex="0" class="qh qi fm qj bh qk"><div class="py pz qa"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*fyfZ6m2NrMq1NgRQ%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*fyfZ6m2NrMq1NgRQ%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*fyfZ6m2NrMq1NgRQ%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*fyfZ6m2NrMq1NgRQ%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*fyfZ6m2NrMq1NgRQ%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*fyfZ6m2NrMq1NgRQ%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*fyfZ6m2NrMq1NgRQ%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*fyfZ6m2NrMq1NgRQ 640w, https://miro.medium.com/v2/resize:fit:720/0*fyfZ6m2NrMq1NgRQ 720w, https://miro.medium.com/v2/resize:fit:750/0*fyfZ6m2NrMq1NgRQ 750w, https://miro.medium.com/v2/resize:fit:786/0*fyfZ6m2NrMq1NgRQ 786w, https://miro.medium.com/v2/resize:fit:828/0*fyfZ6m2NrMq1NgRQ 828w, https://miro.medium.com/v2/resize:fit:1100/0*fyfZ6m2NrMq1NgRQ 1100w, https://miro.medium.com/v2/resize:fit:1400/0*fyfZ6m2NrMq1NgRQ 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh fx ql c" width="700" height="438" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="9514" class="pw-post-body-paragraph oc od io oe b of og oh oi oj ok ol om gp on oo op gs oq or os gv ot ou ov ow hp bk">Another challenge was to accommodate Netflix’s <a class="ag hc" href="https://lpc.events/event/11/contributions/932/attachments/908/1764/LPC%202021_%20Talking%20IPv6%20to%20IPv4%20Without%20NAT_2.pdf" rel="noopener ugc nofollow" target="_blank">IPv6 to IPv4 translation mechanism</a> on Titus. To facilitate IPv6 migration, Netflix developed a mechanism that enables IPv6-only containers to communicate with IPv4 destinations without incurring NAT64 overhead. This mechanism intercepts connect syscalls and replaces the underlying socket with one that uses a shared IPv4 address assigned to the container host. This confuses FlowExporter because the kernel reports the same local IPv4 address for sockets created by different container workloads. To disambiguate, local port information is additionally required. We modified Titus to write a mapping of (local IPv4 address, local port) to the workload ID into an eBPF map whenever a connect syscall is intercepted. FlowExporter’s eBPF programs then use this map to correctly attribute sockets created by the translation mechanism.</p><p id="9cd6" class="pw-post-body-paragraph oc od io oe b of og oh oi oj ok ol om gp on oo op gs oq or os gv ot ou ov ow hp bk">With these problems solved, we can now accurately attribute the local IP address of every flow.</p><h1 id="d60f" class="ox oy io bf oz pa pb pc gm pd pe pf go pg ph pi pj pk pl pm pn po pp pq pr ps bk">Attributing Remote IP Addresses</h1><p id="0edb" class="pw-post-body-paragraph oc od io oe b of pt oh oi oj pu ol om gp pv oo op gs pw or os gv px ou ov ow hp bk">Once the local IP address attribution problem is solved, accurately attributing remote IP addresses becomes feasible. Now, each flow reported by FlowExporter includes the local IP address, the local workload identity, and connection start/end timestamps. As FlowCollector receives these flows, it can learn the time ranges during which each workload owns a given IP address. For instance, if FlowCollector sees a flow with local IP address 10.0.0.1 associated with workload X that starts at t1 and ends at t2, it can deduce that 10.0.0.1 belonged to workload X from t1 to t2. Since Netflix uses <a class="ag hc" href="https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/set-time.html" rel="noopener ugc nofollow" target="_blank">Amazon Time Sync</a> across its fleet, the timestamps (captured by FlowExporter) are reliable.</p><p id="7cd7" class="pw-post-body-paragraph oc od io oe b of og oh oi oj ok ol om gp on oo op gs oq or os gv ot ou ov ow hp bk">The FlowCollector service cluster consists of many nodes. Every node must be capable of attributing arbitrary remote IP addresses and, therefore, requires knowledge of all workload IP addresses and their recent ownership records. To represent this knowledge, each node maintains an in-memory hashmap that maps an IP address to a list of time ranges, as illustrated by the following Go structs:</p><pre class="qb qc qd qe qf qm qn qo bp qp bb bk">type IPAddressTracker struct {<br />    ipToTimeRanges map[netip.Addr]timeRanges<br />}type timeRanges []timeRangetype timeRange struct {<br />    workloadID   string<br />    start        time.Time<br />    end          time.Time<br />}</pre><p id="5dda" class="pw-post-body-paragraph oc od io oe b of og oh oi oj ok ol om gp on oo op gs oq or os gv ot ou ov ow hp bk">To populate the hashmap, FlowCollector extracts the local IP address, local workload identity, start time, and end time from each received flow and creates/extends the corresponding time ranges in the map. The time ranges for each IP address are sorted in ascending order, and they are non-overlapping since an IP address cannot belong to two different workloads simultaneously.</p><p id="c274" class="pw-post-body-paragraph oc od io oe b of og oh oi oj ok ol om gp on oo op gs oq or os gv ot ou ov ow hp bk">Since each flow is only sent to one FlowCollector node, each node must share the time ranges it learned from received flows with other nodes. We implemented a broadcasting mechanism using Kafka, where each node publishes learned time ranges to all other nodes. Although more efficient broadcasting implementations exist, the Kafka-based approach is simple and has worked well for us.</p><p id="68ab" class="pw-post-body-paragraph oc od io oe b of og oh oi oj ok ol om gp on oo op gs oq or os gv ot ou ov ow hp bk">Now, FlowCollector can attribute remote IP addresses by looking them up in the populated map, which returns a list of time ranges. It then uses the flow’s start timestamp to determine the corresponding time range and associated workload identity. If the start time does not fall within any time range, FlowCollector will retry after a delay, eventually giving up if the retry fails. Such failures may occur when flows are lost or broadcast messages are delayed. For our use cases, it is acceptable to leave a small percentage of flows unattributed, but any misattribution is unacceptable.</p><p id="1a45" class="pw-post-body-paragraph oc od io oe b of og oh oi oj ok ol om gp on oo op gs oq or os gv ot ou ov ow hp bk">This new method achieves accurate attribution thanks to the continuous heartbeats, each associated with a reliable time range of IP address ownership. It handles transient issues gracefully — a few delayed or lost heartbeats do not lead to misattribution. In contrast, the previous method relied solely on discrete IP address assignment and unassignment events. Lacking heartbeats, it had to presume an IP address remained assigned until notified otherwise (which can be hours or days later), making it vulnerable to misattribution when the notifications were delayed.</p><figure class="qb qc qd qe qf qg py pz paragraph-image"><div role="button" tabindex="0" class="qh qi fm qj bh qk"><div class="py pz qa"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*o8tJzaxRlWDBIYBS%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*o8tJzaxRlWDBIYBS%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*o8tJzaxRlWDBIYBS%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*o8tJzaxRlWDBIYBS%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*o8tJzaxRlWDBIYBS%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*o8tJzaxRlWDBIYBS%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*o8tJzaxRlWDBIYBS%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*o8tJzaxRlWDBIYBS 640w, https://miro.medium.com/v2/resize:fit:720/0*o8tJzaxRlWDBIYBS 720w, https://miro.medium.com/v2/resize:fit:750/0*o8tJzaxRlWDBIYBS 750w, https://miro.medium.com/v2/resize:fit:786/0*o8tJzaxRlWDBIYBS 786w, https://miro.medium.com/v2/resize:fit:828/0*o8tJzaxRlWDBIYBS 828w, https://miro.medium.com/v2/resize:fit:1100/0*o8tJzaxRlWDBIYBS 1100w, https://miro.medium.com/v2/resize:fit:1400/0*o8tJzaxRlWDBIYBS 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh fx ql c" width="700" height="520" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="704c" class="pw-post-body-paragraph oc od io oe b of og oh oi oj ok ol om gp on oo op gs oq or os gv ot ou ov ow hp bk">One detail is that when FlowCollector receives a flow, it cannot attribute its remote IP address right away because it requires the latest observed time ranges for the remote IP address. Since FlowExporter reports flows in batches every minute, FlowCollector must wait until it receives the flow batch from the remote workload FlowExporter for the last minute, which may not have arrived yet. To address this, FlowCollector temporarily stores received flows on disk for one minute before attributing their remote IP addresses. This introduces a 1-minute delay, but it is much shorter than the 15-minute delay with the previous approach.</p><p id="827b" class="pw-post-body-paragraph oc od io oe b of og oh oi oj ok ol om gp on oo op gs oq or os gv ot ou ov ow hp bk">In addition to producing accurate attribution, the new method is also cost-effective thanks to its simplicity and in-memory lookups. Because the in-memory state can be quickly rebuilt when a FlowCollector node starts up, no persistent storage is required. With 30 c7i.2xlarge instances, we can process 5 million flows per second across the entire Netflix fleet.</p><h1 id="7687" class="ox oy io bf oz pa pb pc gm pd pe pf go pg ph pi pj pk pl pm pn po pp pq pr ps bk">Attributing Cross-Regional IP Addresses</h1><p id="26ad" class="pw-post-body-paragraph oc od io oe b of pt oh oi oj pu ol om gp pv oo op gs pw or os gv px ou ov ow hp bk">For simplicity, we have so far glossed over one topic: regionalization. Netflix’s cloud microservices operate across multiple AWS regions. To optimize flow reporting and minimize cross-regional traffic, a FlowCollector cluster runs in each major region, and FlowExporter agents send flows to their corresponding regional FlowCollector. When FlowCollector receives a flow, its local IP address is guaranteed to be within the region.</p><p id="cdfc" class="pw-post-body-paragraph oc od io oe b of og oh oi oj ok ol om gp on oo op gs oq or os gv ot ou ov ow hp bk">To minimize cross-region traffic, the broadcasting mechanism is limited to FlowCollector nodes within the same region. Consequently, the IP address time ranges map contains only IP addresses from that region. However, cross-regional flows have a remote IP address in a different region. To attribute these flows, the receiving FlowCollector node forwards them to nodes in the corresponding region. FlowCollector determines the region for a remote IP address by looking up a trie built from all Netflix VPC CIDRs. This approach is more efficient than broadcasting IP address time range updates across all regions, as only 1% of Netflix flows are cross-regional.</p><h1 id="f8f3" class="ox oy io bf oz pa pb pc gm pd pe pf go pg ph pi pj pk pl pm pn po pp pq pr ps bk">Attributing Non-Workload IP Addresses</h1><p id="5a96" class="pw-post-body-paragraph oc od io oe b of pt oh oi oj pu ol om gp pv oo op gs pw or os gv px ou ov ow hp bk">So far, FlowCollector can accurately attribute IP addresses belonging to Netflix’s cloud workloads. However, not all flow IP addresses fall into this category. For instance, a significant portion of flows goes through AWS ELBs. For these flows, their remote IP addresses are associated with the ELBs, where we cannot run FlowExporter. Consequently, FlowCollector cannot determine their identities by simply observing the received flows. To attribute these remote IP addresses, we continue to use IP address change events from Sonar, which crawls AWS resources to detect changes in IP address assignments. Although this data stream may contain inaccurate timestamps and be delayed, misattribution is not a main concern since ELB IP address reassignment occurs very infrequently.</p><h1 id="4953" class="ox oy io bf oz pa pb pc gm pd pe pf go pg ph pi pj pk pl pm pn po pp pq pr ps bk">Verifying Correctness</h1><p id="6cc9" class="pw-post-body-paragraph oc od io oe b of pt oh oi oj pu ol om gp pv oo op gs pw or os gv px ou ov ow hp bk">Verifying that the new method has eliminated misattribution is challenging due to the lack of a definitive source of truth for workload dependencies to validate flow logs against; the flow logs themselves are intended to serve as this source of truth, after all. To build confidence, we analyzed the flow logs of a large service with well-understood dependencies. A large footprint is necessary, as misattribution is more prevalent in services with numerous instances, and there must be a reliable method to determine the dependencies for this service without relying on flow logs.</p><p id="c413" class="pw-post-body-paragraph oc od io oe b of og oh oi oj ok ol om gp on oo op gs oq or os gv ot ou ov ow hp bk">Netflix’s cloud gateway, <a class="ag hc" href="https://github.com/Netflix/zuul" rel="noopener ugc nofollow" target="_blank">Zuul</a>, served this purpose perfectly due to its extensive footprint (handling all cloud ingress traffic), its large number of downstream dependencies, and our ability to derive its dependencies from its routing configurations as the source of truth for comparison with flow logs. We found no misattribution for flows through Zuul over a two-week window. This provided strong confidence that the new attribution method has eliminated misattribution. In the previous approach, approximately 40% of Zuul’s dependencies reported by the flow logs were misattributed.</p><h1 id="57d4" class="ox oy io bf oz pa pb pc gm pd pe pf go pg ph pi pj pk pl pm pn po pp pq pr ps bk">Conclusion</h1><p id="bb9c" class="pw-post-body-paragraph oc od io oe b of pt oh oi oj pu ol om gp pv oo op gs pw or os gv px ou ov ow hp bk">With misattribution solved, eBPF flow logs now deliver dependable, fleet-wide insights into Netflix’s service topology and network health. This advancement unlocks numerous exciting opportunities in areas such as service dependency auditing, security analysis, and incident triage, while helping Netflix engineers develop a better understanding of our ever-evolving distributed systems.</p><h1 id="4224" class="ox oy io bf oz pa pb pc gm pd pe pf go pg ph pi pj pk pl pm pn po pp pq pr ps bk">Acknowledgments</h1><p id="28be" class="pw-post-body-paragraph oc od io oe b of pt oh oi oj pu ol om gp pv oo op gs pw or os gv px ou ov ow hp bk">We would like to thank <a class="ag hc" href="https://www.linkedin.com/in/mdubcovsky/" rel="noopener ugc nofollow" target="_blank">Martin Dubcovsky</a>, <a class="ag hc" href="https://www.linkedin.com/in/joannekoong/" rel="noopener ugc nofollow" target="_blank">Joanne Koong</a>, <a class="ag hc" href="https://www.linkedin.com/in/troshko/" rel="noopener ugc nofollow" target="_blank">Taras Roshko</a>, <a class="ag hc" href="https://www.linkedin.com/in/nabilschear/" rel="noopener ugc nofollow" target="_blank">Nabil Schear</a>, <a class="ag hc" href="https://www.linkedin.com/in/jacobmeyers35/" rel="noopener ugc nofollow" target="_blank">Jacob Meyers</a>, <a class="ag hc" href="https://www.linkedin.com/in/parshap/" rel="noopener ugc nofollow" target="_blank">Parsha Pourkhomami</a>, <a class="ag hc" href="https://www.linkedin.com/in/hechaoli/" rel="noopener ugc nofollow" target="_blank">Hechao Li</a>, <a class="ag hc" href="https://www.linkedin.com/in/donavanfritz/" rel="noopener ugc nofollow" target="_blank">Donavan Fritz</a>, <a class="ag hc" href="https://www.linkedin.com/in/rob-gulewich-0335b52/" rel="noopener ugc nofollow" target="_blank">Rob Gulewich</a>, <a class="ag hc" href="https://www.linkedin.com/in/amanda-li-410286166/" rel="noopener ugc nofollow" target="_blank">Amanda Li</a>, <a class="ag hc" href="https://www.linkedin.com/in/jdsalem/" rel="noopener ugc nofollow" target="_blank">John Salem</a>, <a class="ag hc" href="https://www.linkedin.com/in/haananth/" rel="noopener ugc nofollow" target="_blank">Hariharan Ananthakrishnan</a>, <a class="ag hc" href="https://www.linkedin.com/in/joshmachine/" rel="noopener ugc nofollow" target="_blank">Keerti Lakshminarayan</a>, and other stunning colleagues for their feedback, inspiration, and contributions to the success of this effort.</p></div>]]></description>
      <link>https://netflixtechblog.com/how-netflix-accurately-attributes-ebpf-flow-logs-afe6d644a3bc</link>
      <guid>https://netflixtechblog.com/how-netflix-accurately-attributes-ebpf-flow-logs-afe6d644a3bc</guid>
      <pubDate>Tue, 08 Apr 2025 19:50:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Globalizing Productions with Netflix’s Media Production Suite]]></title>
      <description><![CDATA[<div><div></div><p id="5f94" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk"><a class="ag ih" href="https://www.linkedin.com/in/jesse-korosi-44790985/" rel="noopener ugc nofollow" target="_blank"><strong class="pd ju">Jesse Korosi</strong></a>,<a class="ag ih" href="https://www.linkedin.com/in/thijsvdkamp/" rel="noopener ugc nofollow" target="_blank"><strong class="pd ju">Thijs van de Kamp</strong></a><strong class="pd ju">, </strong><a class="ag ih" href="https://www.linkedin.com/in/mayralvega/" rel="noopener ugc nofollow" target="_blank"><strong class="pd ju">Mayra Vega</strong></a>,<a class="ag ih" href="https://www.linkedin.com/in/laurafuturo/" rel="noopener ugc nofollow" target="_blank"><strong class="pd ju">Laura Futuro</strong></a>,<a class="ag ih" href="https://www.linkedin.com/in/margoline/" rel="noopener ugc nofollow" target="_blank"><strong class="pd ju">Anton Margoline</strong></a></p><p id="4329" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk">The journey from script to screen is full of challenges in the ever-evolving world of film and television. The industry has always innovated, and over the last decade, it started moving towards cloud-based workflows. However, unlocking cloud innovation and all its benefits on a global scale has proven to be difficult. The opportunity is clear: streamline complex media management logistics, eliminate tedious, non-creative task-based work and enable productions to focus on what matters most–creative storytelling. With these challenges in mind, Netflix has developed a suite of tools by filmmakers for filmmakers: the Media Production Suite (MPS).</p><figure class="pz qa qb qc qd qe pw px paragraph-image"><div role="button" tabindex="0" class="qf qg gs qh bh qi"><div class="pw px py"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*CUGxiNprXnLcOmhI%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*CUGxiNprXnLcOmhI%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*CUGxiNprXnLcOmhI%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*CUGxiNprXnLcOmhI%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*CUGxiNprXnLcOmhI%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*CUGxiNprXnLcOmhI%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*CUGxiNprXnLcOmhI%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*CUGxiNprXnLcOmhI 640w, https://miro.medium.com/v2/resize:fit:720/0*CUGxiNprXnLcOmhI 720w, https://miro.medium.com/v2/resize:fit:750/0*CUGxiNprXnLcOmhI 750w, https://miro.medium.com/v2/resize:fit:786/0*CUGxiNprXnLcOmhI 786w, https://miro.medium.com/v2/resize:fit:828/0*CUGxiNprXnLcOmhI 828w, https://miro.medium.com/v2/resize:fit:1100/0*CUGxiNprXnLcOmhI 1100w, https://miro.medium.com/v2/resize:fit:1400/0*CUGxiNprXnLcOmhI 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh hc qj c" width="700" height="210" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><h1 id="ca35" class="qk ql jt bf qm qn qo qp hp qq qr qs hr qt qu qv qw qx qy qz ra rb rc rd re rf bk"><strong class="am">What are we solving for?</strong></h1><p id="0953" class="pw-post-body-paragraph pb pc jt pd b pe rg pg ph pi rh pk pl hs ri pn po hv rj pq pr hy rk pt pu pv iu bk">Significant time and resources are devoted to managing media logistics throughout the production lifecycle. An average Netflix title produces around ~200 Terabytes of Original Camera Files (OCF), with outliers up to 700 Terabytes, not including any work-in-progress files, VFX assets, 3D assets, etc. The data produced on set is traditionally copied to physical tape stock like LTO. This workflow has been considered the industry norm for a long time and may be cost-effective, but comes with trade-offs. Aside from needing to physically ship and track all movement of the tape stock, storing media on a physical tape makes it harder to search, play and share media assets; slowing down accessibility to production media when needed, especially when titles need to collaborate with talent and vendors all over the world.</p><p id="8c3a" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk">Even when workflows are fully digital, the distribution of media between multiple departments and vendors can still be challenging. A lack of automation and standardization often results in a labour-intensive process across post-production and VFX with a lot of dependencies that introduce potential human errors and security risks. Many productions utilize a large variety of vendors, making this collaboration a large technical puzzle. As file sizes grow and workflows become more complex, these issues are magnified, leading to inefficiencies that slow down post-production and reduce the available time spent on creative work.</p><p id="beaa" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk">Moving media into the cloud introduces new challenges for production and post ramping up to meet the operational and technological hurdles this poses. For some post-production facilities, it’s not uncommon to see a wall of portable hard drives at their facility, with media being hand-carried between vendors because alternatives are not available. The need for a centralized, cloud-based solution that transcends these barriers is more pressing than ever. This results in a willingness to embrace new and innovative ideas, even if exploratory, and introduce drastic workflow changes to productions in pursuit of creative evolution.</p><p id="a501" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk">At Netflix, we believe that great stories can come from anywhere, but we have seen that technical limitations in traditional workflows reduce access to media and restrict filmmakers’ access to talent. Besides the need for robust cloud storage for their media, artists need access to powerful workstations and real-time playback. Depending on the market, or production budget, cutting-edge technology might not be available or affordable.</p><p id="1726" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk">What if we started charting a course to break free from many of these technical limitations and found ways to enhance creativity? Industry trade shows like the International Broadcast Convention (IBC) and the National Association of Broadcasters Show (NAB) highlight a strong global trend: instead of bringing media to the artist/applications (traditional workflow) we see the shift towards bringing people and applications to the media (cloud workflows and remote workstations). The concept of cloud-based workflows is not new, as many technology leaders in our industry have been experimenting in this space for more than a decade. However, executing this vision at a Netflix scale with hundreds of titles a year has not been done before…</p><h1 id="4f2c" class="qk ql jt bf qm qn qo qp hp qq qr qs hr qt qu qv qw qx qy qz ra rb rc rd re rf bk"><strong class="am">The challenge of building a global technology to solve this</strong></h1><p id="8b30" class="pw-post-body-paragraph pb pc jt pd b pe rg pg ph pi rh pk pl hs ri pn po hv rj pq pr hy rk pt pu pv iu bk">Building solutions at a global scale poses significant challenges. The art of making movies and series lacks equal access to technology, best practices, and global standardization. Different countries worldwide are at different phases of innovation based on local needs and nuances. While some regions boast over a century of cinematic history and have a strong industry, others are just beginning to carve their niche. This vast gap presents a unique challenge: developing global technology that caters to both established and emerging markets, each with distinct languages and workflows.</p><p id="a919" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk">The large diversity of needs by talent and vendors globally creates a standardization challenge and can be seen when productions use a global talent pool. Many mature post-production and VFX facilities have built scripts and automation that flow between various artists and personnel within their facility; allowing a more streamlined workflow, even though the customization is time-consuming. E.g., Transcoding, or transcriptions that automatically run when files are dropped in a hot folder, with the expectation that certain sidecar metadata files will accompany them with a specific organizational structure. Embracing and integrating new workflows introduces the fear of disrupting a well-established process, increasing additional pressure on the profit margins of vendors. Small workflow changes that may seem arbitrary may actually have a large impact on vendors. Therefore, innovation should provide meaningful benefits to a title in order to get adopted at scale. Reliability, a proven track record, strong support, and an incredibly low tolerance for bugs, or issues are top of mind in well-established markets.</p><p id="c8c2" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk">In developing this suite, we recognized the necessity of addressing the vast array of titles that flow through Netflix without the luxury of expanding into a massive operational entity. Consequently, automation became imperative. The intricacies of color and framing management, along with deliverables, must be seamlessly controlled and effortlessly managed by the user, without the need for manual intervention. Therefore, we cannot lean into humans configuring JSON files behind the scenes to map camera formats into deliverables. By embracing open standards, we not only streamline these processes but also facilitate smoother collaboration across diverse markets and countries, ensuring that our global productions can operate with unparalleled efficiency and cohesion. To ensure this, we’ve decided to lean heavily into standards like <a class="ag ih" href="https://www.oscars.org/science-technology/sci-tech-projects/aces" rel="noopener ugc nofollow" target="_blank">ACES</a>, <a class="ag ih" href="https://acescentral.com/knowledge-base-2/when-is-amf-used/" rel="noopener ugc nofollow" target="_blank">AMF</a>, <a class="ag ih" href="https://theasc.com/society/ascmitc/asc-media-hash-list" rel="noopener ugc nofollow" target="_blank">ASC MHL</a>, <a class="ag ih" href="https://theasc.com/society/ascmitc/asc-framing-decision-list" rel="noopener ugc nofollow" target="_blank">ASC FDL</a>, and <a class="ag ih" href="https://github.com/OpenTimelineIO" rel="noopener ugc nofollow" target="_blank">OTIO</a>. ACES and AMF for color pipeline management. ASC MHL for any file management/verifications. ASC FDL will serve as our framing interoperability and OTIO for any timeline interchange. Leaning into standards like this means that many things can be automated at scale and more importantly, high-complexity workflows can be offered to markets or shows that don’t normally have access to them. As an example, if a show is shot on various camera formats all framed and recorded at different resolutions, with different lenses and different safeties on each frame. The task of normalizing all of these for a VFX vendor into one common container with a normalized center extracted frame is often only offered to very high-end titles, considering it takes a human behind the curtain to create all of these mappings. But by leaning into a standard like the FDL, it means this can now easily be automated, and the control for these mappings, put directly in the hands of users.</p><h1 id="56f4" class="qk ql jt bf qm qn qo qp hp qq qr qs hr qt qu qv qw qx qy qz ra rb rc rd re rf bk"><strong class="am">Our Answer — Content Hub’s Media Production Suite (MPS)</strong></h1><figure class="pz qa qb qc qd qe"><div class="rl cr l gs"><figcaption class="ro go rp pw px rq rr bf b bg z cm">Introducing Content Hub Media Production Suite video</figcaption></div></figure><p id="e1a8" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk">Building a global scalable solution that could be utilized in a diversity of markets has been an exciting challenge. We set out to provide customizable and feature-rich tooling for advanced users while remaining intuitive and streamlined enough for less experienced filmmakers. With collaboration from Netflix teams, vendors, and talent across the globe, we’ve taken a bold step forward in enabling a suite of tools inside Netflix Content Hub that democratizes technology: the Media Production Suite. While leveraging our scale economies and access to resources, we can now unlock global talent pools for our productions, drastically reduce non-creative task-based work, streamline workflows, and level the playing field between our markets, ultimately maximizing the time available for what matters most; creative work!</p><h1 id="657d" class="qk ql jt bf qm qn qo qp hp qq qr qs hr qt qu qv qw qx qy qz ra rb rc rd re rf bk">So what is it?</h1><p id="2737" class="pw-post-body-paragraph pb pc jt pd b pe rg pg ph pi rh pk pl hs ri pn po hv rj pq pr hy rk pt pu pv iu bk">1. <strong class="pd ju">Netflix Hybrid Infrastructure</strong>: Netflix has invested in a hybrid infrastructure, a mix of cloud-based and physically distributed capabilities operating in multiple locations across the world and close to our productions to optimize user performance. This infrastructure is available for Netflix shows and is foundational under Content Hub’s Media Production Suite tooling. Local storage and compute services are connected through the Netflix Open Connect network (Netflix Content Delivery Network) to the infrastructure of Amazon Web Services (AWS). The system facilitates large volumes of camera and sound media and is built for speed. In order to ensure that productions have sufficient upload speeds to get their media into the cloud, Netflix has started to roll out Content Hub Ingest Centers globally to provide high-speed internet connectivity where required. With all media centralized, MPS eliminates the need for physical media transport and reduces the risk of human error. This approach not only streamlines operations but also enhances security and accessibility.</p><p id="507e" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk">2. <strong class="pd ju">Automation and Tooling</strong>: In addition to the Netflix Hybrid infrastructure layer, MPS consists of a suite of tools that tap into the media in the Netflix ecosystem.</p><p id="8ae6" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk"><strong class="pd ju">Footage Ingest</strong> — An application that allows users to upload media/files into Content Hub.</p><p id="e3f8" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk"><strong class="pd ju">Media Library</strong> — A central library that allows users to search, preview, share and download media.</p><p id="7864" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk"><strong class="pd ju">Dailies</strong> — A workflow, backed by an operational team, offering automated Quality Control of your footage, sound sync, application of color, rendering, and delivering dailies directly to editorial.</p><p id="4061" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk"><strong class="pd ju">Remote Workstations</strong> — Offering access to remote editorial workstations and storage for post-production needs.</p><p id="0465" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk"><strong class="pd ju">VFX Pulls</strong> — An automated method for converting and delivering visual effects plates, associated color, and framing files to VFX vendors.</p><p id="dc0b" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk"><strong class="pd ju">Conform Pulls</strong> — An automated method for consolidating, trimming, and delivering all OCF to picture-finishing vendors.</p><p id="0506" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk"><strong class="pd ju">Media Downloader</strong> — An automated download tool that initiates a download once media has been made available in the Netflix cloud.</p><p id="8a17" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk">While each of the individual tools within MPS is at different states of maturity, over 350 titles have made use of at least one of the tools noted above. Input has been taken from all over the world while developing, with users ranging from UCAN (United States/Canada), EMEA (Europe, Middle East, and Africa), SEA (South East Asia), LATAM (Latin America), and APAC (Asia Pacific).</p><h1 id="c95b" class="qk ql jt bf qm qn qo qp hp qq qr qs hr qt qu qv qw qx qy qz ra rb rc rd re rf bk"><strong class="am">Senna: Early Adoption and Insightful Feedback Driving MPS Evolution</strong></h1><figure class="pz qa qb qc qd qe pw px paragraph-image"><div role="button" tabindex="0" class="qf qg gs qh bh qi"><div class="pw px rs"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*NLeFw4FDGx2jsZg7%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*NLeFw4FDGx2jsZg7%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*NLeFw4FDGx2jsZg7%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*NLeFw4FDGx2jsZg7%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*NLeFw4FDGx2jsZg7%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*NLeFw4FDGx2jsZg7%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*NLeFw4FDGx2jsZg7%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*NLeFw4FDGx2jsZg7 640w, https://miro.medium.com/v2/resize:fit:720/0*NLeFw4FDGx2jsZg7 720w, https://miro.medium.com/v2/resize:fit:750/0*NLeFw4FDGx2jsZg7 750w, https://miro.medium.com/v2/resize:fit:786/0*NLeFw4FDGx2jsZg7 786w, https://miro.medium.com/v2/resize:fit:828/0*NLeFw4FDGx2jsZg7 828w, https://miro.medium.com/v2/resize:fit:1100/0*NLeFw4FDGx2jsZg7 1100w, https://miro.medium.com/v2/resize:fit:1400/0*NLeFw4FDGx2jsZg7 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh hc qj c" width="700" height="402" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="ro go rp pw px rq rr bf b bg z cm"><em class="ku">Media from the Brazilian-produced series ‘Senna’ being reviewed in MPS</em></figcaption></figure><p id="c6d3" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk">The Brazilian-produced series <em class="rt">Senna</em>, which follows the life of legendary Formula 1 driver Ayrton Senna, utilized MPS to reshape their content creation workflow, overcome geographical barriers, and unlock innovation to support world-class storytelling for a global audience. <em class="rt">Senna</em> is a groundbreaking series, not just for its storytelling but for its production journey across Argentina, Uruguay, Brazil, and the United Kingdom. With editorial teams spread across Porto Alegre and Spain, and VFX studios collaborating across locations in Brazil, Canada, the United States, and India, all orchestrated by our subsidiary Scanline VFX. The series exemplifies the global nature of modern filmmaking and was the perfect fit for Netflix’s new Content Hub Media Production Suite (MPS).</p><p id="2fcb" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk">At the heart of <em class="rt">Senna’s</em> workflow orchestration is MPS. While each of the tools within MPS is based on an opt-in model, in order to use many of the downstream services, the first step is ensuring that the original camera files (OCF) and original sound files (OSF) are uploaded. “<em class="rt">We knew we were going to shoot in different places,</em>” said Post Supervisor Gabriel Queiroz,<em class="rt">“to have all this material cloud-based, it’s definitely one of the most important things for us. It would be hard to bring all this media physically from Argentina or wherever to Brazil. It will take us a lot of time.”</em> With <em class="rt">Senna</em> shooting across locations, allowing production the capability of uploading their OCF and OSF resulted in no longer requiring shuttling hard drives on airplanes, creating LTO tapes, &amp; managing physical shipments for their negative. And yes, you read that correctly; when utilizing MPS, we don’t require LTO tapes to be written unless there are title-specific needs.</p><p id="838c" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk">With <em class="rt">Senna</em> beginning production back in June of 2023, our investment in MPS was still very early stages, and the tooling was considered beta. However, with the help, feedback, and partnership from this production, it was quickly realized that the investment was worth doubling down on. Since the early version used on <em class="rt">Senna</em>, Netflix has been spinning up ingest centers around the world, where drives can be dropped off, and within a matter of hours, all original camera files are uploaded into the Netflix ecosystem. While creating the ability to upload is not a novel concept, behind the scenes, it’s far from simple. Once a drive has been plugged in and our Netflix Footage Ingest application is opened, a validation is run, ensuring all expected media from set is on the drive. After media has been uploaded and checksums are run validating media integrity, all media is inspected, metadata is extracted, and assets are created for viewing/sharing/downloading with playable proxies. All media is then automatically backed up to a second tier of cloud-based storage for the final archive.</p><p id="e51e" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk">Traditionally, if you wanted to check in with your post vendor on how things are going for each of these media management steps noted above, or whether or not you can clear on set camera cards if you haven’t gotten a completion notification, you would have to pick up the phone and call your vendor. For <em class="rt">Senna</em>, anyone who wanted visibility on progress, simply logged in to Content Hub and could see any activity in the Footage Ingest dashboard, as well as look up any information needed on past uploads.</p><figure class="pz qa qb qc qd qe pw px paragraph-image"><div role="button" tabindex="0" class="qf qg gs qh bh qi"><div class="pw px rs"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*7gaSH8-YpnOTKnpu%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*7gaSH8-YpnOTKnpu%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*7gaSH8-YpnOTKnpu%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*7gaSH8-YpnOTKnpu%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*7gaSH8-YpnOTKnpu%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*7gaSH8-YpnOTKnpu%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*7gaSH8-YpnOTKnpu%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*7gaSH8-YpnOTKnpu 640w, https://miro.medium.com/v2/resize:fit:720/0*7gaSH8-YpnOTKnpu 720w, https://miro.medium.com/v2/resize:fit:750/0*7gaSH8-YpnOTKnpu 750w, https://miro.medium.com/v2/resize:fit:786/0*7gaSH8-YpnOTKnpu 786w, https://miro.medium.com/v2/resize:fit:828/0*7gaSH8-YpnOTKnpu 828w, https://miro.medium.com/v2/resize:fit:1100/0*7gaSH8-YpnOTKnpu 1100w, https://miro.medium.com/v2/resize:fit:1400/0*7gaSH8-YpnOTKnpu 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh hc qj c" width="700" height="402" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="ro go rp pw px rq rr bf b bg z cm"><em class="ku">Remote monitoring media being uploaded and archived using the MPS Footage Ingest workflow</em></figcaption></figure><p id="f4ad" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk">While many services in MPS are available once media has been uploaded, <em class="rt">Senna’s</em> use of MPS focused on VFX. With <em class="rt">Senna</em> shooting a high volume of footage and the show having a high volume of VFX shots, according to Post Supervisor Gabriel Queiroz <em class="rt">“Using MPS was basically a no-brainer, </em>[having]<em class="rt"> used the tool before, I knew what it could bring to the project. And to be honest, with the amount of footage that we have, it was just so much material and with the amount of vendors we have, knowing that we would have to deliver all this footage to all these kinds of vendors, including outside of Brazil and to different parts of the world.”</em></p><p id="770f" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk">With a traditional workflow, utilizing available resources in Latin America, VFX Pulls would have been done manually. This process is prone to human error and more importantly, for a show like <em class="rt">Senna</em>, too slow and would have resulted in different I/O methods for every vendor.</p><figure class="pz qa qb qc qd qe pw px paragraph-image"><div role="button" tabindex="0" class="qf qg gs qh bh qi"><div class="pw px ru"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*_CzN6mkamROqqxjo%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*_CzN6mkamROqqxjo%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*_CzN6mkamROqqxjo%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*_CzN6mkamROqqxjo%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*_CzN6mkamROqqxjo%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*_CzN6mkamROqqxjo%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*_CzN6mkamROqqxjo%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*_CzN6mkamROqqxjo 640w, https://miro.medium.com/v2/resize:fit:720/0*_CzN6mkamROqqxjo 720w, https://miro.medium.com/v2/resize:fit:750/0*_CzN6mkamROqqxjo 750w, https://miro.medium.com/v2/resize:fit:786/0*_CzN6mkamROqqxjo 786w, https://miro.medium.com/v2/resize:fit:828/0*_CzN6mkamROqqxjo 828w, https://miro.medium.com/v2/resize:fit:1100/0*_CzN6mkamROqqxjo 1100w, https://miro.medium.com/v2/resize:fit:1400/0*_CzN6mkamROqqxjo 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh hc qj c" width="700" height="399" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="ro go rp pw px rq rr bf b bg z cm"><em class="ku">Illustrating a traditional VFX Editor having to manage various I/O methods</em></figcaption></figure><p id="3595" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk">By utilizing MPS, the Assistant Editor was able to log into Content Hub, upload an EDL, and have their VFX Pulls automatically transcoded, color files consolidated and all media placed into a Google Drive style folder built directly in Content Hub (called Workspaces). The VFX Editor was able to make any additional tweaks they wanted to the directory before farming out each of the shots to whichever vendor they were meant for. When it came time for the VFX vendors to then send shots back to editorial or DI, this was also done through MPS. Having one standard method for I/O for all VFX file sharing meant that Editorial and DI did not have to manage a different file transfer/workflow for every single vendor that was onboarded.</p><figure class="pz qa qb qc qd qe pw px paragraph-image"><div role="button" tabindex="0" class="qf qg gs qh bh qi"><div class="pw px ru"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*V01AyZdu0si2N0z5%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*V01AyZdu0si2N0z5%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*V01AyZdu0si2N0z5%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*V01AyZdu0si2N0z5%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*V01AyZdu0si2N0z5%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*V01AyZdu0si2N0z5%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*V01AyZdu0si2N0z5%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*V01AyZdu0si2N0z5 640w, https://miro.medium.com/v2/resize:fit:720/0*V01AyZdu0si2N0z5 720w, https://miro.medium.com/v2/resize:fit:750/0*V01AyZdu0si2N0z5 750w, https://miro.medium.com/v2/resize:fit:786/0*V01AyZdu0si2N0z5 786w, https://miro.medium.com/v2/resize:fit:828/0*V01AyZdu0si2N0z5 828w, https://miro.medium.com/v2/resize:fit:1100/0*V01AyZdu0si2N0z5 1100w, https://miro.medium.com/v2/resize:fit:1400/0*V01AyZdu0si2N0z5 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh hc qj c" width="700" height="399" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="ro go rp pw px rq rr bf b bg z cm"><em class="ku">Illustrating a more streamlined workflow for VFX vendors when using MPS</em></figcaption></figure><p id="c2be" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk">After picture was locked and it was time for <em class="rt">Senna</em> to do their Online, the DI facility Quanta was able to utilize the Conform Pull service within MPS. The Conform Pull service allowed their team to upload an EDL, which ran a QC on all of the media from within their edit to ensure a smooth conform and then consolidated, trimmed, and packaged up all of the media they needed for the online. Since this early beta and thanks to learnings from many shows like Senna, advancements have been made in the system’s ability to match back to source media for both Conform and VFX Pulls. Rather than requiring an exact match between EDL and source OCF, there are several variations of fuzzy matching that can take place, as well as a current investigation in using one of our perceptual matching algorithms, allowing for a perceptual conform using computer vision, instead of solely relying on metadata.</p><figure class="pz qa qb qc qd qe"><div class="rl cr l gs"><figcaption class="ro go rp pw px rq rr bf b bg z cm">Inside Senna with Content Hub Media Production Suite video</figcaption></div></figure><h1 id="60c3" class="qk ql jt bf qm qn qo qp hp qq qr qs hr qt qu qv qw qx qy qz ra rb rc rd re rf bk">Conclusion</h1><p id="a3cb" class="pw-post-body-paragraph pb pc jt pd b pe rg pg ph pi rh pk pl hs ri pn po hv rj pq pr hy rk pt pu pv iu bk">The Media Production Suite (MPS) represents a transformative leap in how we approach media production at Netflix. By embracing open standards, we have crafted a scalable solution that not only makes economic sense but also democratizes access to advanced production tools across the globe. This approach allows us to eliminate tedious tasks, enabling our teams to focus on what truly matters: creative storytelling. By fostering global collaboration and leveraging the power of cloud-based workflows, we’re not just enhancing efficiency but also elevating the quality of our productions. As we continue to innovate and refine our processes, we remain committed to breaking down barriers and unlocking the full potential of creative talent worldwide. The future of filmmaking is here, and with MPS, we are leading the charge toward a more connected and creatively empowered industry.</p></div>]]></description>
      <link>https://netflixtechblog.com/globalizing-productions-with-netflixs-media-production-suite-fc3c108c0a22</link>
      <guid>https://netflixtechblog.com/globalizing-productions-with-netflixs-media-production-suite-fc3c108c0a22</guid>
      <pubDate>Mon, 31 Mar 2025 18:21:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Foundation Model for Personalized Recommendation]]></title>
      <description><![CDATA[<div><div></div><p id="af55" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk">By <a class="ag ih" href="https://www.linkedin.com/in/markhsiao/" rel="noopener ugc nofollow" target="_blank">Ko-Jen Hsiao</a>, <a class="ag ih" href="https://www.linkedin.com/in/yesufeng/" rel="noopener ugc nofollow" target="_blank">Yesu Feng</a> and <a class="ag ih" href="https://www.linkedin.com/in/sudarshanlamkhede/" rel="noopener ugc nofollow" target="_blank">Sudarshan Lamkhede</a></p><h1 id="400b" class="pw px jt bf py pz qa qb hp qc qd qe hr qf qg qh qi qj qk ql qm qn qo qp qq qr bk">Motivation</h1><p id="4d49" class="pw-post-body-paragraph pb pc jt pd b pe qs pg ph pi qt pk pl hs qu pn po hv qv pq pr hy qw pt pu pv iu bk">Netflix’s personalized recommender system is a complex system, boasting a variety of specialized machine learned models each catering to distinct needs including “Continue Watching” and “Today’s Top Picks for You.” (Refer to our recent <a class="ag ih" href="https://videorecsys.com/slides/mark_talk3.pdf" rel="noopener ugc nofollow" target="_blank">overview</a> for more details). However, as we expanded our set of personalization algorithms to meet increasing business needs, maintenance of the recommender system became quite costly. Furthermore, it was difficult to transfer innovations from one model to another, given that most are independently trained despite using common data sources. This scenario underscored the need for a new recommender system architecture where member preference learning is centralized, enhancing accessibility and utility across different models.</p><p id="b0d8" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk">Particularly, these models predominantly extract features from members’ recent interaction histories on the platform. Yet, many are confined to a brief temporal window due to constraints in serving latency or training costs. This limitation has inspired us to develop a foundation model for recommendation. This model aims to assimilate information both from members’ comprehensive interaction histories and our content at a very large scale. It facilitates the distribution of these learnings to other models, either through shared model weights for fine tuning or directly through embeddings.</p><p id="34e6" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk">The impetus for constructing a foundational recommendation model is based on the paradigm shift in natural language processing (NLP) to large language models (LLMs). In NLP, the trend is moving away from numerous small, specialized models towards a single, large language model that can perform a variety of tasks either directly or with minimal fine-tuning. Key insights from this shift include:</p><ol class=""><li id="c797" class="pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv qx qy qz bk"><strong class="pd ju">A Data-Centric Approach</strong>: Shifting focus from model-centric strategies, which heavily rely on feature engineering, to a data-centric one. This approach prioritizes the accumulation of large-scale, high-quality data and, where feasible, aims for end-to-end learning.</li><li id="00ed" class="pb pc jt pd b pe ra pg ph pi rb pk pl hs rc pn po hv rd pq pr hy re pt pu pv qx qy qz bk"><strong class="pd ju">Leveraging Semi-Supervised Learning</strong>: The next-token prediction objective in LLMs has proven remarkably effective. It enables large-scale semi-supervised learning using unlabeled data while also equipping the model with a surprisingly deep understanding of world knowledge.</li></ol><p id="215e" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk">These insights have shaped the design of our foundation model, enabling a transition from maintaining numerous small, specialized models to building a scalable, efficient system. By scaling up semi-supervised training data and model parameters, we aim to develop a model that not only meets current needs but also adapts dynamically to evolving demands, ensuring sustainable innovation and resource efficiency.</p><h1 id="cbeb" class="pw px jt bf py pz qa qb hp qc qd qe hr qf qg qh qi qj qk ql qm qn qo qp qq qr bk">Data</h1><p id="9678" class="pw-post-body-paragraph pb pc jt pd b pe qs pg ph pi qt pk pl hs qu pn po hv qv pq pr hy qw pt pu pv iu bk">At Netflix, user engagement spans a wide spectrum, from casual browsing to committed movie watching. With over 300 million users at the end of 2024, this translates into hundreds of billions of interactions — an immense dataset comparable in scale to the token volume of large language models (LLMs). However, as in LLMs, the quality of data often outweighs its sheer volume. To harness this data effectively, we employ a process of interaction tokenization, ensuring meaningful events are identified and redundancies are minimized.</p><p id="6600" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk"><strong class="pd ju">Tokenizing User Interactions</strong>: Not all raw user actions contribute equally to understanding preferences. Tokenization helps define what constitutes a meaningful “token” in a sequence. Drawing an analogy to Byte Pair Encoding (BPE) in NLP, we can think of tokenization as merging adjacent actions to form new, higher-level tokens. However, unlike language tokenization, creating these new tokens requires careful consideration of what information to retain. For instance, the total watch duration might need to be summed or engagement types aggregated to preserve critical details.</p><figure class="ri rj rk rl rm rn rf rg paragraph-image"><div role="button" tabindex="0" class="ro rp gs rq bh rr"><div class="rf rg rh"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*1dhdoLxKnf_fcZOq%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*1dhdoLxKnf_fcZOq%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*1dhdoLxKnf_fcZOq%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*1dhdoLxKnf_fcZOq%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*1dhdoLxKnf_fcZOq%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*1dhdoLxKnf_fcZOq%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*1dhdoLxKnf_fcZOq%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*1dhdoLxKnf_fcZOq 640w, https://miro.medium.com/v2/resize:fit:720/0*1dhdoLxKnf_fcZOq 720w, https://miro.medium.com/v2/resize:fit:750/0*1dhdoLxKnf_fcZOq 750w, https://miro.medium.com/v2/resize:fit:786/0*1dhdoLxKnf_fcZOq 786w, https://miro.medium.com/v2/resize:fit:828/0*1dhdoLxKnf_fcZOq 828w, https://miro.medium.com/v2/resize:fit:1100/0*1dhdoLxKnf_fcZOq 1100w, https://miro.medium.com/v2/resize:fit:1400/0*1dhdoLxKnf_fcZOq 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh hc rs c" width="700" height="281" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="rt go ru rf rg rv rw bf b bg z cm"><strong class="bf py">Figure 1.</strong>Tokenization of user interaction history by merging actions on the same title, preserving important information.</figcaption></figure><p id="eb1f" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk">This tradeoff between granular data and sequence compression is akin to the balance in LLMs between vocabulary size and context window. In our case, the goal is to balance the length of interaction history against the level of detail retained in individual tokens. Overly lossy tokenization risks losing valuable signals, while too granular a sequence can exceed practical limits on processing time and memory.</p><p id="55a5" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk">Even with such strategies, interaction histories from active users can span thousands of events, exceeding the capacity of transformer models with standard self attention layers. In recommendation systems, context windows during inference are often limited to hundreds of events — not due to model capability but because these services typically require millisecond-level latency. This constraint is more stringent than what is typical in LLM applications, where longer inference times (seconds) are more tolerable.</p><p id="c640" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk">To address this during training, we implement two key solutions:</p><ol class=""><li id="0868" class="pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv qx qy qz bk"><strong class="pd ju">Sparse Attention Mechanisms</strong>: By leveraging sparse attention techniques such as low-rank compression, the model can extend its context window to several hundred events while maintaining computational efficiency. This enables it to process more extensive interaction histories and derive richer insights into long-term preferences.</li><li id="75fa" class="pb pc jt pd b pe ra pg ph pi rb pk pl hs rc pn po hv rd pq pr hy re pt pu pv qx qy qz bk"><a class="ag ih" href="https://arxiv.org/abs/2409.14517" rel="noopener ugc nofollow" target="_blank"><strong class="pd ju">Sliding Window Sampling</strong></a>: During training, we sample overlapping windows of interactions from the full sequence. This ensures the model is exposed to different segments of the user’s history over multiple epochs, allowing it to learn from the entire sequence without requiring an impractically large context window.</li></ol><p id="00ba" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk">At inference time, when multi-step decoding is needed, we can deploy KV caching to efficiently reuse past computations and maintain low latency.</p><p id="e2ee" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk">These approaches collectively allow us to balance the need for detailed, long-term interaction modeling with the practical constraints of model training and inference, enhancing both the precision and scalability of our recommendation system.</p><p id="c410" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk"><strong class="pd ju">Information in Each ‘Token’</strong>: While the first part of our tokenization process focuses on structuring sequences of interactions, the next critical step is defining the rich information contained within each token. Unlike LLMs, which typically rely on a single embedding space to represent input tokens, our interaction events are packed with heterogeneous details. These include attributes of the action itself (such as locale, time, duration, and device type) as well as information about the content (such as item ID and metadata like genre and release country). Most of these features, especially categorical ones, are directly embedded within the model, embracing an end-to-end learning approach. However, certain features require special attention. For example, timestamps need additional processing to capture both absolute and relative notions of time, with absolute time being particularly important for understanding time-sensitive behaviors.</p><p id="5ece" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk">To enhance prediction accuracy in sequential recommendation systems, we organize token features into two categories:</p><ol class=""><li id="a88b" class="pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv qx qy qz bk"><strong class="pd ju">Request-Time Features</strong>: These are features available at the moment of prediction, such as log-in time, device, or location.</li><li id="5441" class="pb pc jt pd b pe ra pg ph pi rb pk pl hs rc pn po hv rd pq pr hy re pt pu pv qx qy qz bk"><strong class="pd ju">Post-Action Features</strong>: These are details available after an interaction has occurred, such as the specific show interacted with or the duration of the interaction.</li></ol><p id="5152" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk">To predict the next interaction, we combine request-time features from the current step with post-action features from the <a class="ag ih" href="https://ojs.aaai.org/aimagazine/index.php/aimagazine/article/view/18140" rel="noopener ugc nofollow" target="_blank">previous step</a>. This blending of contextual and historical information ensures each token in the sequence carries a comprehensive representation, capturing both the immediate context and user behavior patterns over time.</p><h1 id="0db0" class="pw px jt bf py pz qa qb hp qc qd qe hr qf qg qh qi qj qk ql qm qn qo qp qq qr bk">Considerations for Model Objective and Architecture</h1><p id="6f3d" class="pw-post-body-paragraph pb pc jt pd b pe qs pg ph pi qt pk pl hs qu pn po hv qv pq pr hy qw pt pu pv iu bk">As previously mentioned, our default approach employs the autoregressive next-token prediction objective, similar to GPT. This strategy effectively leverages the vast scale of unlabeled user interaction data. The adoption of this objective in recommendation systems has shown multiple successes [1–3]. However, given the distinct differences between language tasks and recommendation tasks, we have made several critical modifications to the objective.</p><p id="91bb" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk">Firstly, during the pretraining phase of typical LLMs, such as GPT, every target token is generally treated with equal weight. In contrast, in our model, not all user interactions are of equal importance. For instance, a 5-minute trailer play should not carry the same weight as a 2-hour full movie watch. A greater challenge arises when trying to align long-term user satisfaction with specific interactions and recommendations. To address this, we can adopt a multi-token prediction objective during training, where the model predicts the next <em class="rx">n</em> tokens at each step instead of a single token[4]. This approach encourages the model to capture longer-term dependencies and avoid myopic predictions focused solely on immediate next events.</p><p id="a9d5" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk">Secondly, we can use multiple fields in our input data as auxiliary prediction objectives in addition to predicting the next item ID, which remains the primary target. For example, we can derive genres from the items in the original sequence and use this genre sequence as an auxiliary target. This approach serves several purposes: it acts as a regularizer to reduce overfitting on noisy item ID predictions, provides additional insights into user intentions or long-term genre preferences, and, when structured hierarchically, can improve the accuracy of predicting the target item ID. By first predicting auxiliary targets, such as genre or original language, the model effectively narrows down the candidate list, simplifying subsequent item ID prediction.</p><h1 id="ab80" class="pw px jt bf py pz qa qb hp qc qd qe hr qf qg qh qi qj qk ql qm qn qo qp qq qr bk">Unique Challenges for Recommendation FM</h1><p id="0a8b" class="pw-post-body-paragraph pb pc jt pd b pe qs pg ph pi qt pk pl hs qu pn po hv qv pq pr hy qw pt pu pv iu bk">In addition to the infrastructure challenges posed by training bigger models with substantial amounts of user interaction data that are common when trying to build foundation models, there are several unique hurdles specific to recommendations to make them viable. One of unique challenges is entity cold-starting.</p><p id="a38b" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk">At Netflix, our mission is to entertain the world. New titles are added to the catalog frequently. Therefore the recommendation foundation models require a cold start capability, which means the models need to estimate members’ preferences for newly launched titles before anyone has engaged with them. To enable this, our foundation model training framework is built with the following two capabilities: Incremental training and being able to do inference with unseen entities.</p><ol class=""><li id="d462" class="pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv qx qy qz bk"><strong class="pd ju">Incremental training </strong>: Foundation models are trained on extensive datasets, including every member’s history of plays and actions, making frequent retraining impractical. However, our catalog and member preferences continually evolve. Unlike large language models, which can be incrementally trained with stable token vocabularies, our recommendation models require new embeddings for new titles, necessitating expanded embedding layers and output components. To address this, we warm-start new models by reusing parameters from previous models and initializing new parameters for new titles. For example, new title embeddings can be initialized by adding slight random noise to existing average embeddings or by using a weighted combination of similar titles’ embeddings based on metadata. This approach allows new titles to start with relevant embeddings, facilitating faster fine-tuning. In practice, the initialization method becomes less critical when more member interaction data is used for fine-tuning.</li><li id="09a9" class="pb pc jt pd b pe ra pg ph pi rb pk pl hs rc pn po hv rd pq pr hy re pt pu pv qx qy qz bk"><strong class="pd ju">Dealing with unseen entities </strong>: Even with incremental training, it’s not always guaranteed to learn efficiently on new entities (ex: newly launched titles). It’s also possible that there will be some new entities that are not included/seen in the training data even if we fine-tune foundation models on a frequent basis. Therefore, it’s also important to let foundation models use metadata information of entities and inputs, not just member interaction data. Thus, our foundation model combines both learnable item id embeddings and learnable embeddings from metadata. The following diagram demonstrates this idea.</li></ol><figure class="ri rj rk rl rm rn rf rg paragraph-image"><div role="button" tabindex="0" class="ro rp gs rq bh rr"><div class="rf rg ry"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*7qnfUGWgXtVUjhP9%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*7qnfUGWgXtVUjhP9%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*7qnfUGWgXtVUjhP9%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*7qnfUGWgXtVUjhP9%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*7qnfUGWgXtVUjhP9%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*7qnfUGWgXtVUjhP9%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*7qnfUGWgXtVUjhP9%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*7qnfUGWgXtVUjhP9 640w, https://miro.medium.com/v2/resize:fit:720/0*7qnfUGWgXtVUjhP9 720w, https://miro.medium.com/v2/resize:fit:750/0*7qnfUGWgXtVUjhP9 750w, https://miro.medium.com/v2/resize:fit:786/0*7qnfUGWgXtVUjhP9 786w, https://miro.medium.com/v2/resize:fit:828/0*7qnfUGWgXtVUjhP9 828w, https://miro.medium.com/v2/resize:fit:1100/0*7qnfUGWgXtVUjhP9 1100w, https://miro.medium.com/v2/resize:fit:1400/0*7qnfUGWgXtVUjhP9 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh hc rs c" width="700" height="389" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="rt go ru rf rg rv rw bf b bg z cm"><strong class="bf py">Figure 2. </strong>Titles are associated with various metadata, such as genres, storylines, and tones. Each type of metadata could be represented by averaging its respective embeddings, which are then concatenated to form the overall metadata-based embedding for the title.</figcaption></figure><p id="c1b6" class="pw-post-body-paragraph pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv iu bk">To create the final title embedding, we combine this metadata-based embedding with a fully-learnable ID-based embedding using a mixing layer. Instead of simply summing these embeddings, we use an attention mechanism based on the “age” of the entity. This approach allows new titles with limited interaction data to rely more on metadata, while established titles can depend more on ID-based embeddings. Since titles with similar metadata can have different user engagement, their embeddings should reflect these differences. Introducing some randomness during training encourages the model to learn from metadata rather than relying solely on ID embeddings. This method ensures that newly-launched or pre-launch titles have reasonable embeddings even with no user interaction data.</p><h1 id="5fdf" class="pw px jt bf py pz qa qb hp qc qd qe hr qf qg qh qi qj qk ql qm qn qo qp qq qr bk">Downstream Applications and Challenges</h1><p id="06ff" class="pw-post-body-paragraph pb pc jt pd b pe qs pg ph pi qt pk pl hs qu pn po hv qv pq pr hy qw pt pu pv iu bk">Our recommendation foundation model is designed to understand long-term member preferences and can be utilized in various ways by downstream applications:</p><ol class=""><li id="eb13" class="pb pc jt pd b pe pf pg ph pi pj pk pl hs pm pn po hv pp pq pr hy ps pt pu pv qx qy qz bk"><strong class="pd ju">Direct Use as a Predictive Model </strong>The model is primarily trained to predict the next entity a user will interact with. It includes multiple predictor heads for different tasks, such as forecasting member preferences for various genres. These can be directly applied to meet diverse business needs..</li><li id="b2fa" class="pb pc jt pd b pe ra pg ph pi rb pk pl hs rc pn po hv rd pq pr hy re pt pu pv qx qy qz bk"><strong class="pd ju">Utilizing embeddings </strong>The model generates valuable embeddings for members and entities like videos, games, and genres. These embeddings are calculated in batch jobs and stored for use in both offline and online applications. They can serve as features in other models or be used for candidate generation, such as retrieving appealing titles for a user. High-quality title embeddings also support title-to-title recommendations. However, one important consideration is that the embedding space has arbitrary, uninterpretable dimensions and is incompatible across different model training runs. This poses challenges for downstream consumers, who must adapt to each retraining and redeployment, risking bugs due to invalidated assumptions about the embedding structure. To address this, we apply an orthogonal low-rank transformation to stabilize the user/item embedding space, ensuring consistent meaning of embedding dimensions, even as the base foundation model is retrained and redeployed.</li><li id="fb28" class="pb pc jt pd b pe ra pg ph pi rb pk pl hs rc pn po hv rd pq pr hy re pt pu pv qx qy qz bk"><strong class="pd ju">Fine-Tuning with Specific Data </strong>The model’s adaptability allows for fine-tuning with application-specific data. Users can integrate the full model or subgraphs into their own models, fine-tuning them with less data and computational power. This approach achieves performance comparable to previous models, despite the initial foundation model requiring significant resources.</li></ol><h1 id="0565" class="pw px jt bf py pz qa qb hp qc qd qe hr qf qg qh qi qj qk ql qm qn qo qp qq qr bk">Scaling Foundation Models for Netflix Recommendations</h1><p id="e598" class="pw-post-body-paragraph pb pc jt pd b pe qs pg ph pi qt pk pl hs qu pn po hv qv pq pr hy qw pt pu pv iu bk">In scaling up our foundation model for Netflix recommendations, we draw inspiration from the success of large language models (LLMs). Just as LLMs have demonstrated the power of scaling in improving performance, we find that scaling is crucial for enhancing generative recommendation tasks. Successful scaling demands robust evaluation, efficient training algorithms, and substantial computing resources. Evaluation must effectively differentiate model performance and identify areas for improvement. Scaling involves data, model, and context scaling, incorporating user engagement, external reviews, multimedia assets, and high-quality embeddings. Our experiments confirm that the scaling law also applies to our foundation model, with consistent improvements observed as we increase data and model size.</p><figure class="ri rj rk rl rm rn rf rg paragraph-image"><div role="button" tabindex="0" class="ro rp gs rq bh rr"><div class="rf rg rz"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*dEypYqp643q6GcVzn3IIww.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*dEypYqp643q6GcVzn3IIww.png" /><img alt="" class="bh hc rs c" width="700" height="456" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="rt go ru rf rg rv rw bf b bg z cm"><strong class="bf py">Figure 3. </strong>The relationship between model parameter size and relative performance improvement. The plot demonstrates the scaling law in recommendation modeling, showing a trend of increased performance with larger model sizes. The x-axis is logarithmically scaled to highlight growth across different magnitudes.</figcaption></figure><h1 id="0f22" class="pw px jt bf py pz qa qb hp qc qd qe hr qf qg qh qi qj qk ql qm qn qo qp qq qr bk">Conclusion</h1><p id="baf8" class="pw-post-body-paragraph pb pc jt pd b pe qs pg ph pi qt pk pl hs qu pn po hv qv pq pr hy qw pt pu pv iu bk">In conclusion, our Foundation Model for Personalized Recommendation represents a significant step towards creating a unified, data-centric system that leverages large-scale data to increase the quality of recommendations for our members. This approach borrows insights from Large Language Models (LLMs), particularly the principles of semi-supervised learning and end-to-end training, aiming to harness the vast scale of unlabeled user interaction data. Addressing unique challenges, like cold start and presentation bias, the model also acknowledges the distinct differences between language tasks and recommendation. The Foundation Model allows various downstream applications, from direct use as a predictive model to generate user and entity embeddings for other applications, and can be fine-tuned for specific canvases. We see promising results from downstream integrations. This move from multiple specialized models to a more comprehensive system marks an exciting development in the field of personalized recommendation systems.</p><h1 id="51d7" class="pw px jt bf py pz qa qb hp qc qd qe hr qf qg qh qi qj qk ql qm qn qo qp qq qr bk">Acknowledgements</h1><p id="8a46" class="pw-post-body-paragraph pb pc jt pd b pe qs pg ph pi qt pk pl hs qu pn po hv qv pq pr hy qw pt pu pv iu bk">Contributors to this work (name in alphabetical order): <a class="ag ih" href="https://www.linkedin.com/in/aileisun/" rel="noopener ugc nofollow" target="_blank">Ai-Lei Sun</a> <a class="ag ih" href="https://www.linkedin.com/in/aishafenton/" rel="noopener ugc nofollow" target="_blank">Aish Fenton</a> <a class="ag ih" href="https://www.linkedin.com/in/annecocos/" rel="noopener ugc nofollow" target="_blank">Anne Cocos</a> <a class="ag ih" href="https://www.linkedin.com/in/foranuj/" rel="noopener ugc nofollow" target="_blank">Anuj Shah</a> <a class="ag ih" href="https://www.linkedin.com/in/arashaghevli/" rel="noopener ugc nofollow" target="_blank">Arash Aghevli</a> <a class="ag ih" href="https://www.linkedin.com/in/baolin-li-659426115/" rel="noopener ugc nofollow" target="_blank">Baolin Li</a> <a class="ag ih" href="https://www.linkedin.com/in/bowei-yan-0080a326/" rel="noopener ugc nofollow" target="_blank">Bowei Yan</a> <a class="ag ih" href="https://www.linkedin.com/in/danielzheng256/" rel="noopener ugc nofollow" target="_blank">Dan Zheng</a> <a class="ag ih" href="https://www.linkedin.com/in/dwliang/" rel="noopener ugc nofollow" target="_blank">Dawen Liang</a> <a class="ag ih" href="https://www.linkedin.com/in/ding-tong-2812785a/" rel="noopener ugc nofollow" target="_blank">Ding Tong</a> <a class="ag ih" href="https://www.linkedin.com/in/divya-gadde-3ba01551/" rel="noopener ugc nofollow" target="_blank">Divya Gadde</a> <a class="ag ih" href="https://www.linkedin.com/in/emma-yanyang-kong-6904b457/" rel="noopener ugc nofollow" target="_blank">Emma Kong</a> <a class="ag ih" href="https://www.linkedin.com/in/gary-y-62175170/" rel="noopener ugc nofollow" target="_blank">Gary Yeh</a> <a class="ag ih" href="https://www.linkedin.com/in/inbar-naor-6b973a50/" rel="noopener ugc nofollow" target="_blank">Inbar Naor</a> <a class="ag ih" href="https://www.linkedin.com/in/jinwangw/" rel="noopener ugc nofollow" target="_blank">Jin Wang</a> <a class="ag ih" href="https://www.linkedin.com/in/jbasilico/" rel="noopener ugc nofollow" target="_blank">Justin Basilico</a> <a class="ag ih" href="https://www.linkedin.com/in/kabir-nagrecha/overlay/about-this-profile/" rel="noopener ugc nofollow" target="_blank">Kabir Nagrecha</a> <a class="ag ih" href="https://www.linkedin.com/in/kzielnicki/" rel="noopener ugc nofollow" target="_blank">Kevin Zielnicki</a> <a class="ag ih" href="https://www.linkedin.com/in/linasbaltrunas/" rel="noopener ugc nofollow" target="_blank">Linas Baltrunas</a> <a class="ag ih" href="https://www.linkedin.com/in/lingyi-liu-4b866016/" rel="noopener ugc nofollow" target="_blank">Lingyi Liu</a> <a class="ag ih" href="https://www.linkedin.com/in/lequn-luke-wang-9226b2129/" rel="noopener ugc nofollow" target="_blank">Luke Wang</a> <a class="ag ih" href="https://www.linkedin.com/in/matan-appelbaum-39472b96/" rel="noopener ugc nofollow" target="_blank">Matan Appelbaum</a> <a class="ag ih" href="https://www.linkedin.com/in/tuzhucheng/" rel="noopener ugc nofollow" target="_blank">Michael Tu</a> <a class="ag ih" href="https://www.linkedin.com/in/moumitab/" rel="noopener ugc nofollow" target="_blank">Moumita Bhattacharya</a> <a class="ag ih" href="https://www.linkedin.com/in/pabloadelgado/" rel="noopener ugc nofollow" target="_blank">Pablo Delgado</a> <a class="ag ih" href="https://www.linkedin.com/in/qiuling-xu-a445b815a/" rel="noopener ugc nofollow" target="_blank">Qiuling Xu</a> <a class="ag ih" href="https://www.linkedin.com/in/rakeshkomuravelli/" rel="noopener ugc nofollow" target="_blank">Rakesh Komuravelli</a> <a class="ag ih" href="https://www.linkedin.com/in/raveeshbhalla/" rel="noopener ugc nofollow" target="_blank">Raveesh Bhalla</a> <a class="ag ih" href="https://www.linkedin.com/in/rob-story-b21a4912/" rel="noopener ugc nofollow" target="_blank">Rob Story</a> <a class="ag ih" href="https://www.linkedin.com/in/rogermenezes/" rel="noopener ugc nofollow" target="_blank">Roger Menezes</a> <a class="ag ih" href="https://www.linkedin.com/in/sejoon-oh/" rel="noopener ugc nofollow" target="_blank">Sejoon Oh</a> <a class="ag ih" href="https://www.linkedin.com/in/shahrzad-naseri-1b988760/" rel="noopener ugc nofollow" target="_blank">Shahrzad Naseri</a> <a class="ag ih" href="https://www.linkedin.com/in/swanandjoshi7/" rel="noopener ugc nofollow" target="_blank">Swanand Joshi</a> <a class="ag ih" href="https://www.linkedin.com/in/trungnguyen324/" rel="noopener ugc nofollow" target="_blank">Trung Nguyen</a> <a class="ag ih" href="https://www.linkedin.com/in/vito-ostuni-0b576027/" rel="noopener ugc nofollow" target="_blank">Vito Ostuni </a><a class="ag ih" href="https://www.linkedin.com/in/thomasweiwang/" rel="noopener ugc nofollow" target="_blank">Wei Wang</a> <a class="ag ih" href="https://www.linkedin.com/in/zhezhangncsu/" rel="noopener ugc nofollow" target="_blank">Zhe Zhang</a></p><h1 id="9c30" class="pw px jt bf py pz qa qb hp qc qd qe hr qf qg qh qi qj qk ql qm qn qo qp qq qr bk">Reference</h1><ol class=""><li id="d518" class="pb pc jt pd b pe qs pg ph pi qt pk pl hs qu pn po hv qv pq pr hy qw pt pu pv qx qy qz bk">C. K. Kang and J. McAuley, “Self-Attentive Sequential Recommendation,” <em class="rx">2018 IEEE International Conference on Data Mining (ICDM)</em>, Singapore, 2018, pp. 197–206, doi: 10.1109/ICDM.2018.00035.</li><li id="3a76" class="pb pc jt pd b pe ra pg ph pi rb pk pl hs rc pn po hv rd pq pr hy re pt pu pv qx qy qz bk">F. Sun et al., “BERT4Rec: Sequential Recommendation with Bidirectional Encoder Representations from Transformer,” <em class="rx">Proceedings of the 28th ACM International Conference on Information and Knowledge Management (CIKM ‘19)</em>, Beijing, China, 2019, pp. 1441–1450, doi: 10.1145/3357384.3357895.</li><li id="6b8c" class="pb pc jt pd b pe ra pg ph pi rb pk pl hs rc pn po hv rd pq pr hy re pt pu pv qx qy qz bk">J. Zhai et al., “Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations,” <em class="rx">arXiv preprint arXiv:2402.17152</em>, 2024.</li><li id="9071" class="pb pc jt pd b pe ra pg ph pi rb pk pl hs rc pn po hv rd pq pr hy re pt pu pv qx qy qz bk">F. Gloeckle, B. Youbi Idrissi, B. Rozière, D. Lopez-Paz, and G. Synnaeve, “Better &amp; Faster Large Language Models via Multi-token Prediction,” arXiv preprint arXiv:2404.19737, Apr. 2024.</li></ol></div>]]></description>
      <link>https://netflixtechblog.com/foundation-model-for-personalized-recommendation-1a0bd8e02d39</link>
      <guid>https://netflixtechblog.com/foundation-model-for-personalized-recommendation-1a0bd8e02d39</guid>
      <pubDate>Sat, 29 Mar 2025 01:51:00 +0100</pubDate>
    </item>
    <item>
      <title><![CDATA[HDR10+ Now Streaming on Netflix]]></title>
      <description><![CDATA[<div><div></div><p id="bf52" class="pw-post-body-paragraph or os jk ot b ou ov ow ox oy oz pa pb hj pc pd pe hm pf pg ph hp pi pj pk pl il bk"><a class="ag hy" href="https://www.linkedin.com/in/rquero" rel="noopener ugc nofollow" target="_blank">Roger Quero</a>, <a class="ag hy" href="https://www.linkedin.com/in/liwei-guo" rel="noopener ugc nofollow" target="_blank">Liwei Guo</a>, <a class="ag hy" href="https://www.linkedin.com/in/jeffrwatts/" rel="noopener ugc nofollow" target="_blank">Jeff Watts</a>, <a class="ag hy" href="https://www.linkedin.com/in/joseph-mccormick-7b386026" rel="noopener ugc nofollow" target="_blank">Joseph McCormick</a>, <a class="ag hy" href="https://www.linkedin.com/in/agataopalach/" rel="noopener ugc nofollow" target="_blank">Agata Opalach</a>, <a class="ag hy" href="https://www.linkedin.com/in/anush-moorthy-b8451142/" rel="noopener ugc nofollow" target="_blank">Anush Moorthy</a></p><p id="f88c" class="pw-post-body-paragraph or os jk ot b ou ov ow ox oy oz pa pb hj pc pd pe hm pf pg ph hp pi pj pk pl il bk">We are excited to announce that we are now streaming HDR10+ content on our service for AV1-enabled devices, enhancing the viewing experience for certified HDR10+ devices, which previously only received HDR10 content. The dynamic metadata included in our HDR10+ content improves the quality and accuracy of the picture when viewed on these devices.</p><h1 id="4a1e" class="pm pn jk bf po pp pq pr hg ps pt pu hi pv pw px py pz qa qb qc qd qe qf qg qh bk">Delighting Members with Even Better Picture Quality</h1><p id="ba30" class="pw-post-body-paragraph or os jk ot b ou qi ow ox oy qj pa pb hj qk pd pe hm ql pg ph hp qm pj pk pl il bk">Nearly a decade ago, we made a bold move to be a pioneering adopter of High Dynamic Range (HDR) technology. HDR enables images to have more details, vivid colors, and improved realism. We began producing our shows and movies in HDR, encoding them in HDR, and streaming them in HDR for our members. We were confident that it would greatly enhance our members’ viewing experience, and unlock new creative visions — and we were right! In the last five years, HDR streaming has increased by more than 300%, while the number of HDR-configured devices watching Netflix has more than doubled. Since launching HDR with season one of <em class="qn">Marco Polo</em>, Netflix now has over 11,000 hours of HDR titles for members to immerse themselves in.</p><p id="c033" class="pw-post-body-paragraph or os jk ot b ou ov ow ox oy oz pa pb hj pc pd pe hm pf pg ph hp pi pj pk pl il bk">We continue to enhance member joy while maintaining creative vision by adding support for HDR10+. This will further augment Netflix’s growing HDR ecosystem, preserve creative intent on even more devices, and provide a more immersive viewing experience.</p><p id="1789" class="pw-post-body-paragraph or os jk ot b ou ov ow ox oy oz pa pb hj pc pd pe hm pf pg ph hp pi pj pk pl il bk">We enabled HDR10+ on Netflix using the <a class="ag hy" href="https://aomedia.org/specifications/av1/" rel="noopener ugc nofollow" target="_blank">AV1 video codec</a> that was standardized by the Alliance for Open Media (AOM) in 2018. AV1 is one of the most efficient codecs available today. We <a class="ag hy" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/bringing-av1-streaming-to-netflix-members-tvs-b7fc88e42320">previously enabled</a> AV1 encoding for SDR content, and saw tremendous value for our members, including higher and more consistent visual quality, lower play delay and increased streaming at the highest resolution. AV1-SDR is already the second most streamed codec at Netflix, behind H.264/AVC, which has been around for over 20 years! With the addition of HDR10+ streams to AV1, we expect the day is not far when AV1 will be the most streamed codec at Netflix.</p><p id="849f" class="pw-post-body-paragraph or os jk ot b ou ov ow ox oy oz pa pb hj pc pd pe hm pf pg ph hp pi pj pk pl il bk">To enhance our offering, we have been adding HDR10+ streams to both new releases and existing popular HDR titles. AV1-HDR10+ now accounts for 50% of all eligible viewing hours. We will continue expanding our HDR10+ offerings with the goal of providing an HDR10+ experience for all HDR titles by the end of this year¹.</p><h1 id="d8ad" class="pm pn jk bf po pp pq pr hg ps pt pu hi pv pw px py pz qa qb qc qd qe qf qg qh bk"><strong class="am">Industry Adopted Formats</strong></h1><p id="947b" class="pw-post-body-paragraph or os jk ot b ou qi ow ox oy qj pa pb hj qk pd pe hm ql pg ph hp qm pj pk pl il bk">Today, the industry recognizes three prevalent HDR formats: Dolby Vision, HDR10, and HDR10+. For all three HDR Formats, metadata is embedded in the content, serving as instructions to guide the playback device — whether it’s a TV, mobile device, or computer — on how to display the image.</p><p id="c1f5" class="pw-post-body-paragraph or os jk ot b ou ov ow ox oy oz pa pb hj pc pd pe hm pf pg ph hp pi pj pk pl il bk">HDR10 is the most widely adopted HDR format, supported by all HDR devices. HDR10 uses static metadata that is defined once for the entire content detailing aspects such as the maximum content light level (MaxCLL), maximum frame average light level (MaxFALL), as well as characteristics of the mastering display used for color grading. This metadata only allows for a one-size-fits-all tone mapping of the content for display devices. It cannot account for dynamic contrast across scenes, which most content contains.</p><p id="2353" class="pw-post-body-paragraph or os jk ot b ou ov ow ox oy oz pa pb hj pc pd pe hm pf pg ph hp pi pj pk pl il bk">HDR10+ and Dolby Vision improve on this with dynamic metadata that provides content image statistics on a per-frame basis, enabling optimized tone mapping adjustments for each scene. This achieves greater perceptual fidelity to the original, preserving creative intent.</p><h1 id="5572" class="pm pn jk bf po pp pq pr hg ps pt pu hi pv pw px py pz qa qb qc qd qe qf qg qh bk"><strong class="am">HDR10 vs. HDR10+</strong></h1><p id="4df0" class="pw-post-body-paragraph or os jk ot b ou qi ow ox oy qj pa pb hj qk pd pe hm ql pg ph hp qm pj pk pl il bk">The figure below shows screen grabs of two AV1-encoded frames of the same content displayed using HDR10 (top) and HDR10+ (bottom).</p><figure class="qr qs qt qu qv qw qo qp paragraph-image"><div role="button" tabindex="0" class="qx qy gj qz bh ra"><div class="qo qp qq"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*AjnQRaY7VFZoonX5SI36IA.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*AjnQRaY7VFZoonX5SI36IA.png" /><img alt="" class="bh gt rb c" width="700" height="418" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><figure class="qr qs qt qu qv qw qo qp paragraph-image"><div role="button" tabindex="0" class="qx qy gj qz bh ra"><div class="qo qp rc"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*gsW42hweG6RMbWwQjy1etg.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*gsW42hweG6RMbWwQjy1etg.png" /><img alt="" class="bh gt rb c" width="700" height="420" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="0ce5" class="pw-post-body-paragraph or os jk ot b ou ov ow ox oy oz pa pb hj pc pd pe hm pf pg ph hp pi pj pk pl il bk"><em class="qn">Photographs of devices displaying the same frame with HDR10 metadata (top) and HDR10+ metadata (bottom). Notice the preservation of the flashlight detail in the HDR10+ capture, and the over-exposure of the region under the flashlight in the HDR10 one².</em></p><p id="5357" class="pw-post-body-paragraph or os jk ot b ou ov ow ox oy oz pa pb hj pc pd pe hm pf pg ph hp pi pj pk pl il bk">As seen in the flashlight on the table, the highlight details are clipped in the HDR10 content, but are recovered in HDR10+. Further, the region under the flashlight is overexposed in the HDR10 content, while HDR10+ renders that region with greater fidelity to the source. The reason HDR10+, with its dynamic metadata, shines in this example is that the scenes preceding and following the scene with this frame have markedly different luminance statistics. The static HDR10 metadata is unable to account for the change in the content. While this is a simple example, the dynamic metadata in HDR10+ demonstrates such value across any set of scenes. This consistency allows our members to stay immersed in the content, and better preserves creative intent.</p><h1 id="925d" class="pm pn jk bf po pp pq pr hg ps pt pu hi pv pw px py pz qa qb qc qd qe qf qg qh bk"><strong class="am">Receiving HDR10+</strong></h1><p id="0211" class="pw-post-body-paragraph or os jk ot b ou qi ow ox oy qj pa pb hj qk pd pe hm ql pg ph hp qm pj pk pl il bk">At the time of launch, these requirements must be satisfied to receive HDR10+:</p><p id="8822" class="pw-post-body-paragraph or os jk ot b ou ov ow ox oy oz pa pb hj pc pd pe hm pf pg ph hp pi pj pk pl il bk">1.Member must have a Netflix Premium plan subscription</p><p id="e1e4" class="pw-post-body-paragraph or os jk ot b ou ov ow ox oy oz pa pb hj pc pd pe hm pf pg ph hp pi pj pk pl il bk">2. Title must be available in HDR10+ format</p><p id="4103" class="pw-post-body-paragraph or os jk ot b ou ov ow ox oy oz pa pb hj pc pd pe hm pf pg ph hp pi pj pk pl il bk">3. Member device must support AV1 &amp; HDR10+. Here are some examples of compatible devices:</p><ul class=""><li id="f4b3" class="or os jk ot b ou ov ow ox oy oz pa pb hj pc pd pe hm pf pg ph hp pi pj pk pl rd re rf bk">SmartTVs, mobile phones, and tablets that meet Netflix certification for HDR10+</li><li id="dff8" class="or os jk ot b ou rg ow ox oy rh pa pb hj ri pd pe hm rj pg ph hp rk pj pk pl rd re rf bk">Source device (such as set-top boxes, streaming devices, MVPDs, etc.) that meets Netflix certification for HDR10+, connected to an HDR10+ compliant display via HDMI</li></ul><p id="3a60" class="pw-post-body-paragraph or os jk ot b ou ov ow ox oy oz pa pb hj pc pd pe hm pf pg ph hp pi pj pk pl il bk">4. For TV or streaming devices, ensure that the HDR toggle is enabled in our Netflix application settings: <a class="ag hy" href="https://help.netflix.com/en/node/100220" rel="noopener ugc nofollow" target="_blank">https://help.netflix.com/en/node/100220</a></p><p id="2e59" class="pw-post-body-paragraph or os jk ot b ou ov ow ox oy oz pa pb hj pc pd pe hm pf pg ph hp pi pj pk pl il bk">Additional guidance: <a class="ag hy" href="https://help.netflix.com/en/node/13444" rel="noopener ugc nofollow" target="_blank">https://help.netflix.com/en/node/13444</a></p><h1 id="62e5" class="pm pn jk bf po pp pq pr hg ps pt pu hi pv pw px py pz qa qb qc qd qe qf qg qh bk">Summary</h1><p id="ffa0" class="pw-post-body-paragraph or os jk ot b ou qi ow ox oy qj pa pb hj qk pd pe hm ql pg ph hp qm pj pk pl il bk">More HDR content is watched every day on Netflix. Expanding the Netflix HDR ecosystem to include HDR10+ increases the accessibility of HDR content with dynamic metadata to more members, improves the viewing experience, and preserves the creative intent of our content creators. The commitment to innovation and quality underscores our dedication to delivering an immersive and authentic viewing experience for all our members.</p><h1 id="0734" class="pm pn jk bf po pp pq pr hg ps pt pu hi pv pw px py pz qa qb qc qd qe qf qg qh bk">Acknowledgements</h1><p id="5054" class="pw-post-body-paragraph or os jk ot b ou qi ow ox oy qj pa pb hj qk pd pe hm ql pg ph hp qm pj pk pl il bk">Launching HDR10+ was a collaborative effort involving multiple teams at Netflix, and we are grateful to everyone who contributed to making this idea a reality. We would like to extend our thanks to the following teams for their crucial roles in this launch:</p><ul class=""><li id="7449" class="or os jk ot b ou ov ow ox oy oz pa pb hj pc pd pe hm pf pg ph hp pi pj pk pl rd re rf bk">The various Client and Partner Engineering teams at Netflix that manage the Netflix experience across different device platforms.<br />Special acknowledgments: <a class="ag hy" href="https://www.linkedin.com/in/akshaygarg05/" rel="noopener ugc nofollow" target="_blank">Akshay Garg</a>, <a class="ag hy" href="https://www.linkedin.com/in/dashap/" rel="noopener ugc nofollow" target="_blank">Dasha Polyakova</a>, <a class="ag hy" href="https://www.linkedin.com/in/wei-vivian-li/" rel="noopener ugc nofollow" target="_blank">Vivian Li</a>, <a class="ag hy" href="https://www.linkedin.com/in/benjamintoofer/" rel="noopener ugc nofollow" target="_blank">Ben Toofer</a>, <a class="ag hy" href="https://www.linkedin.com/in/allanzp/" rel="noopener ugc nofollow" target="_blank">Allan Zhou</a>, <a class="ag hy" href="https://www.linkedin.com/in/artemdanylenko/" rel="noopener ugc nofollow" target="_blank">Artem Danylenko</a></li><li id="23d0" class="or os jk ot b ou rg ow ox oy rh pa pb hj ri pd pe hm rj pg ph hp rk pj pk pl rd re rf bk">The Encoding Technologies team that is responsible for producing optimized encodings to enable high-quality experiences for our members. Special acknowledgments: <a class="ag hy" href="https://www.linkedin.com/in/adithyaprakash/" rel="noopener ugc nofollow" target="_blank">Adithya Prakash</a>, <a class="ag hy" href="https://www.linkedin.com/in/carvalhovinicius/" rel="noopener ugc nofollow" target="_blank">Vinicius Carvalho</a></li><li id="fd90" class="or os jk ot b ou rg ow ox oy rh pa pb hj ri pd pe hm rj pg ph hp rk pj pk pl rd re rf bk">The Content Operations &amp; Innovation teams responsible for producing and delivering HDR content to Netflix, maintaining the intent of creative vision from production to streaming. Special acknowledgements: <a class="ag hy" href="https://www.linkedin.com/in/michael-keegan-072a4950/" rel="noopener ugc nofollow" target="_blank">Michael Keegan</a></li></ul><h2 id="5696" class="rl pn jk bf po hf rm ez hg hh rn fb hi hj ro hk hl hm rp hn ho hp rq hq hr rr bk">Footnotes</h2><ol class=""><li id="ddfa" class="or os jk ot b ou qi ow ox oy qj pa pb hj qk pd pe hm ql pg ph hp qm pj pk pl rs re rf bk">While we have enabled HDR10+ for distribution i.e., for what our members consume on their devices, we continue to accept only Dolby Vision masters on the ingest side, i.e., for all content delivery to Netflix as per our <a class="ag hy" href="https://partnerhelp.netflixstudios.com/hc/en-us/sections/360012197873-Branded-Delivery-Specifications" rel="noopener ugc nofollow" target="_blank">delivery specification</a>. In addition to HDR10+, we continue to serve HDR10 and DolbyVision. Our encoding pipeline is designed with flexibility and extensibility where all these HDR formats could be derived from a single DolbyVision deliverable efficiently at scale.</li><li id="d2d9" class="or os jk ot b ou rg ow ox oy rh pa pb hj ri pd pe hm rj pg ph hp rk pj pk pl rs re rf bk">We recognize that it is hard to convey visual improvements in HDR video using still photographs converted to SDR. We encourage the reader to stream Netflix content in HDR10+ and check for yourself!</li></ol></div>]]></description>
      <link>https://netflixtechblog.com/hdr10-now-streaming-on-netflix-c9ab1f4bd72b</link>
      <guid>https://netflixtechblog.com/hdr10-now-streaming-on-netflix-c9ab1f4bd72b</guid>
      <pubDate>Mon, 24 Mar 2025 19:39:00 +0100</pubDate>
    </item>
    <item>
      <title><![CDATA[Title Launch Observability at Netflix Scale]]></title>
      <description><![CDATA[<div><div><h2 id="645e" class="pw-subtitle-paragraph hr gt gu bf b hs ht hu hv hw hx hy hz ia ib ic id ie if ig cq du">Part 3: System Strategies and Architecture</h2><div></div><p id="a132" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk"><strong class="ni gv">By:</strong> <a class="af oc" href="https://www.linkedin.com/in/varun-khaitan/" rel="noopener ugc nofollow" target="_blank">Varun Khaitan</a></p><p id="6090" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">With special thanks to my stunning colleagues: <a class="af oc" href="https://www.linkedin.com/in/mallikarao/" rel="noopener ugc nofollow" target="_blank">Mallika Rao</a>, <a class="af oc" href="https://www.linkedin.com/in/esmir-mesic/" rel="noopener ugc nofollow" target="_blank">Esmir Mesic</a>, <a class="af oc" href="https://www.linkedin.com/in/hugodesmarques/" rel="noopener ugc nofollow" target="_blank">Hugo Marques</a></p><p id="6db9" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">This blog post is a continuation of <a class="af oc" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/title-launch-observability-at-netflix-scale-19ea916be1ed">Part 2</a>, where we cleared the ambiguity around title launch observability at Netflix. In this installment, we will explore the strategies, tools, and methodologies that were employed to achieve comprehensive title observability at scale.</p><h1 id="69a2" class="od oe gu bf of og oh hu oi oj ok hx ol om on oo op oq or os ot ou ov ow ox oy bk">Defining the observability endpoint</h1><p id="bc9b" class="pw-post-body-paragraph ng nh gu ni b hs oz nk nl hv pa nn no np pb nr ns nt pc nv nw nx pd nz oa ob gn bk">To create a comprehensive solution, we decided to introduce observability endpoints first. Each microservice involved in our <strong class="ni gv">Personalization stack</strong> that integrated with our observability solution had to introduce a new “Title Health” endpoint. Our goal was for each new endpoint to adhere to a few principles:</p><ol class=""><li id="1298" class="ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob pe pf pg bk">Accurate reflection of production behavior</li><li id="cade" class="ng nh gu ni b hs ph nk nl hv pi nn no np pj nr ns nt pk nv nw nx pl nz oa ob pe pf pg bk">Standardization across all endpoints</li><li id="7204" class="ng nh gu ni b hs ph nk nl hv pi nn no np pj nr ns nt pk nv nw nx pl nz oa ob pe pf pg bk">Answering the Insight Triad: “Healthy” or not, why not and how to fix it.</li></ol><p id="0cac" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk"><strong class="ni gv">Accurately Reflecting Production Behavior</strong></p><p id="4dbf" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">A key part of our solution is insights into production behavior, which necessitates our requests to the endpoint result in traffic to the real service functions that mimics the same pathways the traffic would take if it came from the usual callers.</p><p id="377d" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">In order to allow for this mimicking, many systems implement an “event” handling, where they convert our request into a call to the real service with properties enabled to log when titles are filtered out of their response and why. Building services that adhere to software best practices, such as Object-Oriented Programming (OOP), the SOLID principles, and modularization, is crucial to have success at this stage. Without these practices, service endpoints may become tightly coupled to business logic, making it challenging and costly to add a new endpoint that seamlessly integrates with the observability solution while following the same production logic.</p><figure class="pp pq pr ps pt pu pm pn paragraph-image"><div role="button" tabindex="0" class="pv pw fj px bh py"><div class="pm pn po"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*8s2gCb2Pqw2Q0Frq%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*8s2gCb2Pqw2Q0Frq%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*8s2gCb2Pqw2Q0Frq%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*8s2gCb2Pqw2Q0Frq%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*8s2gCb2Pqw2Q0Frq%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*8s2gCb2Pqw2Q0Frq%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*8s2gCb2Pqw2Q0Frq%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*8s2gCb2Pqw2Q0Frq 640w, https://miro.medium.com/v2/resize:fit:720/0*8s2gCb2Pqw2Q0Frq 720w, https://miro.medium.com/v2/resize:fit:750/0*8s2gCb2Pqw2Q0Frq 750w, https://miro.medium.com/v2/resize:fit:786/0*8s2gCb2Pqw2Q0Frq 786w, https://miro.medium.com/v2/resize:fit:828/0*8s2gCb2Pqw2Q0Frq 828w, https://miro.medium.com/v2/resize:fit:1100/0*8s2gCb2Pqw2Q0Frq 1100w, https://miro.medium.com/v2/resize:fit:1400/0*8s2gCb2Pqw2Q0Frq 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh mn pz c" width="700" height="406" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="qa ff qb pm pn qc qd bf b bg z du"><em class="qe">A service with modular business logic facilitates the seamless addition of an observability endpoint.</em></figcaption></figure><p id="2da9" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk"><strong class="ni gv">Standardization</strong></p><p id="e9b6" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">To standardize communication between our observability service and the personalization stack’s observability endpoints, we’ve developed a stable proto request/response format. This centralized format, defined and maintained by our team, ensures all endpoints adhere to a consistent protocol. As a result, requests are uniformly handled, and responses are processed cohesively. This standardization enhances adoption within the personalization stack, simplifies the system, and improves understanding and debuggability for engineers.</p><figure class="pp pq pr ps pt pu pm pn paragraph-image"><div role="button" tabindex="0" class="pv pw fj px bh py"><div class="pm pn qf"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*P-0nxUAHve77yBtv%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*P-0nxUAHve77yBtv%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*P-0nxUAHve77yBtv%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*P-0nxUAHve77yBtv%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*P-0nxUAHve77yBtv%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*P-0nxUAHve77yBtv%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*P-0nxUAHve77yBtv%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*P-0nxUAHve77yBtv 640w, https://miro.medium.com/v2/resize:fit:720/0*P-0nxUAHve77yBtv 720w, https://miro.medium.com/v2/resize:fit:750/0*P-0nxUAHve77yBtv 750w, https://miro.medium.com/v2/resize:fit:786/0*P-0nxUAHve77yBtv 786w, https://miro.medium.com/v2/resize:fit:828/0*P-0nxUAHve77yBtv 828w, https://miro.medium.com/v2/resize:fit:1100/0*P-0nxUAHve77yBtv 1100w, https://miro.medium.com/v2/resize:fit:1400/0*P-0nxUAHve77yBtv 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh mn pz c" width="700" height="152" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="qa ff qb pm pn qc qd bf b bg z du"><em class="qe">The request schema for the observability endpoint.</em></figcaption></figure><p id="97c5" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk"><strong class="ni gv">The Insight Triad API</strong></p><p id="95d2" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">To efficiently understand the health of a title and triage issues quickly, all implementations of the observability endpoint must answer: is the title eligible for this phase of promotion, if not — why is it not eligible, and what can be done to fix any problems.</p><p id="a1a7" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">The end-users of this observability system are Launch Managers, whose job it is to ensure smooth title launches. As such, they must be able to quickly see whether there is a problem, what the problem is, and how to solve it. Teams implementing the endpoint must provide as much information as possible so that a non-engineer (Launch Manager) can understand the root cause of the issue and fix any title setup issues as they arise. They must also provide enough information for partner engineers to identify the problem with the underlying service in cases of system-level issues.</p><p id="29f8" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">These requirements are captured in the following protobuf object that defines the endpoint response.</p><figure class="pp pq pr ps pt pu pm pn paragraph-image"><div role="button" tabindex="0" class="pv pw fj px bh py"><div class="pm pn qg"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*aeo7vs3h2Z5JKH5t%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*aeo7vs3h2Z5JKH5t%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*aeo7vs3h2Z5JKH5t%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*aeo7vs3h2Z5JKH5t%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*aeo7vs3h2Z5JKH5t%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*aeo7vs3h2Z5JKH5t%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*aeo7vs3h2Z5JKH5t%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*aeo7vs3h2Z5JKH5t 640w, https://miro.medium.com/v2/resize:fit:720/0*aeo7vs3h2Z5JKH5t 720w, https://miro.medium.com/v2/resize:fit:750/0*aeo7vs3h2Z5JKH5t 750w, https://miro.medium.com/v2/resize:fit:786/0*aeo7vs3h2Z5JKH5t 786w, https://miro.medium.com/v2/resize:fit:828/0*aeo7vs3h2Z5JKH5t 828w, https://miro.medium.com/v2/resize:fit:1100/0*aeo7vs3h2Z5JKH5t 1100w, https://miro.medium.com/v2/resize:fit:1400/0*aeo7vs3h2Z5JKH5t 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh mn pz c" width="700" height="150" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="qa ff qb pm pn qc qd bf b bg z du"><em class="qe">The response schema for the observability endpoint.</em></figcaption></figure><h1 id="b501" class="od oe gu bf of og oh hu oi oj ok hx ol om on oo op oq or os ot ou ov ow ox oy bk">High level architecture</h1><p id="34a0" class="pw-post-body-paragraph ng nh gu ni b hs oz nk nl hv pa nn no np pb nr ns nt pc nv nw nx pd nz oa ob gn bk">We’ve distilled our comprehensive solution into the following key steps, capturing the essence of our approach:</p><ol class=""><li id="7946" class="ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob pe pf pg bk">Establish observability endpoints across all services within our Personalization and Discovery Stack.</li><li id="d917" class="ng nh gu ni b hs ph nk nl hv pi nn no np pj nr ns nt pk nv nw nx pl nz oa ob pe pf pg bk">Implement proactive monitoring for each of these endpoints.</li><li id="feab" class="ng nh gu ni b hs ph nk nl hv pi nn no np pj nr ns nt pk nv nw nx pl nz oa ob pe pf pg bk">Track real-time title impressions from the Netflix UI.</li><li id="08ec" class="ng nh gu ni b hs ph nk nl hv pi nn no np pj nr ns nt pk nv nw nx pl nz oa ob pe pf pg bk">Store the data in an optimized, highly distributed datastore.</li><li id="109e" class="ng nh gu ni b hs ph nk nl hv pi nn no np pj nr ns nt pk nv nw nx pl nz oa ob pe pf pg bk">Offer easy-to-integrate APIs for our dashboard, enabling stakeholders to track specific titles effectively.</li><li id="51d0" class="ng nh gu ni b hs ph nk nl hv pi nn no np pj nr ns nt pk nv nw nx pl nz oa ob pe pf pg bk">“Time Travel” to validate ahead of time.</li></ol><figure class="pp pq pr ps pt pu pm pn paragraph-image"><div role="button" tabindex="0" class="pv pw fj px bh py"><div class="pm pn qh"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*1h2cwZDfmz8nis_h%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*1h2cwZDfmz8nis_h%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*1h2cwZDfmz8nis_h%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*1h2cwZDfmz8nis_h%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*1h2cwZDfmz8nis_h%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*1h2cwZDfmz8nis_h%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*1h2cwZDfmz8nis_h%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*1h2cwZDfmz8nis_h 640w, https://miro.medium.com/v2/resize:fit:720/0*1h2cwZDfmz8nis_h 720w, https://miro.medium.com/v2/resize:fit:750/0*1h2cwZDfmz8nis_h 750w, https://miro.medium.com/v2/resize:fit:786/0*1h2cwZDfmz8nis_h 786w, https://miro.medium.com/v2/resize:fit:828/0*1h2cwZDfmz8nis_h 828w, https://miro.medium.com/v2/resize:fit:1100/0*1h2cwZDfmz8nis_h 1100w, https://miro.medium.com/v2/resize:fit:1400/0*1h2cwZDfmz8nis_h 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh mn pz c" width="700" height="445" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="qa ff qb pm pn qc qd bf b bg z du"><em class="qe">Observability stack high level architecture diagram</em></figcaption></figure><p id="9f94" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">In the following sections, we will explore each of these concepts and components as illustrated in the diagram above.</p><h1 id="ea0e" class="od oe gu bf of og oh hu oi oj ok hx ol om on oo op oq or os ot ou ov ow ox oy bk">Key Features</h1><h2 id="5bfa" class="qi oe gu bf of qj qk dy oi ql qm ea ol np qn qo qp nt qq qr qs nx qt qu qv qw bk">Proactive monitoring through scheduled collectors jobs</h2><p id="2070" class="pw-post-body-paragraph ng nh gu ni b hs oz nk nl hv pa nn no np pb nr ns nt pc nv nw nx pd nz oa ob gn bk">Our Title Health microservice runs a scheduled collector job every 30 minutes for most of our personalization stack.</p><p id="9419" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">For each Netflix row we support (such as Trending Now, Coming Soon, etc.), there is a dedicated collector. These collectors retrieve the relevant list of titles from our catalog that qualify for a specific row by interfacing with our catalog services. These services are informed about the expected subset of titles for each row, for which we are assessing title health.</p><p id="0bb9" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">Once a collector retrieves its list of candidate titles, it orchestrates batched calls to assigned row services using the above standardized schema to retrieve all the relevant health information of the titles. Additionally, some collectors will instead poll our kafka queue for impressions data.</p><h2 id="07cf" class="qi oe gu bf of qj qk dy oi ql qm ea ol np qn qo qp nt qq qr qs nx qt qu qv qw bk">Real-time Title Impressions and Kafka Queue</h2><p id="c0e6" class="pw-post-body-paragraph ng nh gu ni b hs oz nk nl hv pa nn no np pb nr ns nt pc nv nw nx pd nz oa ob gn bk">In addition to evaluating title health via our personalization stack services, we also keep an eye on how our recommendation algorithms treat titles by reviewing impressions data. It’s essential that our algorithms treat all titles equitably, for each one has limitless potential.</p><p id="1754" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">This data is processed from a real-time impressions stream into a Kafka queue, which our title health system regularly polls. Specialized collectors access the Kafka queue every two minutes to retrieve impressions data. This data is then aggregated in minute(s) intervals, calculating the number of impressions titles receive in near-real-time, and presented as an additional health status indicator for stakeholders.</p><h2 id="958b" class="qi oe gu bf of qj qk dy oi ql qm ea ol np qn qo qp nt qq qr qs nx qt qu qv qw bk">Data storage and distribution through Hollow Feeds</h2><p id="0fe5" class="pw-post-body-paragraph ng nh gu ni b hs oz nk nl hv pa nn no np pb nr ns nt pc nv nw nx pd nz oa ob gn bk"><a class="af oc" href="https://hollow.how/" rel="noopener ugc nofollow" target="_blank">Netflix Hollow</a> is an Open Source java library and toolset for disseminating in-memory datasets from a single producer to many consumers for high performance read-only access. Given the shape of our data, hollow feeds are an excellent strategy to distribute the data across our service boxes.</p><p id="7314" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">Once collectors gather health data from partner services in the personalization stack or from our impressions stream, this data is stored in a dedicated Hollow feed for each collector. Hollow offers numerous features that help us monitor the overall health of a Netflix row, including ensuring there are no large-scale issues across a feed publish. It also allows us to track the history of each title by maintaining a per-title data history, calculate differences between previous and current data versions, and roll back to earlier versions if a problematic data change is detected.</p><h2 id="321f" class="qi oe gu bf of qj qk dy oi ql qm ea ol np qn qo qp nt qq qr qs nx qt qu qv qw bk">Observability Dashboard using Health Check Engine</h2><p id="d7ac" class="pw-post-body-paragraph ng nh gu ni b hs oz nk nl hv pa nn no np pb nr ns nt pc nv nw nx pd nz oa ob gn bk">We maintain several dashboards that utilize our title health service to present the status of titles to stakeholders. These user interfaces access an endpoint in our service, enabling them to request the current status of a title across all supported rows. This endpoint efficiently reads from all available Hollow Feeds to obtain the current status, thanks to Hollow’s in-memory capabilities. The results are returned in a standardized format, ensuring easy support for future UIs.</p><p id="5c46" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">Additionally, we have other endpoints that can summarize the health of a title across subsets of sections to highlight specific member experiences.</p><figure class="pp pq pr ps pt pu pm pn paragraph-image"><div role="button" tabindex="0" class="pv pw fj px bh py"><div class="pm pn qg"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*dBFS1pBlqNoCUHwV%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*dBFS1pBlqNoCUHwV%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*dBFS1pBlqNoCUHwV%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*dBFS1pBlqNoCUHwV%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*dBFS1pBlqNoCUHwV%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*dBFS1pBlqNoCUHwV%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*dBFS1pBlqNoCUHwV%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*dBFS1pBlqNoCUHwV 640w, https://miro.medium.com/v2/resize:fit:720/0*dBFS1pBlqNoCUHwV 720w, https://miro.medium.com/v2/resize:fit:750/0*dBFS1pBlqNoCUHwV 750w, https://miro.medium.com/v2/resize:fit:786/0*dBFS1pBlqNoCUHwV 786w, https://miro.medium.com/v2/resize:fit:828/0*dBFS1pBlqNoCUHwV 828w, https://miro.medium.com/v2/resize:fit:1100/0*dBFS1pBlqNoCUHwV 1100w, https://miro.medium.com/v2/resize:fit:1400/0*dBFS1pBlqNoCUHwV 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh mn pz c" width="700" height="116" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="qa ff qb pm pn qc qd bf b bg z du">Message depicting a dashboard request.</figcaption></figure><h2 id="e3b2" class="qi oe gu bf of qj qk dy oi ql qm ea ol np qn qo qp nt qq qr qs nx qt qu qv qw bk">Time Traveling: Catching before launch</h2><figure class="pp pq pr ps pt pu pm pn paragraph-image"><div role="button" tabindex="0" class="pv pw fj px bh py"><div class="pm pn qg"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*Zz2Y8yjPAsbG5WVR%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*Zz2Y8yjPAsbG5WVR%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*Zz2Y8yjPAsbG5WVR%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*Zz2Y8yjPAsbG5WVR%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*Zz2Y8yjPAsbG5WVR%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*Zz2Y8yjPAsbG5WVR%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*Zz2Y8yjPAsbG5WVR%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*Zz2Y8yjPAsbG5WVR 640w, https://miro.medium.com/v2/resize:fit:720/0*Zz2Y8yjPAsbG5WVR 720w, https://miro.medium.com/v2/resize:fit:750/0*Zz2Y8yjPAsbG5WVR 750w, https://miro.medium.com/v2/resize:fit:786/0*Zz2Y8yjPAsbG5WVR 786w, https://miro.medium.com/v2/resize:fit:828/0*Zz2Y8yjPAsbG5WVR 828w, https://miro.medium.com/v2/resize:fit:1100/0*Zz2Y8yjPAsbG5WVR 1100w, https://miro.medium.com/v2/resize:fit:1400/0*Zz2Y8yjPAsbG5WVR 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh mn pz c" width="700" height="541" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="3f42" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">Titles launching at Netflix go through several phases of pre-promotion before ultimately launching on our platform. For each of these phases, the first several hours of promotion are critical for the reach and effective personalization of a title, especially once the title has launched. Thus, to prevent issues as titles go through the launch lifecycle, our observability system needs to be capable of simulating traffic ahead of time so that relevant teams can catch and fix issues before they impact members. We call this capability <strong class="ni gv">“Time Travel”</strong>.</p><p id="6a4a" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">Many of the metadata and assets involved in title setup have specific timelines for when they become available to members. To determine if a title will be viewable at the start of an experience, we must simulate a request to a partner service as if it were from a future time when those specific metadata or assets are available. This is achieved by including a future timestamp in our request to the observability endpoint, corresponding to when the title is expected to appear for a given experience. The endpoint then communicates with any further downstream services using the context of that future timestamp.</p><figure class="pp pq pr ps pt pu pm pn paragraph-image"><div role="button" tabindex="0" class="pv pw fj px bh py"><div class="pm pn qx"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*jrdqpJmp0lzna6Zc%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*jrdqpJmp0lzna6Zc%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*jrdqpJmp0lzna6Zc%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*jrdqpJmp0lzna6Zc%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*jrdqpJmp0lzna6Zc%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*jrdqpJmp0lzna6Zc%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*jrdqpJmp0lzna6Zc%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*jrdqpJmp0lzna6Zc 640w, https://miro.medium.com/v2/resize:fit:720/0*jrdqpJmp0lzna6Zc 720w, https://miro.medium.com/v2/resize:fit:750/0*jrdqpJmp0lzna6Zc 750w, https://miro.medium.com/v2/resize:fit:786/0*jrdqpJmp0lzna6Zc 786w, https://miro.medium.com/v2/resize:fit:828/0*jrdqpJmp0lzna6Zc 828w, https://miro.medium.com/v2/resize:fit:1100/0*jrdqpJmp0lzna6Zc 1100w, https://miro.medium.com/v2/resize:fit:1400/0*jrdqpJmp0lzna6Zc 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh mn pz c" width="700" height="118" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="qa ff qb pm pn qc qd bf b bg z du">An example request with a future timestamp.</figcaption></figure><h1 id="a91c" class="od oe gu bf of og oh hu oi oj ok hx ol om on oo op oq or os ot ou ov ow ox oy bk">Conclusion</h1><p id="0f5b" class="pw-post-body-paragraph ng nh gu ni b hs oz nk nl hv pa nn no np pb nr ns nt pc nv nw nx pd nz oa ob gn bk">Throughout this series, we’ve explored the journey of enhancing title launch observability at Netflix. In <a class="af oc" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/title-launch-observability-at-netflix-scale-c88c586629eb">Part 1</a>, we identified the challenges of managing vast content launches and the need for scalable solutions to ensure each title’s success. <a class="af oc" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/title-launch-observability-at-netflix-scale-19ea916be1ed">Part 2</a> highlighted the strategic approach to navigating ambiguity, introducing “Title Health” as a framework to align teams and prioritize core issues. In this final part, we detailed the sophisticated system strategies and architecture, including observability endpoints, proactive monitoring, and “Time Travel” capabilities; all designed to ensure a thrilling viewing experience.</p><p id="ad08" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">By investing in these innovative solutions, we enhance the discoverability and success of each title, fostering trust with content creators and partners. This journey not only bolsters our operational capabilities but also lays the groundwork for future innovations, ensuring that every story reaches its intended audience and that every member enjoys their favorite titles on Netflix.</p><p id="4e68" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">Thank you for joining us on this exploration, and stay tuned for more insights and innovations as we continue to entertain the world.</p></div></div>]]></description>
      <link>https://netflixtechblog.com/title-launch-observability-at-netflix-scale-8efe69ebd653</link>
      <guid>https://netflixtechblog.com/title-launch-observability-at-netflix-scale-8efe69ebd653</guid>
      <pubDate>Wed, 05 Mar 2025 02:24:00 +0100</pubDate>
    </item>
    <item>
      <title><![CDATA[Introducing Impressions at Netflix]]></title>
      <description><![CDATA[<div><div><h2 id="097f" class="pw-subtitle-paragraph hr gt gu bf b hs ht hu hv hw hx hy hz ia ib ic id ie if ig cq du">Part 1: Creating the Source of Truth for Impressions</h2><div></div><p id="ce57" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk"><strong class="ni gv">By:</strong> <a class="af oc" href="https://www.linkedin.com/in/tulikabhatt/" rel="noopener ugc nofollow" target="_blank">Tulika Bhatt</a></p><p id="e560" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">Imagine scrolling through Netflix, where each movie poster or promotional banner competes for your attention. Every image you hover over isn’t just a visual placeholder; it’s a critical data point that fuels our sophisticated personalization engine. At Netflix, we call these images ‘impressions,’ and they play a pivotal role in transforming your interaction from simple browsing into an immersive binge-watching experience, all tailored to your unique tastes.</p><p id="3c19" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">Capturing these moments and turning them into a personalized journey is no simple feat. It requires a state-of-the-art system that can track and process these impressions while maintaining a detailed history of each profile’s exposure. This nuanced integration of data and technology empowers us to offer bespoke content recommendations.</p><p id="d15d" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">In this multi-part blog series, we take you behind the scenes of our system that processes billions of impressions daily. We will explore the challenges we encounter and unveil how we are building a resilient solution that transforms these client-side impressions into a personalized content discovery experience for every Netflix viewer.</p><figure class="og oh oi oj ok ol od oe paragraph-image"><div role="button" tabindex="0" class="om on fj oo bh op"><div class="od oe of"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*T6tQiUj-VDtyEhd1%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*T6tQiUj-VDtyEhd1%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*T6tQiUj-VDtyEhd1%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*T6tQiUj-VDtyEhd1%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*T6tQiUj-VDtyEhd1%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*T6tQiUj-VDtyEhd1%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*T6tQiUj-VDtyEhd1%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*T6tQiUj-VDtyEhd1 640w, https://miro.medium.com/v2/resize:fit:720/0*T6tQiUj-VDtyEhd1 720w, https://miro.medium.com/v2/resize:fit:750/0*T6tQiUj-VDtyEhd1 750w, https://miro.medium.com/v2/resize:fit:786/0*T6tQiUj-VDtyEhd1 786w, https://miro.medium.com/v2/resize:fit:828/0*T6tQiUj-VDtyEhd1 828w, https://miro.medium.com/v2/resize:fit:1100/0*T6tQiUj-VDtyEhd1 1100w, https://miro.medium.com/v2/resize:fit:1400/0*T6tQiUj-VDtyEhd1 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh mn oq c" width="700" height="330" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="or ff os od oe ot ou bf b bg z du">Impressions on homepage</figcaption></figure><h1 id="9ac6" class="ov ow gu bf ox oy oz hu pa pb pc hx pd pe pf pg ph pi pj pk pl pm pn po pp pq bk">Why do we need impression history?</h1><h2 id="16ca" class="pr ow gu bf ox ps pt dy pa pu pv ea pd np pw px py nt pz qa qb nx qc qd qe qf bk">Enhanced Personalization</h2><p id="9340" class="pw-post-body-paragraph ng nh gu ni b hs qg nk nl hv qh nn no np qi nr ns nt qj nv nw nx qk nz oa ob gn bk">To tailor recommendations more effectively, it’s crucial to track what content a user has already encountered. Having impression history helps us achieve this by allowing us to identify content that has been displayed on the homepage but not engaged with, helping us deliver fresh, engaging recommendations.</p><h2 id="b9ca" class="pr ow gu bf ox ps pt dy pa pu pv ea pd np pw px py nt pz qa qb nx qc qd qe qf bk">Frequency Capping</h2><p id="549e" class="pw-post-body-paragraph ng nh gu ni b hs qg nk nl hv qh nn no np qi nr ns nt qj nv nw nx qk nz oa ob gn bk">By maintaining a history of impressions, we can implement frequency capping to prevent over-exposure to the same content. This ensures users aren’t repeatedly shown identical options, keeping the viewing experience vibrant and reducing the risk of frustration or disengagement.</p><h2 id="dc85" class="pr ow gu bf ox ps pt dy pa pu pv ea pd np pw px py nt pz qa qb nx qc qd qe qf bk">Highlighting New Releases</h2><p id="ba61" class="pw-post-body-paragraph ng nh gu ni b hs qg nk nl hv qh nn no np qi nr ns nt qj nv nw nx qk nz oa ob gn bk">For new content, impression history helps us monitor initial user interactions and adjust our merchandising efforts accordingly. We can experiment with different content placements or promotional strategies to boost visibility and engagement.</p><h2 id="277c" class="pr ow gu bf ox ps pt dy pa pu pv ea pd np pw px py nt pz qa qb nx qc qd qe qf bk">Analytical Insights</h2><p id="1d77" class="pw-post-body-paragraph ng nh gu ni b hs qg nk nl hv qh nn no np qi nr ns nt qj nv nw nx qk nz oa ob gn bk">Additionally, impression history offers insightful information for addressing a number of platform-related analytics queries. Analyzing impression history, for example, might help determine how well a specific row on the home page is functioning or assess the effectiveness of a merchandising strategy.</p><h1 id="6595" class="ov ow gu bf ox oy oz hu pa pb pc hx pd pe pf pg ph pi pj pk pl pm pn po pp pq bk">Architecture Overview</h1><p id="3dda" class="pw-post-body-paragraph ng nh gu ni b hs qg nk nl hv qh nn no np qi nr ns nt qj nv nw nx qk nz oa ob gn bk">The first pivotal step in managing impressions begins with the creation of a Source-of-Truth (SOT) dataset. This foundational dataset is essential, as it supports various downstream workflows and enables a multitude of use cases.</p><h2 id="50f2" class="pr ow gu bf ox ps pt dy pa pu pv ea pd np pw px py nt pz qa qb nx qc qd qe qf bk">Collecting Raw Impression Events</h2><p id="79ce" class="pw-post-body-paragraph ng nh gu ni b hs qg nk nl hv qh nn no np qi nr ns nt qj nv nw nx qk nz oa ob gn bk">As Netflix members explore our platform, their interactions with the user interface spark a vast array of raw events. These events are promptly relayed from the client side to our servers, entering a centralized event processing queue. This queue ensures we are consistently capturing raw events from our global user base.</p><p id="2110" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">After raw events are collected into a centralized queue, a custom event extractor processes this data to identify and extract all impression events. These extracted events are then routed to an Apache Kafka topic for immediate processing needs and simultaneously stored in an Apache Iceberg table for long-term retention and historical analysis. This dual-path approach leverages Kafka’s capability for low-latency streaming and Iceberg’s efficient management of large-scale, immutable datasets, ensuring both real-time responsiveness and comprehensive historical data availability.</p><figure class="og oh oi oj ok ol od oe paragraph-image"><div role="button" tabindex="0" class="om on fj oo bh op"><div class="od oe ql"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*4NRQp10pg9KK_GKU%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*4NRQp10pg9KK_GKU%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*4NRQp10pg9KK_GKU%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*4NRQp10pg9KK_GKU%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*4NRQp10pg9KK_GKU%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*4NRQp10pg9KK_GKU%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*4NRQp10pg9KK_GKU%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*4NRQp10pg9KK_GKU 640w, https://miro.medium.com/v2/resize:fit:720/0*4NRQp10pg9KK_GKU 720w, https://miro.medium.com/v2/resize:fit:750/0*4NRQp10pg9KK_GKU 750w, https://miro.medium.com/v2/resize:fit:786/0*4NRQp10pg9KK_GKU 786w, https://miro.medium.com/v2/resize:fit:828/0*4NRQp10pg9KK_GKU 828w, https://miro.medium.com/v2/resize:fit:1100/0*4NRQp10pg9KK_GKU 1100w, https://miro.medium.com/v2/resize:fit:1400/0*4NRQp10pg9KK_GKU 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh mn oq c" width="700" height="303" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="or ff os od oe ot ou bf b bg z du">Collecting raw impression events</figcaption></figure><h2 id="48cf" class="pr ow gu bf ox ps pt dy pa pu pv ea pd np pw px py nt pz qa qb nx qc qd qe qf bk">Filtering &amp; Enriching Raw Impressions</h2><p id="829b" class="pw-post-body-paragraph ng nh gu ni b hs qg nk nl hv qh nn no np qi nr ns nt qj nv nw nx qk nz oa ob gn bk">Once the raw impression events are queued, a stateless Apache Flink job takes charge, meticulously processing this data. It filters out any invalid entries and enriches the valid ones with additional metadata, such as show or movie title details, and the specific page and row location where each impression was presented to users. This refined output is then structured using an Avro schema, establishing a definitive source of truth for Netflix’s impression data. The enriched data is seamlessly accessible for both real-time applications via Kafka and historical analysis through storage in an Apache Iceberg table. This dual availability ensures immediate processing capabilities alongside comprehensive long-term data retention.</p><figure class="og oh oi oj ok ol od oe paragraph-image"><div role="button" tabindex="0" class="om on fj oo bh op"><div class="od oe qm"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*Lhs-gvhMuIyKylHt%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*Lhs-gvhMuIyKylHt%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*Lhs-gvhMuIyKylHt%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*Lhs-gvhMuIyKylHt%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*Lhs-gvhMuIyKylHt%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*Lhs-gvhMuIyKylHt%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*Lhs-gvhMuIyKylHt%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*Lhs-gvhMuIyKylHt 640w, https://miro.medium.com/v2/resize:fit:720/0*Lhs-gvhMuIyKylHt 720w, https://miro.medium.com/v2/resize:fit:750/0*Lhs-gvhMuIyKylHt 750w, https://miro.medium.com/v2/resize:fit:786/0*Lhs-gvhMuIyKylHt 786w, https://miro.medium.com/v2/resize:fit:828/0*Lhs-gvhMuIyKylHt 828w, https://miro.medium.com/v2/resize:fit:1100/0*Lhs-gvhMuIyKylHt 1100w, https://miro.medium.com/v2/resize:fit:1400/0*Lhs-gvhMuIyKylHt 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh mn oq c" width="700" height="447" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="or ff os od oe ot ou bf b bg z du">Impression Source-of-Truth architecture</figcaption></figure><h2 id="d39b" class="pr ow gu bf ox ps pt dy pa pu pv ea pd np pw px py nt pz qa qb nx qc qd qe qf bk">Ensuring High Quality Impressions</h2><p id="d682" class="pw-post-body-paragraph ng nh gu ni b hs qg nk nl hv qh nn no np qi nr ns nt qj nv nw nx qk nz oa ob gn bk">Maintaining the highest quality of impressions is a top priority. We accomplish this by gathering detailed column-level metrics that offer insights into the state and quality of each impression. These metrics include everything from validating identifiers to checking that essential columns are properly filled. The data collected feeds into a comprehensive quality dashboard and supports a tiered threshold-based alerting system. These alerts promptly notify us of any potential issues, enabling us to swiftly address regressions. Additionally, while enriching the data, we ensure that all columns are in agreement with each other, offering in-place corrections wherever possible to deliver accurate data.</p><figure class="og oh oi oj ok ol od oe paragraph-image"><div role="button" tabindex="0" class="om on fj oo bh op"><div class="od oe qn"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*VWssCnOIabEqo02H%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*VWssCnOIabEqo02H%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*VWssCnOIabEqo02H%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*VWssCnOIabEqo02H%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*VWssCnOIabEqo02H%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*VWssCnOIabEqo02H%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*VWssCnOIabEqo02H%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*VWssCnOIabEqo02H 640w, https://miro.medium.com/v2/resize:fit:720/0*VWssCnOIabEqo02H 720w, https://miro.medium.com/v2/resize:fit:750/0*VWssCnOIabEqo02H 750w, https://miro.medium.com/v2/resize:fit:786/0*VWssCnOIabEqo02H 786w, https://miro.medium.com/v2/resize:fit:828/0*VWssCnOIabEqo02H 828w, https://miro.medium.com/v2/resize:fit:1100/0*VWssCnOIabEqo02H 1100w, https://miro.medium.com/v2/resize:fit:1400/0*VWssCnOIabEqo02H 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh mn oq c" width="700" height="428" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="or ff os od oe ot ou bf b bg z du">Dashboard showing mismatch count between two columns- entityId and videoId</figcaption></figure><h1 id="fed2" class="ov ow gu bf ox oy oz hu pa pb pc hx pd pe pf pg ph pi pj pk pl pm pn po pp pq bk">Configuration</h1><p id="9417" class="pw-post-body-paragraph ng nh gu ni b hs qg nk nl hv qh nn no np qi nr ns nt qj nv nw nx qk nz oa ob gn bk">We handle a staggering volume of 1 to 1.5 million impression events globally every second, with each event approximately 1.2KB in size. To efficiently process this massive influx in real-time, we employ Apache Flink for its low-latency stream processing capabilities, which seamlessly integrates both batch and stream processing to facilitate efficient backfilling of historical data and ensure consistency across real-time and historical analyses. Our Flink configuration includes 8 task managers per region, each equipped with 8 CPU cores and 32GB of memory, operating at a parallelism of 48, allowing us to handle the necessary scale and speed for seamless performance delivery. The Flink job’s sink is equipped with a data mesh connector, as detailed in our <a class="af oc" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/data-mesh-a-data-movement-and-processing-platform-netflix-1288bcab2873">Data Mesh platform</a> which has two outputs: Kafka and Iceberg. This setup allows for efficient streaming of real-time data through Kafka and the preservation of historical data in Iceberg, providing a comprehensive and flexible data processing and storage solution.</p><figure class="og oh oi oj ok ol od oe paragraph-image"><div role="button" tabindex="0" class="om on fj oo bh op"><div class="od oe of"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*B-hm-UJMBV7-WOb6%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*B-hm-UJMBV7-WOb6%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*B-hm-UJMBV7-WOb6%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*B-hm-UJMBV7-WOb6%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*B-hm-UJMBV7-WOb6%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*B-hm-UJMBV7-WOb6%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*B-hm-UJMBV7-WOb6%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*B-hm-UJMBV7-WOb6 640w, https://miro.medium.com/v2/resize:fit:720/0*B-hm-UJMBV7-WOb6 720w, https://miro.medium.com/v2/resize:fit:750/0*B-hm-UJMBV7-WOb6 750w, https://miro.medium.com/v2/resize:fit:786/0*B-hm-UJMBV7-WOb6 786w, https://miro.medium.com/v2/resize:fit:828/0*B-hm-UJMBV7-WOb6 828w, https://miro.medium.com/v2/resize:fit:1100/0*B-hm-UJMBV7-WOb6 1100w, https://miro.medium.com/v2/resize:fit:1400/0*B-hm-UJMBV7-WOb6 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh mn oq c" width="700" height="235" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="or ff os od oe ot ou bf b bg z du">Raw impressions records per second</figcaption></figure><p id="8d4a" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">We utilize the ‘island model’ for deploying our Flink jobs, where all dependencies for a given application reside within a single region. This approach ensures high availability by isolating regions, so if one becomes degraded, others remain unaffected, allowing traffic to be shifted between regions to maintain service continuity. Thus, all data in one region is processed by the Flink job deployed within that region.</p><h1 id="9ca3" class="ov ow gu bf ox oy oz hu pa pb pc hx pd pe pf pg ph pi pj pk pl pm pn po pp pq bk">Future Work</h1><h2 id="0a4f" class="pr ow gu bf ox ps pt dy pa pu pv ea pd np pw px py nt pz qa qb nx qc qd qe qf bk">Addressing the Challenge of Unschematized Events</h2><p id="ef75" class="pw-post-body-paragraph ng nh gu ni b hs qg nk nl hv qh nn no np qi nr ns nt qj nv nw nx qk nz oa ob gn bk">Allowing raw events to land on our centralized processing queue unschematized offers significant flexibility, but it also introduces challenges. Without a defined schema, it can be difficult to determine whether missing data was intentional or due to a logging error. We are investigating solutions to introduce schema management that maintains flexibility while providing clarity.</p><h2 id="8f38" class="pr ow gu bf ox ps pt dy pa pu pv ea pd np pw px py nt pz qa qb nx qc qd qe qf bk">Automating Performance Tuning with Autoscalers</h2><p id="c962" class="pw-post-body-paragraph ng nh gu ni b hs qg nk nl hv qh nn no np qi nr ns nt qj nv nw nx qk nz oa ob gn bk">Tuning the performance of our Apache Flink jobs is currently a manual process. The next step is to integrate with autoscalers, which can dynamically adjust resources based on workload demands. This integration will not only optimize performance but also ensure more efficient resource utilization.</p><h2 id="fa29" class="pr ow gu bf ox ps pt dy pa pu pv ea pd np pw px py nt pz qa qb nx qc qd qe qf bk">Improving Data Quality Alerts</h2><p id="64ef" class="pw-post-body-paragraph ng nh gu ni b hs qg nk nl hv qh nn no np qi nr ns nt qj nv nw nx qk nz oa ob gn bk">Right now, there’s a lot of business rules dictating when a data quality alert needs to be fired. This leads to a lot of false positives that require manual judgement. A lot of times it is difficult to track changes leading to regression due to inadequate data lineage information. We are investing in building a comprehensive data quality platform that more intelligently identifies anomalies in our impression stream, keeps track of data lineage and data governance, and also, generates alerts notifying producers of any regressions. This approach will enhance efficiency, reduce manual oversight, and ensure a higher standard of data integrity.</p><h1 id="8555" class="ov ow gu bf ox oy oz hu pa pb pc hx pd pe pf pg ph pi pj pk pl pm pn po pp pq bk">Conclusion</h1><p id="1482" class="pw-post-body-paragraph ng nh gu ni b hs qg nk nl hv qh nn no np qi nr ns nt qj nv nw nx qk nz oa ob gn bk">Creating a reliable source of truth for impressions is a complex but essential task that enhances personalization and discovery experience. Stay tuned for the next part of this series, where we’ll delve into how we use this SOT dataset to create a microservice that provides impression histories. We invite you to share your thoughts in the comments and continue with us on this journey of discovering impressions.</p><h1 id="22e6" class="ov ow gu bf ox oy oz hu pa pb pc hx pd pe pf pg ph pi pj pk pl pm pn po pp pq bk">Acknowledgments</h1><p id="0174" class="pw-post-body-paragraph ng nh gu ni b hs qg nk nl hv qh nn no np qi nr ns nt qj nv nw nx qk nz oa ob gn bk">We are genuinely grateful to our amazing colleagues whose contributions were essential to the success of Impressions: Julian Jaffe, Bryan Keller, Yun Wang, Brandon Bremen, Kyle Alford, Ron Brown and Shriya Arora.</p></div></div>]]></description>
      <link>https://netflixtechblog.com/introducing-impressions-at-netflix-e2b67c88c9fb</link>
      <guid>https://netflixtechblog.com/introducing-impressions-at-netflix-e2b67c88c9fb</guid>
      <pubDate>Sat, 15 Feb 2025 02:13:00 +0100</pubDate>
    </item>
    <item>
      <title><![CDATA[Title Launch Observability at Netflix Scale]]></title>
      <description><![CDATA[<div><div><h2 id="af30" class="pw-subtitle-paragraph hr gt gu bf b hs ht hu hv hw hx hy hz ia ib ic id ie if ig cq du">Part 2: Navigating Ambiguity</h2><div></div><p id="77d1" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk"><strong class="ni gv">By:</strong> <a class="af oc" href="https://www.linkedin.com/in/varun-khaitan/" rel="noopener ugc nofollow" target="_blank">Varun Khaitan</a></p><p id="72bf" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">With special thanks to my stunning colleagues: <a class="af oc" href="https://www.linkedin.com/in/mallikarao/" rel="noopener ugc nofollow" target="_blank">Mallika Rao</a>, <a class="af oc" href="https://www.linkedin.com/in/esmir-mesic/" rel="noopener ugc nofollow" target="_blank">Esmir Mesic</a>, <a class="af oc" href="https://www.linkedin.com/in/hugodesmarques/" rel="noopener ugc nofollow" target="_blank">Hugo Marques</a></p><p id="90f9" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">Building on the foundation laid in <a class="af oc" href="https://medium.com/netflix-techblog/title-launch-observability-at-netflix-scale-c88c586629eb" rel="noopener">Part 1</a>, where we explored the “what” behind the challenges of title launch observability at Netflix, this post shifts focus to the “how.” How do we ensure every title launches seamlessly and remains discoverable by the right audience?</p><p id="8c5c" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">In the dynamic world of technology, it’s tempting to leap into problem-solving mode. But the key to lasting success lies in taking a step back — understanding the broader context before diving into solutions. This thoughtful approach doesn’t just address immediate hurdles; it builds the resilience and scalability needed for the future. Let’s explore how this mindset drives results.</p><h1 id="afff" class="od oe gu bf of og oh hu oi oj ok hx ol om on oo op oq or os ot ou ov ow ox oy bk">Understanding the Bigger Picture</h1><p id="1ecd" class="pw-post-body-paragraph ng nh gu ni b hs oz nk nl hv pa nn no np pb nr ns nt pc nv nw nx pd nz oa ob gn bk">Let’s take a comprehensive look at all the elements involved and how they interconnect. We should aim to address questions such as: What is vital to the business? Which aspects of the problem are essential to resolve? And how did we arrive at this point?</p><p id="b0f8" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">This process involves:</p><ol class=""><li id="bdf3" class="ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob pe pf pg bk"><strong class="ni gv">Identifying Stakeholders: </strong>Determine who is impacted by the issue and whose input is crucial for a successful resolution. In this case, the main stakeholders are:<p>-<strong class="ni gv"><em class="ph"> Title Launch Operators<br />Role:</em></strong><em class="ph"> Responsible for setting up the title and its metadata into our systems.<br /></em><strong class="ni gv"><em class="ph">Challenge:</em></strong><em class="ph"> Don’t understand the cascading effects of their setup on these perceived black box personalization systems</em></p><p>-<strong class="ni gv"><em class="ph"> Personalization System Engineers</em></strong><em class="ph"><br /></em><strong class="ni gv"><em class="ph">Role: </em></strong><em class="ph">Develop and operate the personalization systems.<br /></em><strong class="ni gv"><em class="ph">Challenge:</em></strong><em class="ph"> End up spending unplanned cycles on title launch and personalization investigations.</em></p><p>- <strong class="ni gv"><em class="ph">Product Managers </em></strong><em class="ph"><br /></em><strong class="ni gv"><em class="ph">Role: </em></strong><em class="ph">Ensure we put forward the best experience for our members.<br /></em><strong class="ni gv"><em class="ph">Challenge: </em></strong><em class="ph">Members may not connect with the most relevant title.</em></p><p>- <strong class="ni gv"><em class="ph">Creative Representatives</em></strong><em class="ph"> <br /></em><strong class="ni gv"><em class="ph">Role:</em></strong><em class="ph"> Mediator between the content creators and Netflix.<br /></em><strong class="ni gv"><em class="ph">Challenge: </em></strong><em class="ph">Build trust in the Netflix brand with content creators.</em></p></li><li id="fb92" class="ng nh gu ni b hs pi nk nl hv pj nn no np pk nr ns nt pl nv nw nx pm nz oa ob pe pf pg bk"><strong class="ni gv">Mapping the Current Landscape:</strong> By charting the existing landscape, we can pinpoint areas ripe for improvement and steer clear of redundant efforts. Beyond the scattered solutions and makeshift scripts, it became evident that there was no established solution for title launch observability. This suggests that this area has been neglected for quite some time and likely requires significant investment. This situation presents both challenges and opportunities; while it may be more difficult to make initial progress, there are plenty of easy wins to capitalize on.</li><li id="d0ea" class="ng nh gu ni b hs pi nk nl hv pj nn no np pk nr ns nt pl nv nw nx pm nz oa ob pe pf pg bk"><strong class="ni gv">Clarifying the Core Problem:</strong> By clearly defining the problem, we can ensure that our solutions address the root cause rather than just the symptoms. While there were many issues and problems we could address, the core problem here was to make sure every title was treated fairly by our personalization stack. If we can ensure fair treatment with confidence and bring that visibility to all our stakeholders, we can address all their challenges.</li><li id="3881" class="ng nh gu ni b hs pi nk nl hv pj nn no np pk nr ns nt pl nv nw nx pm nz oa ob pe pf pg bk"><strong class="ni gv">Assessing Business Priorities: </strong>Understanding what is most important to the organization helps prioritize actions and resources effectively. In this context, we’re focused on developing systems that ensure successful title launches, build trust between content creators and our brand, and reduce engineering operational overhead. While this is a critical business need and we definitely should solve it, it’s essential to evaluate how it stacks up against other priorities across different areas of the organization.</li></ol><h1 id="4208" class="od oe gu bf of og oh hu oi oj ok hx ol om on oo op oq or os ot ou ov ow ox oy bk">Defining Title Health</h1><p id="2dfd" class="pw-post-body-paragraph ng nh gu ni b hs oz nk nl hv pa nn no np pb nr ns nt pc nv nw nx pd nz oa ob gn bk">Navigating such an ambiguous space required a shared understanding to foster clarity and collaboration. To address this, we introduced the term “Title Health,” a concept designed to help us communicate effectively and capture the nuances of maintaining each title’s visibility and performance. This shared language became a foundation for discussing the complexities of this domain.</p><p id="70a0" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk"><strong class="ni gv">“Title Health”</strong> encompasses various metrics and indicators that reflect how well a title is performing, in terms of discoverability and member engagement. The three main questions we try to answer are:</p><ol class=""><li id="7e75" class="ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob pe pf pg bk">Is this title visible at all to <strong class="ni gv">any</strong> <strong class="ni gv">member</strong>?</li><li id="777c" class="ng nh gu ni b hs pi nk nl hv pj nn no np pk nr ns nt pl nv nw nx pm nz oa ob pe pf pg bk">Is this title visible to an appropriate <strong class="ni gv">audience size</strong>?</li><li id="0cf3" class="ng nh gu ni b hs pi nk nl hv pj nn no np pk nr ns nt pl nv nw nx pm nz oa ob pe pf pg bk">Is this title reaching <strong class="ni gv">all the appropriate audiences</strong>?</li></ol><p id="b2dd" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">Defining Title Health provided a framework to monitor and optimize each title’s lifecycle. It allowed us to align with partners on principles and requirements before building solutions, ensuring every title reaches its intended audience seamlessly. This common language not only introduced the problem space effectively but also accelerated collaboration and decision-making across teams.</p><h1 id="b793" class="od oe gu bf of og oh hu oi oj ok hx ol om on oo op oq or os ot ou ov ow ox oy bk">Categories of issues</h1><p id="0431" class="pw-post-body-paragraph ng nh gu ni b hs oz nk nl hv pa nn no np pb nr ns nt pc nv nw nx pd nz oa ob gn bk">To build a robust plan for title launch observability, we first needed to categorize the types of issues we encounter. This structured approach allows us to address all aspects of title health comprehensively.</p><p id="86be" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">Currently, these issues are grouped into three primary categories:</p><p id="62b5" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk"><strong class="ni gv">1. Title Setup</strong></p><p id="1422" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">A title’s setup includes essential attributes like metadata (e.g., launch dates, audio and subtitle languages, editorial tags) and assets (e.g., artwork, trailers, supplemental messages). These elements are critical for a title’s eligibility in a row, accurate personalization, and an engaging presentation. Since these attributes feed directly into algorithms, any delays or inaccuracies can ripple through the system.</p><p id="c1fa" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">The observability system must ensure that title setup is complete and validated in a timely manner, identify potential bottlenecks and ensure a smooth launch process.</p><p id="551b" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk"><strong class="ni gv">2. Personalization Systems</strong></p><p id="4456" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">Titles are eligible to be recommended across multiple canvases on product — HomePage, Coming Soon, Messaging, Search and more. Personalization systems handle the recommendation and serving of titles on these canvases, leveraging a vast ecosystem of microservices, caches, databases, code, and configurations to build these product canvases.</p><p id="d32b" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">We aim to validate that titles are eligible in all appropriate product canvases across the end to end personalization stack during all of the title’s launch phases.</p><p id="928f" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk"><strong class="ni gv">3. Algorithms</strong></p><p id="3751" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">Complex algorithms drive each personalized product experience, recommending titles tailored to individual members. Observability here means validating the accuracy of algorithmic recommendations for all titles.<br />Algorithmic performance can be affected by various factors, such as model shortcomings, incomplete or inaccurate input signals, feature anomalies, or interactions between titles. Identifying and addressing these issues ensures that recommendations remain precise and effective.</p><p id="a466" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">By categorizing issues into these areas, we can systematically address challenges and deliver a reliable, personalized experience for every title on our platform.</p><h1 id="7b81" class="od oe gu bf of og oh hu oi oj ok hx ol om on oo op oq or os ot ou ov ow ox oy bk">Issue Analysis</h1><p id="a29f" class="pw-post-body-paragraph ng nh gu ni b hs oz nk nl hv pa nn no np pb nr ns nt pc nv nw nx pd nz oa ob gn bk">Let’s also learn more about how often we see each of these types of issues and how much effort it takes to fix them once they come up.</p><figure class="pq pr ps pt pu pv pn po paragraph-image"><div role="button" tabindex="0" class="pw px fj py bh pz"><div class="pn po pp"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*YyCLwVKiGE_L6fWb%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*YyCLwVKiGE_L6fWb%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*YyCLwVKiGE_L6fWb%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*YyCLwVKiGE_L6fWb%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*YyCLwVKiGE_L6fWb%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*YyCLwVKiGE_L6fWb%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*YyCLwVKiGE_L6fWb%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*YyCLwVKiGE_L6fWb 640w, https://miro.medium.com/v2/resize:fit:720/0*YyCLwVKiGE_L6fWb 720w, https://miro.medium.com/v2/resize:fit:750/0*YyCLwVKiGE_L6fWb 750w, https://miro.medium.com/v2/resize:fit:786/0*YyCLwVKiGE_L6fWb 786w, https://miro.medium.com/v2/resize:fit:828/0*YyCLwVKiGE_L6fWb 828w, https://miro.medium.com/v2/resize:fit:1100/0*YyCLwVKiGE_L6fWb 1100w, https://miro.medium.com/v2/resize:fit:1400/0*YyCLwVKiGE_L6fWb 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh mn qa c" width="700" height="619" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="030a" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">From the above chart, we see that setup issues are the most common but they are also easy to fix since it’s relatively straightforward to go back and rectify a title’s metadata. System issues, which mostly manifest as bugs in our personalization microservices are not uncommon, and they take moderate effort to address. Algorithm issues, while rare, are really difficult to address since these often involve interpreting and retraining complex machine learning models.</p><h1 id="c0e1" class="od oe gu bf of og oh hu oi oj ok hx ol om on oo op oq or os ot ou ov ow ox oy bk">Evaluating Our Options</h1><p id="473f" class="pw-post-body-paragraph ng nh gu ni b hs oz nk nl hv pa nn no np pb nr ns nt pc nv nw nx pd nz oa ob gn bk">Now that we understand more deeply about the problems we want to address and how we should go about prioritizing our resources. Lets go back to the two options we discussed in Part 1, and make an informed decision.</p><figure class="pq pr ps pt pu pv pn po paragraph-image"><div role="button" tabindex="0" class="pw px fj py bh pz"><div class="pn po qb"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*5YRxJT3YI53wgtLs9zO6gg.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*5YRxJT3YI53wgtLs9zO6gg.png" /><img alt="" class="bh mn qa c" width="700" height="263" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="64d2" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">Ultimately, we realized this space demands the full spectrum of features we’ve discussed. But the question remained: <em class="ph">Where do we start?</em> <br />After careful consideration, we chose to focus on proactive issue detection first. Catching problems before launch offered the greatest potential for business impact, ensuring smoother launches, better member experiences, and stronger system reliability.</p><p id="84b9" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">This decision wasn’t just about solving today’s challenges — it was about laying the foundation for a scalable, robust system that can grow with the complexities of our ever-evolving platform.</p><h1 id="ab95" class="od oe gu bf of og oh hu oi oj ok hx ol om on oo op oq or os ot ou ov ow ox oy bk">Up next</h1><p id="f46c" class="pw-post-body-paragraph ng nh gu ni b hs oz nk nl hv pa nn no np pb nr ns nt pc nv nw nx pd nz oa ob gn bk">In the next iteration we will talk about how to design an observability endpoint that works for all personalization systems. What are the main things to keep in mind while creating a microservice API endpoint? How do we ensure standardization? What is the architecture of the systems involved?</p><p id="a208" class="pw-post-body-paragraph ng nh gu ni b hs nj nk nl hv nm nn no np nq nr ns nt nu nv nw nx ny nz oa ob gn bk">Keep an eye out for our next binge-worthy episode!</p></div></div>]]></description>
      <link>https://netflixtechblog.com/title-launch-observability-at-netflix-scale-19ea916be1ed</link>
      <guid>https://netflixtechblog.com/title-launch-observability-at-netflix-scale-19ea916be1ed</guid>
      <pubDate>Tue, 07 Jan 2025 02:25:00 +0100</pubDate>
    </item>
    <item>
      <title><![CDATA[Part 3: A Survey of Analytics Engineering Work at Netflix]]></title>
      <description><![CDATA[<div class="gn go gp gq gr"><div class="ab cb"><div class="ci bh fz ga gb gc"><div><div></div><p id="ca1c" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><em class="nu">This article is the last in a multi-part series sharing a breadth of Analytics Engineering work at Netflix, recently presented as part of our annual internal Analytics Engineering conference. Need to catch up? Check out </em><a class="af nv" href="https://research.netflix.com/publication/part-1-a-survey-of-analytics-engineering-work-at-netflix" rel="noopener ugc nofollow" target="_blank"><em class="nu">Part 1</em></a><em class="nu">, which detailed how we’re empowering Netflix to efficiently produce and effectively deliver high quality, actionable analytic insights across the company and </em><a class="af nv" href="https://research.netflix.com/publication/part-2-a-survey-of-analytics-engineering-work-at-netflix" rel="noopener ugc nofollow" target="_blank"><em class="nu">Part 2</em></a><em class="nu">, which stepped through a few exciting business applications for Analytics Engineering. This post will go into aspects of technical craft.</em></p><h1 id="a095" class="nw nx gu bf ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os ot bk">Dashboard Design Tips</h1><p id="1817" class="pw-post-body-paragraph mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt gn bk"><a class="af nv" href="https://www.linkedin.com/in/rinachang" rel="noopener ugc nofollow" target="_blank">Rina Chang</a>, <a class="af nv" href="https://www.linkedin.com/in/shansusielu/" rel="noopener ugc nofollow" target="_blank">Susie Lu</a></p><p id="4102" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">What is design, and why does it matter? Often people think design is about how things look, but design is actually about how things work. Everything is designed, because we’re all making choices about how things work, but not everything is designed well. Good design doesn’t waste time or mental energy; instead, it helps the user achieve their goals.</p><p id="59f0" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">When applying this to a dashboard application, the easiest way to use design effectively is to leverage existing patterns. (For example, people have learned that blue underlined text on a website means it’s a clickable link.) So knowing the arsenal of available patterns and what they imply is useful when making the choice of when to use which pattern.</p><p id="692c" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">First, to design a dashboard well, you need to understand your user.</p><ul class=""><li id="e325" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt oz pa pb bk">Talk to your users throughout the entire product lifecycle. Talk to them early and often, through whatever means you can.</li><li id="12e2" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk">Understand their needs, ask why, then ask why again. Separate symptoms from problems from solutions.</li><li id="4374" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk">Prioritize and clarify — less is more! Distill what you can build that’s differentiated and provides the most value to your user.</li></ul><p id="dd4d" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Here is a framework for thinking about what your users are trying to achieve. Where do your users fall on these axes? Don’t solve for multiple positions across these axes in a given view; if that exists, then create different views or potentially different dashboards.</p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi pj"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*ar0t2-zF5YVuXnUe%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*ar0t2-zF5YVuXnUe%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*ar0t2-zF5YVuXnUe%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*ar0t2-zF5YVuXnUe%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*ar0t2-zF5YVuXnUe%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*ar0t2-zF5YVuXnUe%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*ar0t2-zF5YVuXnUe%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*ar0t2-zF5YVuXnUe 640w, https://miro.medium.com/v2/resize:fit:720/0*ar0t2-zF5YVuXnUe 720w, https://miro.medium.com/v2/resize:fit:750/0*ar0t2-zF5YVuXnUe 750w, https://miro.medium.com/v2/resize:fit:786/0*ar0t2-zF5YVuXnUe 786w, https://miro.medium.com/v2/resize:fit:828/0*ar0t2-zF5YVuXnUe 828w, https://miro.medium.com/v2/resize:fit:1100/0*ar0t2-zF5YVuXnUe 1100w, https://miro.medium.com/v2/resize:fit:1400/0*ar0t2-zF5YVuXnUe 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="370" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="106c" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Second, understanding your users’ mental models will allow you to choose how to structure your app to match. A few questions to ask yourself when considering the information architecture of your app include:</p><ul class=""><li id="30c3" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt oz pa pb bk">Do you have different user groups trying to accomplish different things? Split them into different apps or different views.</li><li id="feea" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk">What should go together on a single page? All the information needed for a single user type to accomplish their “job.” If there are multiple <a class="af nv" href="https://www.christenseninstitute.org/theory/jobs-to-be-done/" rel="noopener ugc nofollow" target="_blank">jobs to be done</a>, split each out onto its own page.</li><li id="c406" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk">What should go together within a single section on a page? All the information needed to answer a single question.</li><li id="3bd1" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk">Does your dashboard feel too difficult to use? You probably have too much information! When in doubt, keep it simple. If needed, hide complexity under an “Advanced” section.</li></ul><p id="0289" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Here are some general guidelines for page layouts:</p><ul class=""><li id="935c" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt oz pa pb bk">Choose infinite scrolling vs. clicking through multiple pages depending on which option suits your users’ expectations better</li><li id="d91d" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk">Lead with the most-used information first, above the fold</li><li id="a5c0" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk">Create signposts that cue the user to where they are by labeling pages, sections, and links</li><li id="7543" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk">Use cards or borders to visually group related items together</li><li id="94e9" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk">Leverage nesting to create well-understood “scopes of control.” Specifically, users expect a controller object to affect children either: Below it (if horizontal) or To the right of it (if vertical)</li></ul><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi pv"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*KIqd6dZXD_NZyTKR%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*KIqd6dZXD_NZyTKR%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*KIqd6dZXD_NZyTKR%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*KIqd6dZXD_NZyTKR%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*KIqd6dZXD_NZyTKR%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*KIqd6dZXD_NZyTKR%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*KIqd6dZXD_NZyTKR%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*KIqd6dZXD_NZyTKR 640w, https://miro.medium.com/v2/resize:fit:720/0*KIqd6dZXD_NZyTKR 720w, https://miro.medium.com/v2/resize:fit:750/0*KIqd6dZXD_NZyTKR 750w, https://miro.medium.com/v2/resize:fit:786/0*KIqd6dZXD_NZyTKR 786w, https://miro.medium.com/v2/resize:fit:828/0*KIqd6dZXD_NZyTKR 828w, https://miro.medium.com/v2/resize:fit:1100/0*KIqd6dZXD_NZyTKR 1100w, https://miro.medium.com/v2/resize:fit:1400/0*KIqd6dZXD_NZyTKR 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="977" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi pv"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*O52xqUnDsJ8kPCVZ%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*O52xqUnDsJ8kPCVZ%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*O52xqUnDsJ8kPCVZ%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*O52xqUnDsJ8kPCVZ%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*O52xqUnDsJ8kPCVZ%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*O52xqUnDsJ8kPCVZ%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*O52xqUnDsJ8kPCVZ%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*O52xqUnDsJ8kPCVZ 640w, https://miro.medium.com/v2/resize:fit:720/0*O52xqUnDsJ8kPCVZ 720w, https://miro.medium.com/v2/resize:fit:750/0*O52xqUnDsJ8kPCVZ 750w, https://miro.medium.com/v2/resize:fit:786/0*O52xqUnDsJ8kPCVZ 786w, https://miro.medium.com/v2/resize:fit:828/0*O52xqUnDsJ8kPCVZ 828w, https://miro.medium.com/v2/resize:fit:1100/0*O52xqUnDsJ8kPCVZ 1100w, https://miro.medium.com/v2/resize:fit:1400/0*O52xqUnDsJ8kPCVZ 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="977" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi pv"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*6qCQlTNyoabhrkVa%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*6qCQlTNyoabhrkVa%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*6qCQlTNyoabhrkVa%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*6qCQlTNyoabhrkVa%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*6qCQlTNyoabhrkVa%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*6qCQlTNyoabhrkVa%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*6qCQlTNyoabhrkVa%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*6qCQlTNyoabhrkVa 640w, https://miro.medium.com/v2/resize:fit:720/0*6qCQlTNyoabhrkVa 720w, https://miro.medium.com/v2/resize:fit:750/0*6qCQlTNyoabhrkVa 750w, https://miro.medium.com/v2/resize:fit:786/0*6qCQlTNyoabhrkVa 786w, https://miro.medium.com/v2/resize:fit:828/0*6qCQlTNyoabhrkVa 828w, https://miro.medium.com/v2/resize:fit:1100/0*6qCQlTNyoabhrkVa 1100w, https://miro.medium.com/v2/resize:fit:1400/0*6qCQlTNyoabhrkVa 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="977" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="8d9f" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Third, some tips and tricks can help you more easily tackle the unique design challenges that come with making interactive charts.</p><ul class=""><li id="9697" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt oz pa pb bk">Titles: Make sure filters are represented in the title or subtitle of the chart for easy scannability and screenshot-ability.</li><li id="c959" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk">Tooltips: Core details should be on the page, while the context in the tooltip is for deeper information. Annotate multiple points when there are only a handful of lines.</li><li id="c00c" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk">Annotations: Provide annotations on charts to explain shifts in values so all users can access that context.</li><li id="4445" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk">Color: Limit the number of colors you use. Be consistent in how you use colors. Otherwise, colors lose meaning.</li><li id="fe99" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk">Onboarding: Separate out onboarding to your dashboard from routine usage.</li></ul><p id="a37f" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Finally, it is important to note that these are general guidelines, but there is always room for interpretation and/or the use of good judgment to adapt them to suit your own product and use cases. At the end of the day, the most important thing is that a user can leverage the data insights provided by your dashboard to perform their work, and good design is a means to that end.</p><h1 id="f4f7" class="nw nx gu bf ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os ot bk"><strong class="al">Learnings from Deploying an Analytics API at Netflix</strong></h1><p id="31ad" class="pw-post-body-paragraph mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt gn bk"><a class="af nv" href="https://www.linkedin.com/in/devincarullo/" rel="noopener ugc nofollow" target="_blank">Devin Carullo</a></p><p id="430b" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">At Netflix Studio, we operate at the intersection of art and science. Data is a tool that enhances decision-making, complementing the deep expertise and industry knowledge of our creative professionals.</p><p id="970d" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">One example is in production budgeting — namely, determining how much we should spend to produce a given show or movie. Although there was already a process for creating and comparing budgets for new productions against similar past projects, it was highly manual. We developed a tool that automatically selects and compares similar Netflix productions, flagging any anomalies for Production Finance to review.</p><p id="b3ea" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">To ensure success, it was essential that results be delivered in real-time and integrated seamlessly into existing tools. This required close collaboration among product teams, DSE, and front-end and back-end developers. We developed a GraphQL endpoint using Metaflow, integrating it into the existing budgeting product. This solution enabled data to be used more effectively for real-time decision-making.</p><p id="d3db" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">We recently launched our MVP and continue to iterate on the product. Reflecting on our journey, the path to launch was complex and filled with unexpected challenges. As an analytics engineer accustomed to crafting quick solutions, I underestimated the effort required to deploy a production-grade analytics API.</p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi pj"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*KOgCUre0HvjZ82ZH%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*KOgCUre0HvjZ82ZH%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*KOgCUre0HvjZ82ZH%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*KOgCUre0HvjZ82ZH%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*KOgCUre0HvjZ82ZH%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*KOgCUre0HvjZ82ZH%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*KOgCUre0HvjZ82ZH%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*KOgCUre0HvjZ82ZH 640w, https://miro.medium.com/v2/resize:fit:720/0*KOgCUre0HvjZ82ZH 720w, https://miro.medium.com/v2/resize:fit:750/0*KOgCUre0HvjZ82ZH 750w, https://miro.medium.com/v2/resize:fit:786/0*KOgCUre0HvjZ82ZH 786w, https://miro.medium.com/v2/resize:fit:828/0*KOgCUre0HvjZ82ZH 828w, https://miro.medium.com/v2/resize:fit:1100/0*KOgCUre0HvjZ82ZH 1100w, https://miro.medium.com/v2/resize:fit:1400/0*KOgCUre0HvjZ82ZH 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="86" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="pw ff px ph pi py pz bf b bg z du">Fig 1. My vague idea of how my API would work</figcaption></figure><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi pj"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*BBEHaQdU_e57_sjD%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*BBEHaQdU_e57_sjD%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*BBEHaQdU_e57_sjD%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*BBEHaQdU_e57_sjD%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*BBEHaQdU_e57_sjD%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*BBEHaQdU_e57_sjD%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*BBEHaQdU_e57_sjD%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*BBEHaQdU_e57_sjD 640w, https://miro.medium.com/v2/resize:fit:720/0*BBEHaQdU_e57_sjD 720w, https://miro.medium.com/v2/resize:fit:750/0*BBEHaQdU_e57_sjD 750w, https://miro.medium.com/v2/resize:fit:786/0*BBEHaQdU_e57_sjD 786w, https://miro.medium.com/v2/resize:fit:828/0*BBEHaQdU_e57_sjD 828w, https://miro.medium.com/v2/resize:fit:1100/0*BBEHaQdU_e57_sjD 1100w, https://miro.medium.com/v2/resize:fit:1400/0*BBEHaQdU_e57_sjD 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="392" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="pw ff px ph pi py pz bf b bg z du">Fig 2: Our actual solution</figcaption></figure><p id="3a82" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">With hindsight, below are my key learnings.</p><p id="3ed3" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Measure Impact and Necessity of Real-Time Results</strong></p><p id="8dac" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Before implementing real-time analytics, assess whether real-time results are truly necessary for your use case. This can significantly impact the complexity and cost of your solution. Batch processing data may provide a similar impact and take significantly less time. It’s easier to develop and maintain, and tends to be more familiar for analytics engineers, data scientists, and data engineers.</p><p id="70ae" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Additionally, if you are developing a proof of concept, the upfront investment may not be worth it. Scrappy solutions can often be the best choice for analytics work.</p><p id="1905" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Explore All Available Solutions</strong></p><p id="fc32" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">At Netflix, there were multiple established methods for creating an API, but none perfectly suited our specific use case. Metaflow, a tool developed at Netflix for data science projects, already supported REST APIs. However, this approach did not align with the preferred workflow of our engineering partners. Although they could integrate with REST endpoints, this solution presented inherent limitations. Large response sizes rendered the API/front-end integration unreliable, necessitating the addition of filter parameters to reduce the response size.</p><p id="ccc2" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Additionally, the product we were integrating into was using GraphQL, and deviating from this established engineering approach was not ideal. Lastly, given our goal to overlay results throughout the product, GraphQL features, such as federation, proved to be particularly advantageous.</p><p id="7ffd" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">After realizing there wasn’t an existing solution at Netflix for deploying python endpoints with GraphQL, we worked with the Metaflow team to build this feature. This allowed us to continue developing via Metaflow and allowed our engineering partners to stay on their paved path.</p><p id="c621" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Align on Performance Expectations</strong></p><p id="d3b5" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">A major challenge during development was managing API latency. Much of this could have been mitigated by aligning on performance expectations from the outset. Initially, we operated under our assumptions of what constituted an acceptable response time, which differed greatly from the actual needs of our users and our engineering partners.</p><p id="2ae7" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Understanding user expectations is key to designing an effective solution. Our methodology resulted in a full budget analysis taking, on average, 7 seconds. Users were willing to wait for an analysis when they modified a budget, but not every time they accessed one. To address this, we implemented caching using Metaflow, reducing the API response time to approximately 1 second for cached results. Additionally, we set up a nightly batch job to pre-cache results.</p><p id="7495" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">While users were generally okay with waiting for analysis during changes, we had to be mindful of GraphQL’s 30-second limit. This highlighted the importance of continuously monitoring the impact of changes on response times, leading us to our next key learning: rigorous testing.</p><p id="1924" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Real-Time Analysis Requires Rigorous Testing</strong></p><p id="ca7a" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Load Testing: We leveraged Locust to measure the response time of our endpoint and assess how the endpoint responded to reasonable and elevated loads. We were able to use FullStory, which was already being used in the product, to estimate expected calls per minute.</p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi qa"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*xVjhqU2DZV7RYBD0%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*xVjhqU2DZV7RYBD0%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*xVjhqU2DZV7RYBD0%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*xVjhqU2DZV7RYBD0%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*xVjhqU2DZV7RYBD0%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*xVjhqU2DZV7RYBD0%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*xVjhqU2DZV7RYBD0%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*xVjhqU2DZV7RYBD0 640w, https://miro.medium.com/v2/resize:fit:720/0*xVjhqU2DZV7RYBD0 720w, https://miro.medium.com/v2/resize:fit:750/0*xVjhqU2DZV7RYBD0 750w, https://miro.medium.com/v2/resize:fit:786/0*xVjhqU2DZV7RYBD0 786w, https://miro.medium.com/v2/resize:fit:828/0*xVjhqU2DZV7RYBD0 828w, https://miro.medium.com/v2/resize:fit:1100/0*xVjhqU2DZV7RYBD0 1100w, https://miro.medium.com/v2/resize:fit:1400/0*xVjhqU2DZV7RYBD0 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="260" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="pw ff px ph pi py pz bf b bg z du">Fig 3. Locust allows us to simulate concurrent calls and measure response time</figcaption></figure><p id="4782" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Unit Tests &amp; Integration Tests: Code testing is always a good idea, but it can often be overlooked in analytics. It is especially important when you are delivering live analysis to circumvent end users from being the first to see an error or incorrect information. We implemented unit testing and full integration tests, ensuring that our analysis would return correct results.</p><p id="91e6" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">The Importance of Aligning Workflows and Collaboration</strong></p><p id="3443" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">This project marked the first time our team collaborated directly with our engineering partners to integrate a DSE API into their product. Throughout the process, we discovered significant gaps in our understanding of each other’s workflows. Assumptions about each other’s knowledge and processes led to misunderstandings and delays.</p><p id="3d94" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Deployment Paths: Our engineering partners followed a strict deployment path, whereas our approach on the DSE side was more flexible. We typically tested our work on feature branches using Metaflow projects and then pushed results to production. However, this lack of control led to issues, such as inadvertently deploying changes to production before the corresponding product updates were ready and difficulties in managing a test endpoint. Ultimately, we deferred to our engineering partners to establish a deployment path and collaborated with the Metaflow team and data engineers to implement it effectively.</p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi pj"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*BaqggE2wQ2C9Svo8%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*BaqggE2wQ2C9Svo8%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*BaqggE2wQ2C9Svo8%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*BaqggE2wQ2C9Svo8%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*BaqggE2wQ2C9Svo8%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*BaqggE2wQ2C9Svo8%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*BaqggE2wQ2C9Svo8%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*BaqggE2wQ2C9Svo8 640w, https://miro.medium.com/v2/resize:fit:720/0*BaqggE2wQ2C9Svo8 720w, https://miro.medium.com/v2/resize:fit:750/0*BaqggE2wQ2C9Svo8 750w, https://miro.medium.com/v2/resize:fit:786/0*BaqggE2wQ2C9Svo8 786w, https://miro.medium.com/v2/resize:fit:828/0*BaqggE2wQ2C9Svo8 828w, https://miro.medium.com/v2/resize:fit:1100/0*BaqggE2wQ2C9Svo8 1100w, https://miro.medium.com/v2/resize:fit:1400/0*BaqggE2wQ2C9Svo8 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="328" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="pw ff px ph pi py pz bf b bg z du">Fig 4. Our current deployment path</figcaption></figure><p id="6fa7" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Work Planning: While the engineering team operated on sprints, our DSE team planned by quarters. This misalignment in planning cycles is an ongoing challenge that we are actively working to resolve.</p><p id="e86e" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Looking ahead, our team is committed to continuing this partnership with our engineering colleagues. Both teams have invested significant time in building this relationship, and we are optimistic that it will yield substantial benefits in future projects.</p><h1 id="f327" class="nw nx gu bf ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os ot bk">External Speaker: Benn Stancil</h1><p id="fa9b" class="pw-post-body-paragraph mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt gn bk">In addition to the above presentations, we kicked off our Analytics Summit with a keynote talk from <a class="af nv" href="https://www.linkedin.com/in/benn-stancil/" rel="noopener ugc nofollow" target="_blank">Benn Stancil</a>, Founder of Mode Analytics. Benn stepped through a history of the modern data stack, and the group discussed ideas on the future of analytics.</p></div></div></div><div class="ab cb qb qc qd qe" role="separator"><div class="gn go gp gq gr"><div class="ab cb"><div class="ci bh fz ga gb gc"><p id="5e52" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Analytics Engineering is a key contributor to building our deep data culture at Netflix, and we are proud to have a large group of stunning colleagues that are not only applying but advancing our analytical capabilities at Netflix. The 2024 Analytics Summit continued to be a wonderful way to give visibility to one another on work across business verticals, celebrate our collective impact, and highlight what’s to come in analytics practice at Netflix.</p><p id="1920" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">To learn more, follow the <a class="af nv" href="https://research.netflix.com/research-area/analytics" rel="noopener ugc nofollow" target="_blank">Netflix Research Site</a>, and if you are also interested in entertaining the world, have a look at <a class="af nv" href="https://explore.jobs.netflix.net/careers" rel="noopener ugc nofollow" target="_blank">our open roles</a>!</p></div></div></div></div></div>]]></description>
      <link>https://netflixtechblog.com/part-3-a-survey-of-analytics-engineering-work-at-netflix-e67f0aa82183</link>
      <guid>https://netflixtechblog.com/part-3-a-survey-of-analytics-engineering-work-at-netflix-e67f0aa82183</guid>
      <pubDate>Mon, 06 Jan 2025 20:27:00 +0100</pubDate>
    </item>
    <item>
      <title><![CDATA[Part 2: A Survey of Analytics Engineering Work at Netflix]]></title>
      <description><![CDATA[<div class="gn go gp gq gr"><div class="ab cb"><div class="ci bh fz ga gb gc"><div><div></div><p id="c218" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><em class="nu">This article is the second in a multi-part series sharing a breadth of Analytics Engineering work at Netflix, recently presented as part of our annual internal Analytics Engineering conference. Need to catch up? Check out </em><a class="af nv" href="https://research.netflix.com/publication/part-1-a-survey-of-analytics-engineering-work-at-netflix" rel="noopener ugc nofollow" target="_blank"><em class="nu">Part 1</em></a><em class="nu">. In this article, we highlight a few exciting analytic business applications, and in our final article we’ll go into aspects of the technical craft.</em></p><h1 id="6e5b" class="nw nx gu bf ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os ot bk">Game Analytics</h1><p id="cf63" class="pw-post-body-paragraph mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt gn bk"><a class="af nv" href="https://www.linkedin.com/in/yimeng-tang-49566b207/" rel="noopener ugc nofollow" target="_blank">Yimeng Tang</a>, <a class="af nv" href="https://www.linkedin.com/in/clairewilleck/" rel="noopener ugc nofollow" target="_blank">Claire Willeck</a>, <a class="af nv" href="https://www.linkedin.com/in/sagarpalao/" rel="noopener ugc nofollow" target="_blank">Sagar Palao</a></p><h1 id="3399" class="nw nx gu bf ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os ot bk">User Acquisition Incrementality for Netflix Games</h1><p id="af46" class="pw-post-body-paragraph mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt gn bk">Netflix has been launching games for the past three years, during which it has initiated various marketing efforts, including User Acquisition (UA) campaigns, to promote these games across different countries. These UA campaigns typically feature static creatives, launch trailers, and game review videos on platforms like Google, Meta, and TikTok. The primary goals of these campaigns are to encourage more people to install and play the games, making incremental installs and engagement crucial metrics for evaluating their effectiveness.</p><p id="8619" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Most UA campaigns are conducted at the country level, meaning that everyone in the targeted countries can see the ads. However, due to the absence of a control group in these countries, we adopt a synthetic control framework (<a class="af nv" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/round-2-a-survey-of-causal-inference-applications-at-netflix-fd78328ee0bb">blog post</a>) to estimate the counterfactual scenario. This involves creating a weighted combination of countries not exposed to the UA campaign to serve as a counterfactual for the treated countries. To facilitate easier access to incrementality results, we have developed an interactive tool powered by this framework. This tool allows users to directly obtain the lift in game installs and engagement, view plots for both the treated country and the synthetic control unit, and assess the p-value from placebo tests.</p><p id="b9e6" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">To better guide the design and budgeting of future campaigns, we are developing an Incremental Return on Investment model. This model incorporates factors such as the incremental impact, the value of the incremental engagement and incremental signups, and the cost of running the campaign. In addition to using the causal inference framework mentioned earlier to estimate incrementality, we also leverage other frameworks, such as Incremental Account Lifetime Valuation (<a class="af nv" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/a-survey-of-causal-inference-applications-at-netflix-b62d25175e6f">blog post</a>), to assign value to the incremental engagement and signups resulting from the campaigns.</p><h1 id="bd39" class="nw nx gu bf ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os ot bk">Measuring and Validating Incremental Signups for Netflix Games</h1><p id="9ea5" class="pw-post-body-paragraph mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt gn bk">Netflix is a subscription service meaning members buy subscriptions which include games but not the individual games themselves. This makes it difficult to measure the impact of different game launches on acquisition. We only observe signups, not why members signed up.</p><p id="c843" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">This means we need to estimate incremental signups. We adopt an approach developed at Netflix to estimate incremental acquisition (<a class="af nv" href="https://arxiv.org/pdf/2106.15346" rel="noopener ugc nofollow" target="_blank">technical paper</a>). This approach uses simple assumptions to estimate a counterfactual for the rate that new members start playing the game.</p><p id="5128" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Because games differ from series/films, it’s crucial to validate this estimation method for games. Ideally, we would have causal estimates from an A/B test to use for validation, but since that is not available, we use another causal inference design as one of our ensemble of validation approaches. This causal inference design involves a systematic framework we designed to measure game events that relies on synthetic control (<a class="af nv" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/round-2-a-survey-of-causal-inference-applications-at-netflix-fd78328ee0bb">blog post</a>).</p><p id="19cb" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">As we mentioned above, we have been launching User Acquisition (UA) campaigns in select countries to boost game engagement and new memberships. We can use this cross-country variation to form a synthetic control and measure the incremental signups due to the UA campaign. The incremental signups from UA campaigns differ from those attributed to a game, but they should be similar. When our estimated incremental acquisition numbers over a campaign period are similar to the incremental acquisition numbers calculated using synthetic control, we feel more confident in our approach to measuring incremental signups for games.</p><h1 id="1015" class="nw nx gu bf ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os ot bk">Netflix Games Players’ Adventure: Modeled using State Machine</h1><p id="828b" class="pw-post-body-paragraph mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt gn bk">At Netflix Games, we aim to have a high number of members engaging with games each month, referred to as Monthly Active Accounts (MAA). To evaluate our progress toward this objective and to find areas to boost our MAA, we modeled the Netflix players’ journey as a state machine.</p><p id="a279" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">We track a daily state machine showing the probability of account transitions between states.</p><figure class="pc pd pe pf pg ph oz pa paragraph-image"><div role="button" tabindex="0" class="pi pj fj pk bh pl"><div class="oz pa pb"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*j2wKL4S3ywEs9mpf%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*j2wKL4S3ywEs9mpf%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*j2wKL4S3ywEs9mpf%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*j2wKL4S3ywEs9mpf%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*j2wKL4S3ywEs9mpf%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*j2wKL4S3ywEs9mpf%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*j2wKL4S3ywEs9mpf%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*j2wKL4S3ywEs9mpf 640w, https://miro.medium.com/v2/resize:fit:720/0*j2wKL4S3ywEs9mpf 720w, https://miro.medium.com/v2/resize:fit:750/0*j2wKL4S3ywEs9mpf 750w, https://miro.medium.com/v2/resize:fit:786/0*j2wKL4S3ywEs9mpf 786w, https://miro.medium.com/v2/resize:fit:828/0*j2wKL4S3ywEs9mpf 828w, https://miro.medium.com/v2/resize:fit:1100/0*j2wKL4S3ywEs9mpf 1100w, https://miro.medium.com/v2/resize:fit:1400/0*j2wKL4S3ywEs9mpf 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pm c" width="700" height="580" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="5504" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Fig: Netflix Players’ Journey as State machine</p><p id="82b9" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Modeling the players’ journey as a state machine allows us to simulate future states and assess progress toward engagement goals. The most basic operation involves multiplying the daily state-transition matrix with the current state values to determine the next day’s state values.</p><figure class="pc pd pe pf pg ph oz pa paragraph-image"><div role="button" tabindex="0" class="pi pj fj pk bh pl"><div class="oz pa pn"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*ud5xnQi9QM6ELiVP%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*ud5xnQi9QM6ELiVP%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*ud5xnQi9QM6ELiVP%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*ud5xnQi9QM6ELiVP%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*ud5xnQi9QM6ELiVP%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*ud5xnQi9QM6ELiVP%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*ud5xnQi9QM6ELiVP%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*ud5xnQi9QM6ELiVP 640w, https://miro.medium.com/v2/resize:fit:720/0*ud5xnQi9QM6ELiVP 720w, https://miro.medium.com/v2/resize:fit:750/0*ud5xnQi9QM6ELiVP 750w, https://miro.medium.com/v2/resize:fit:786/0*ud5xnQi9QM6ELiVP 786w, https://miro.medium.com/v2/resize:fit:828/0*ud5xnQi9QM6ELiVP 828w, https://miro.medium.com/v2/resize:fit:1100/0*ud5xnQi9QM6ELiVP 1100w, https://miro.medium.com/v2/resize:fit:1400/0*ud5xnQi9QM6ELiVP 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pm c" width="700" height="73" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="0d9c" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">This basic operation allows us to explore various scenarios:</p><ul class=""><li id="03b5" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt po pp pq bk">Constant Trends: If transition rates stay constant, we can predict future states by repeatedly multiplying the daily state-transition matrix to new state values, helping us assess progress towards annual goals under unchanged conditions.</li><li id="fd4f" class="mw mx gu my b mz pr nb nc nd ps nf ng nh pt nj nk nl pu nn no np pv nr ns nt po pp pq bk">Dynamic Scenarios: By modifying transition rates, we can simulate complex scenarios. For instance, mimicking past changes in transition rates from a game launch allows us to predict the impact of similar future launches by altering the transition rate for a specific period.</li><li id="80bf" class="mw mx gu my b mz pr nb nc nd ps nf ng nh pt nj nk nl pu nn no np pv nr ns nt po pp pq bk">Steady State: We can calculate the steady state of the state-transition matrix (excluding new players) to estimate the MAA once all accounts have tried Netflix games and understand long-term retention and reactivation effects.</li></ul><p id="bf61" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Beyond predicting future states, we use the state machine for sensitivity analysis to find which transition rates most impact MAA. By making small changes to each transition rate we calculate the resulting MAA and measure its impact. This guides us in prioritizing efforts on top-of-funnel improvements, member retention, or reactivation.</p><h1 id="166d" class="nw nx gu bf ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os ot bk">Content Cash Modeling</h1><p id="3391" class="pw-post-body-paragraph mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt gn bk"><a class="af nv" href="https://www.linkedin.com/in/alexandra-diamond-b04902219/" rel="noopener ugc nofollow" target="_blank">Alex Diamond</a></p><p id="e03d" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">At Netflix we produce a variety of entertainment: movies, series, documentaries, stand-up specials, and more. Each format has a different production process and different patterns of cash spend, called our “Content Forecast”. Looking into the future, Netflix keeps a plan of how many titles we intend to produce, what kinds, and when. Because we don’t yet know what specific titles that content will eventually become, these generic placeholders are called “TBD Slots.” A sizable portion of our Content Forecast is represented by TBD Slots.</p><p id="4188" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Almost all businesses have a cash forecasting process informing how much cash they need in a given time period to continue executing on their plans. As plans change, the cash forecast will change. Netflix has a cash forecast that projects our cash needs to produce the titles we plan to make. This presents the question: how can we optimally forecast cash needs for TBD Slots, given we don’t have details on what real titles they will become?</p><p id="54e5" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">The large majority of our titles are funded throughout the production process — starting from when we begin developing the title to shooting the actual shows and movies to launch on our Netflix service.</p><p id="b01b" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Since cash spend is driven by what is happening on a production, we model it by breaking down into these three steps:</p><ol class=""><li id="8b6a" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt pw pp pq bk">Determine estimated production phase durations using historical actuals</li><li id="a35d" class="mw mx gu my b mz pr nb nc nd ps nf ng nh pt nj nk nl pu nn no np pv nr ns nt pw pp pq bk">Determine estimated percent of cash spent in each production phase</li><li id="1606" class="mw mx gu my b mz pr nb nc nd ps nf ng nh pt nj nk nl pu nn no np pv nr ns nt pw pp pq bk">Model the shape of cash spend within each phase</li></ol><p id="565e" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Putting these three pieces together allows us to generate a generic estimation of cash spend per day leading up to and beyond a title’s launch date (a proxy for “completion”). We could distribute this spend linearly across each phase, but this approach allows us to capture nuance around patterns of spend that ramp up slowly, or are concentrated at the start and taper off throughout.</p><figure class="pc pd pe pf pg ph oz pa paragraph-image"><div role="button" tabindex="0" class="pi pj fj pk bh pl"><div class="oz pa px"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*B6Abl5okW1BRfvrc%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*B6Abl5okW1BRfvrc%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*B6Abl5okW1BRfvrc%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*B6Abl5okW1BRfvrc%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*B6Abl5okW1BRfvrc%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*B6Abl5okW1BRfvrc%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*B6Abl5okW1BRfvrc%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*B6Abl5okW1BRfvrc 640w, https://miro.medium.com/v2/resize:fit:720/0*B6Abl5okW1BRfvrc 720w, https://miro.medium.com/v2/resize:fit:750/0*B6Abl5okW1BRfvrc 750w, https://miro.medium.com/v2/resize:fit:786/0*B6Abl5okW1BRfvrc 786w, https://miro.medium.com/v2/resize:fit:828/0*B6Abl5okW1BRfvrc 828w, https://miro.medium.com/v2/resize:fit:1100/0*B6Abl5okW1BRfvrc 1100w, https://miro.medium.com/v2/resize:fit:1400/0*B6Abl5okW1BRfvrc 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pm c" width="700" height="292" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="5046" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Before starting any math, we need to ensure a high quality historical dataset. Data quality plays a huge role in this work. For example, if we see 80% of our cash spent before production even started, it might be safe to say that either the production dates (which are manually captured) are incorrect or that title had a unique spending pattern that we don’t want to anticipate our future titles will follow.</p><p id="6932" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">For the first two steps, finding the estimated phase durations and cash percent per phase, we’ve found that simple math works best, for interpretability and consistency. We use a weighted average across our “clean” historical actuals to produce these estimated assumptions.</p><p id="05a0" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">For modeling the shape of spend throughout each phase, we perform constrained optimization to fit a 3rd degree polynomial function. The constraints include:</p><ol class=""><li id="01b8" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt pw pp pq bk">Must pass through the points (0,0) and (1,1). This ensures that 0% through the phase, 0% of that phase’s cash has been spent. Similarly, 100% through the phase, 100% of that phase’s cash has been spent.</li><li id="3032" class="mw mx gu my b mz pr nb nc nd ps nf ng nh pt nj nk nl pu nn no np pv nr ns nt pw pp pq bk">The derivative must be non-negative. This ensures that the function is monotonically increasing, avoiding counterintuitively forecasting any negative spend.</li></ol><p id="5632" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">The optimization’s objective function minimizes the sum of squared residuals and returns the coefficients of the polynomial that will guide the shape of cash spend through each phase.</p><p id="d6c2" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Once we have these coefficients, we can evaluate this polynomial at each day of the expected phase duration, and then multiply the result by the expected cash per phase. With some additional data processing, this yields an expected percent of cash spend each day leading up to and beyond the launch date, which we can base our forecasts on.</p><figure class="pc pd pe pf pg ph oz pa paragraph-image"><div role="button" tabindex="0" class="pi pj fj pk bh pl"><div class="oz pa px"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*ki-M_57G284X4IKo%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*ki-M_57G284X4IKo%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*ki-M_57G284X4IKo%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*ki-M_57G284X4IKo%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*ki-M_57G284X4IKo%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*ki-M_57G284X4IKo%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*ki-M_57G284X4IKo%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*ki-M_57G284X4IKo 640w, https://miro.medium.com/v2/resize:fit:720/0*ki-M_57G284X4IKo 720w, https://miro.medium.com/v2/resize:fit:750/0*ki-M_57G284X4IKo 750w, https://miro.medium.com/v2/resize:fit:786/0*ki-M_57G284X4IKo 786w, https://miro.medium.com/v2/resize:fit:828/0*ki-M_57G284X4IKo 828w, https://miro.medium.com/v2/resize:fit:1100/0*ki-M_57G284X4IKo 1100w, https://miro.medium.com/v2/resize:fit:1400/0*ki-M_57G284X4IKo 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pm c" width="700" height="280" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><h1 id="5ad2" class="nw nx gu bf ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os ot bk">Assistive Speech Recognition in Dubbing Workflows at Netflix</h1><p id="3a21" class="pw-post-body-paragraph mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt gn bk"><a class="af nv" href="https://www.linkedin.com/in/tanguycornuau/" rel="noopener ugc nofollow" target="_blank">Tanguy Cornau</a></p><p id="40e1" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Great stories can come from anywhere and be loved everywhere. At Netflix, we strive to make our titles accessible to a global audience, transcending language barriers to connect with viewers worldwide. One of the key ways we achieve this is through creating dubs in many languages.</p><p id="1dce" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">From the transcription of the original titles all the way to the delivery of the dub audio, we blend innovation with human expertise to preserve the original creative intent.</p><p id="65c6" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Leveraging technologies like Assistive Speech Recognition (ASR), we seek to make the <em class="nu">transcription</em> part of the process more efficient for our linguists. Transcription, in our context, involves creating a verbatim script of the spoken dialogue, along with precise timing information to perfectly align the text with the original video. With ASR, instead of starting the transcription from scratch, linguists get a pre-generated starting point which they can use and edit for complete accuracy.</p><figure class="pc pd pe pf pg ph oz pa paragraph-image"><div role="button" tabindex="0" class="pi pj fj pk bh pl"><div class="oz pa py"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*tdYvT28jMf3Z7QI_%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*tdYvT28jMf3Z7QI_%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*tdYvT28jMf3Z7QI_%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*tdYvT28jMf3Z7QI_%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*tdYvT28jMf3Z7QI_%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*tdYvT28jMf3Z7QI_%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*tdYvT28jMf3Z7QI_%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*tdYvT28jMf3Z7QI_ 640w, https://miro.medium.com/v2/resize:fit:720/0*tdYvT28jMf3Z7QI_ 720w, https://miro.medium.com/v2/resize:fit:750/0*tdYvT28jMf3Z7QI_ 750w, https://miro.medium.com/v2/resize:fit:786/0*tdYvT28jMf3Z7QI_ 786w, https://miro.medium.com/v2/resize:fit:828/0*tdYvT28jMf3Z7QI_ 828w, https://miro.medium.com/v2/resize:fit:1100/0*tdYvT28jMf3Z7QI_ 1100w, https://miro.medium.com/v2/resize:fit:1400/0*tdYvT28jMf3Z7QI_ 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pm c" width="700" height="265" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="c129" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">This efficiency enables linguists to focus more on other creative tasks, such as adding cultural annotations and references, which are crucial for downstream dubbing.</p><p id="a59a" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">With ASR, and other new and enhanced technologies we introduce, rigorous analytics and measurement are essential to their success. To effectively evaluate our ASR system, we’ve established a multi-layered measurement framework that provides comprehensive insights into its performance across many dimensions (for example, the accuracy of the text and timing predictions), offline and online.</p><p id="5e57" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">ASR is expected to perform differently for various languages; therefore, at a high level, we track metrics by original language of the show, allowing us to assess overall ASR effectiveness and identify trends across different linguistic contexts. We further break down performance by various dimensions, e.g. content type, genre, etc… to help us pinpoint specific areas where the ASR system may encounter difficulties. Furthermore, our framework allows us to conduct in-depth analyses of individual titles’ transcription, focusing on critical quality dimensions around text and timing accuracy of ASR suggestions. By zooming in on where the system falls short, we gain valuable insights into specific challenges, enabling us to further refine our understanding of ASR performance.</p><p id="b2bc" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">These measurement layers collectively empower us to continuously monitor, identify improvement areas, and implement targeted enhancements, ensuring that our ASR technology gets more and more accurate, effective, and helpful to linguists across diverse content types and languages. By refining our dubbing workflows through these innovations, we aim to keep improving the quality of our dubs to help great stories travel across the globe and bring joy to our members.</p></div></div></div><div class="ab cb pz qa qb qc" role="separator"><div class="gn go gp gq gr"><div class="ab cb"><div class="ci bh fz ga gb gc"><p id="f9f3" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Analytics Engineering is a key contributor to building our deep data culture at Netflix, and we are proud to have a large group of stunning colleagues that are not only applying but advancing our analytical capabilities at Netflix. The 2024 Analytics Summit continued to be a wonderful way to give visibility to one another on work across business verticals, celebrate our collective impact, and highlight what’s to come in analytics practice at Netflix.</p><p id="d18b" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">To learn more, follow the <a class="af nv" href="https://research.netflix.com/research-area/analytics" rel="noopener ugc nofollow" target="_blank">Netflix Research Site</a>, and if you are also interested in entertaining the world, have a look at <a class="af nv" href="https://explore.jobs.netflix.net/careers" rel="noopener ugc nofollow" target="_blank">our open roles</a>!</p></div></div></div></div></div>]]></description>
      <link>https://netflixtechblog.com/part-2-a-survey-of-analytics-engineering-work-at-netflix-4f1f53b4ab0f</link>
      <guid>https://netflixtechblog.com/part-2-a-survey-of-analytics-engineering-work-at-netflix-4f1f53b4ab0f</guid>
      <pubDate>Thu, 02 Jan 2025 22:07:00 +0100</pubDate>
    </item>
    <item>
      <title><![CDATA[Introducing Configurable Metaflow]]></title>
      <description><![CDATA[<div><div></div><p id="156c" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><a class="af nu" href="https://www.linkedin.com/in/david-j-berg/" rel="noopener ugc nofollow" target="_blank"><em class="nv">David J. Berg</em></a>*<em class="nv">, </em><a class="af nu" href="https://www.linkedin.com/in/david-casler-05a5278/" rel="noopener ugc nofollow" target="_blank"><em class="nv">David Casler</em></a>^, <a class="af nu" href="https://www.linkedin.com/in/romain-cledat-4a211a5/" rel="noopener ugc nofollow" target="_blank"><em class="nv">Romain Cledat</em></a>*<em class="nv">, </em><a class="af nu" href="https://www.linkedin.com/in/qian-huang-emma/" rel="noopener ugc nofollow" target="_blank"><em class="nv">Qian Huang</em></a>*<em class="nv">, </em><a class="af nu" href="https://www.linkedin.com/in/rui-lin-483a83111/" rel="noopener ugc nofollow" target="_blank"><em class="nv">Rui Lin</em></a>*<em class="nv">, </em><a class="af nu" href="https://www.linkedin.com/in/nissanpow/" rel="noopener ugc nofollow" target="_blank"><em class="nv">Nissan Pow</em></a>*<em class="nv">, </em><a class="af nu" href="https://www.linkedin.com/in/nurcansonmez/" rel="noopener ugc nofollow" target="_blank"><em class="nv">Nurcan Sonmez</em></a>*<em class="nv">, </em><a class="af nu" href="https://www.linkedin.com/in/shashanksrikanth/" rel="noopener ugc nofollow" target="_blank"><em class="nv">Shashank Srikanth</em></a>*<em class="nv">, </em><a class="af nu" href="https://www.linkedin.com/in/chaoying-wang/" rel="noopener ugc nofollow" target="_blank"><em class="nv">Chaoying Wang</em></a>*<em class="nv">, </em><a class="af nu" href="https://www.linkedin.com/in/reginalw/" rel="noopener ugc nofollow" target="_blank"><em class="nv">Regina Wang</em></a>*<em class="nv">, </em><a class="af nu" href="https://www.linkedin.com/in/zitingyu/" rel="noopener ugc nofollow" target="_blank"><em class="nv">Darin Yu</em></a>*<br />*: Model Development Team, Machine Learning Platform<br />^: Content Demand Modeling Team</p><p id="8a5d" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">A month ago at QConSF, we showcased how <a class="af nu" href="https://qconsf.com/presentation/nov2024/supporting-diverse-ml-systems-netflix" rel="noopener ugc nofollow" target="_blank">Netflix utilizes Metaflow to power a diverse set of ML and AI use cases</a>, managing thousands of unique Metaflow flows. This followed a previous <a class="af nu" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/supporting-diverse-ml-systems-at-netflix-2d2e6b6d205d">blog</a> on the same topic. Many of these projects are under constant development by dedicated teams with their own business goals and development best practices, such as the system that <a class="af nu" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/supporting-content-decision-makers-with-machine-learning-995b7b76006f">supports our content decision makers</a>, or the system that ranks which language subtitles are most valuable for a specific piece of content.</p><p id="8a83" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">As a central ML and AI platform team, our role is to empower our partner teams with tools that maximize their productivity and effectiveness, while adapting to their specific needs (not the other way around). This has been a guiding design principle with <a class="af nu" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/open-sourcing-metaflow-a-human-centric-framework-for-data-science-fa72e04a5d9">Metaflow since its inception</a>.</p><figure class="nz oa ob oc od oe nw nx paragraph-image"><div role="button" tabindex="0" class="of og fj oh bh oi"><div class="nw nx ny"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*XrOVl25ZLx8_4nHLRxNgDg.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*XrOVl25ZLx8_4nHLRxNgDg.png" /><img alt="" class="bh md oj c" width="700" height="413" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="ok ff ol nw nx om on bf b bg z du">Metaflow infrastructure stack</figcaption></figure><p id="46ac" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Standing on the shoulders of our extensive cloud infrastructure, Metaflow facilitates easy access to data, compute, and <a class="af nu" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/maestro-netflixs-workflow-orchestrator-ee13a06f9c78">production-grade workflow orchestration</a>, as well as built-in best practices for common concerns such as <a class="af nu" href="https://docs.metaflow.org/scaling/tagging" rel="noopener ugc nofollow" target="_blank">collaboration</a>, <a class="af nu" href="https://docs.metaflow.org/metaflow/basics#artifacts" rel="noopener ugc nofollow" target="_blank">versioning</a>, <a class="af nu" href="https://docs.metaflow.org/scaling/dependencies" rel="noopener ugc nofollow" target="_blank">dependency management</a>, and <a class="af nu" href="https://outerbounds.com/blog/metaflow-dynamic-cards" rel="noopener ugc nofollow" target="_blank">observability</a>, which teams use to setup ML/AI experiments and systems that work for them. As a result, Metaflow users at Netflix have been able to run millions of experiments over the past few years without wasting time on low-level concerns.</p><h1 id="15f1" class="oo op gu bf oq or os ot ou ov ow ox oy oz pa pb pc pd pe pf pg ph pi pj pk pl bk">A long standing FAQ: configurable flows</h1><p id="8abd" class="pw-post-body-paragraph mw mx gu my b mz pm nb nc nd pn nf ng nh po nj nk nl pp nn no np pq nr ns nt gn bk">While Metaflow aims to be un-opinionated about some of the upper levels of the stack, some teams within Netflix have developed their own opinionated tooling. As part of Metaflow’s adaptation to their specific needs, we constantly try to understand what has been developed and, more importantly, what gaps these solutions are filling.</p><p id="bd7e" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">In some cases, we determine that the gap being addressed is very team specific, or too opinionated at too high a level in the stack, and we therefore decide to not develop it within Metaflow. In other cases, however, we realize that we can develop an underlying construct that aids in filling that gap. Note that even in that case, we do not always aim to completely fill the gap and instead focus on extracting a more general lower level concept that can be leveraged by that particular user but also by others. One such recurring pattern we noticed at Netflix is the need to deploy sets of closely related flows, often as part of a larger pipeline involving table creations, ETLs, and deployment jobs. Frequently, practitioners want to <a class="af nu" href="https://docs.metaflow.org/production/coordinating-larger-metaflow-projects" rel="noopener ugc nofollow" target="_blank">experiment with variants</a> of these flows, testing new data, new parameterizations, or new algorithms, while keeping the overall structure of the flow or flows intact.</p><p id="64c6" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">A natural solution is to make flows configurable using configuration files, so variants can be defined without changing the code. Thus far, there hasn’t been a built-in solution for configuring flows, so teams have built their bespoke solutions leveraging Metaflow’s <a class="af nu" href="https://docs.metaflow.org/metaflow/basics#advanced-parameters" rel="noopener ugc nofollow" target="_blank">JSON-typed Parameters</a>, <a class="af nu" href="https://docs.metaflow.org/scaling/data#data-in-local-files" rel="noopener ugc nofollow" target="_blank">IncludeFile</a>, and <a class="af nu" href="https://docs.metaflow.org/production/scheduling-metaflow-flows/scheduling-with-aws-step-functions#deploy-time-parameters" rel="noopener ugc nofollow" target="_blank">deploy-time Parameters</a> or deploying their own home-grown solution (often with great pain). However, none of these solutions make it easy to configure all aspects of the flow’s behavior, decorators in particular.</p><figure class="nz oa ob oc od oe nw nx paragraph-image"><div role="button" tabindex="0" class="of og fj oh bh oi"><div class="nw nx pr"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*3f9q7PZgxYX8rRygIOWXyA.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*3f9q7PZgxYX8rRygIOWXyA.png" /><img alt="" class="bh md oj c" width="700" height="434" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="ok ff ol nw nx om on bf b bg z du">Requests for a feature like Metaflow Config</figcaption></figure><p id="0e00" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Outside Netflix, we have seen similar frequently asked questions on the <a class="af nu" href="http://chat.metaflow.org" rel="noopener ugc nofollow" target="_blank">Metaflow community Slack</a> as shown in the user quotes above:</p><ul class=""><li id="8010" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt ps pt pu bk">how can I adjust <a class="af nu" href="https://docs.metaflow.org/scaling/remote-tasks/requesting-resources" rel="noopener ugc nofollow" target="_blank">the @resource requirements</a>, such as CPU or memory, without having to hardcode the values in my flows?</li><li id="129c" class="mw mx gu my b mz pv nb nc nd pw nf ng nh px nj nk nl py nn no np pz nr ns nt ps pt pu bk">how to adjust <a class="af nu" href="https://docs.metaflow.org/production/scheduling-metaflow-flows/scheduling-with-argo-workflows#time-based-triggering" rel="noopener ugc nofollow" target="_blank">the triggering @schedule</a> programmatically, so our production and staging deployments can run at different cadences?</li></ul><h1 id="69ac" class="oo op gu bf oq or os ot ou ov ow ox oy oz pa pb pc pd pe pf pg ph pi pj pk pl bk">New in Metaflow: Configs!</h1><p id="36e8" class="pw-post-body-paragraph mw mx gu my b mz pm nb nc nd pn nf ng nh po nj nk nl pp nn no np pq nr ns nt gn bk">Today, to answer the FAQ, we introduce a new — small but mighty — feature in Metaflow: <a class="af nu" href="https://docs.metaflow.org/metaflow/configuring-flows/introduction" rel="noopener ugc nofollow" target="_blank">a Config object</a>. Configs complement the existing Metaflow constructs of artifacts and Parameters, by allowing you to configure all aspects of the flow, decorators in particular, prior to any run starting. At the end of the day, artifacts, Parameters and Configs are all stored as artifacts by Metaflow but they differ in when they are persisted as shown in the diagram below:</p><figure class="nz oa ob oc od oe nw nx paragraph-image"><div role="button" tabindex="0" class="of og fj oh bh oi"><div class="nw nx qa"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*L-klklqt1n9LKXG0jh-fTw.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*L-klklqt1n9LKXG0jh-fTw.png" /><img alt="" class="bh md oj c" width="700" height="302" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="ok ff ol nw nx om on bf b bg z du">Different data artifacts in Metaflow</figcaption></figure><p id="448b" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Said another way:</p><ul class=""><li id="b920" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt ps pt pu bk">An<strong class="my gv"> artifact</strong> is resolved and persisted to the datastore at the end of each task.</li><li id="73c1" class="mw mx gu my b mz pv nb nc nd pw nf ng nh px nj nk nl py nn no np pz nr ns nt ps pt pu bk">A<strong class="my gv"> parameter</strong> is resolved and persisted at the start of a run; it can therefore be modified up to that point. One common use case is to use <a class="af nu" href="https://docs.metaflow.org/production/event-triggering" rel="noopener ugc nofollow" target="_blank">triggers</a> to pass values to a run right before executing. Parameters can only be used within your step code.</li><li id="61b0" class="mw mx gu my b mz pv nb nc nd pw nf ng nh px nj nk nl py nn no np pz nr ns nt ps pt pu bk">A<strong class="my gv"> config</strong> is resolved and persisted when the flow is deployed. When using a scheduler such as <a class="af nu" href="https://docs.metaflow.org/production/scheduling-metaflow-flows/scheduling-with-argo-workflows" rel="noopener ugc nofollow" target="_blank">Argo Workflows</a>, deployment happens when create’ing the flow. In the case of a local run, “deployment” happens just prior to the execution of the run — think of “deployment” as gathering all that is needed to run the flow. Unlike parameters, configs can be used more widely in your flow code, particularly, they can be used in step or flow level decorators as well as to set defaults for parameters. Configs can of course also be used within your flow.</li></ul><p id="226d" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">As an example, you can specify a Config that reads a pleasantly human-readable configuration file, formatted as <a class="af nu" href="https://toml.io/en/" rel="noopener ugc nofollow" target="_blank">TOML</a>. The Config specifies a triggering ‘@schedule’ and ‘@resource’ requirements, as well as application-specific parameters for this specific deployment:</p><pre class="nz oa ob oc od qb qc qd bp qe bb bk">[schedule]<br />cron = "0 * * * *"[model]<br />optimizer = "adam"<br />learning_rate = 0.5[resources]<br />cpu = 1</pre><p id="be39" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Using the newly released Metaflow 2.13, you can configure a flow with a Config like above, as demonstrated by this flow:</p><pre class="nz oa ob oc od qb qc qd bp qe bb bk">import pprint<br />from metaflow import FlowSpec, step, Config, resources, config_expr, schedule@schedule(cron=config_expr("config.schedule.cron"))<br />class ConfigurableFlow(FlowSpec):<br />    config = Config("config", default="myconfig.toml", parser="tomllib.loads")@resources(cpu=config.resources.cpu)<br />    @step<br />    def start(self):<br />        print("Config loaded:")<br />        pprint.pp(self.config)<br />        self.next(self.end)@step<br />    def end(self):<br />        passif __name__ == "__main__":<br />    ConfigurableFlow()</pre><p id="251b" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">There is a lot going on in the code above, a few highlights:</p><ul class=""><li id="cc56" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt ps pt pu bk">you can refer to configs <em class="nv">before</em> they have been defined using ‘config_expr’.</li><li id="2cf7" class="mw mx gu my b mz pv nb nc nd pw nf ng nh px nj nk nl py nn no np pz nr ns nt ps pt pu bk">you can define arbitrary <a class="af nu" href="https://docs.metaflow.org/metaflow/configuring-flows/parsing-configs" rel="noopener ugc nofollow" target="_blank">parsers</a> — using a string means the parser doesn’t even have to be present remotely!</li></ul><p id="1590" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">From the developer’s point of view, Configs behave like dictionary-like artifacts. For convenience, they support the dot-syntax (when possible) for accessing keys, making it easy to access values in a nested configuration. You can also unpack the whole Config (or a subtree of it) with Python’s standard dictionary unpacking syntax, ‘**config’. The standard dictionary subscript notation is also available.</p><p id="5e1e" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Since Configs turn into dictionary artifacts, they get versioned and persisted automatically as artifacts. You can <a class="af nu" href="https://docs.metaflow.org/metaflow/client" rel="noopener ugc nofollow" target="_blank">access Configs of any past runs easily through the Client API</a>. As a result, your data, models, code, Parameters, Configs, and <a class="af nu" href="https://docs.metaflow.org/scaling/dependencies" rel="noopener ugc nofollow" target="_blank">execution environments</a> are all stored as a consistent bundle — neatly organized in <a class="af nu" href="https://docs.metaflow.org/scaling/tagging" rel="noopener ugc nofollow" target="_blank">Metaflow namespaces</a> — paving the way for easily reproducible, consistent, low-boilerplate, and now easily configurable experiments and robust production deployments.</p><h1 id="9d64" class="oo op gu bf oq or os ot ou ov ow ox oy oz pa pb pc pd pe pf pg ph pi pj pk pl bk">More than a humble config file</h1><p id="ee02" class="pw-post-body-paragraph mw mx gu my b mz pm nb nc nd pn nf ng nh po nj nk nl pp nn no np pq nr ns nt gn bk">While you can get far by accompanying your flow with a simple config file (stored in your favorite format, thanks to <a class="af nu" href="https://docs.metaflow.org/metaflow/configuring-flows/parsing-configs" rel="noopener ugc nofollow" target="_blank">user-definable parsers</a>), Configs unlock a number of advanced use cases. Consider these examples from the updated documentation:</p><ul class=""><li id="626e" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt ps pt pu bk">You can <a class="af nu" href="https://docs.metaflow.org/metaflow/configuring-flows/basic-configuration#mixing-configs-and-parameters" rel="noopener ugc nofollow" target="_blank"><strong class="my gv">choose the right level of runtime configurability</strong></a> versus fixed deployments by mixing Parameters and Configs. For instance, you can use a Config to define a default value for a parameter which can be <a class="af nu" href="https://docs.metaflow.org/production/event-triggering/external-events#passing-parameters-in-events" rel="noopener ugc nofollow" target="_blank">overridden by a real-time event</a> as a run is triggered.</li><li id="83e8" class="mw mx gu my b mz pv nb nc nd pw nf ng nh px nj nk nl py nn no np pz nr ns nt ps pt pu bk">You can define a custom parser to <a class="af nu" href="https://docs.metaflow.org/metaflow/configuring-flows/parsing-configs#validating-configs-with-pydantic" rel="noopener ugc nofollow" target="_blank"><strong class="my gv">validate the configuration</strong></a>, e.g. using the popular <a class="af nu" href="https://docs.pydantic.dev/latest/" rel="noopener ugc nofollow" target="_blank">Pydantic</a> library.</li><li id="19d3" class="mw mx gu my b mz pv nb nc nd pw nf ng nh px nj nk nl py nn no np pz nr ns nt ps pt pu bk">You are not limited to using a single file: you can leverage a configuration manager like <a class="af nu" href="https://omegaconf.readthedocs.io/en/2.3_branch/" rel="noopener ugc nofollow" target="_blank">OmegaConf</a> or <a class="af nu" href="https://hydra.cc/" rel="noopener ugc nofollow" target="_blank">Hydra</a> to <a class="af nu" href="https://docs.metaflow.org/metaflow/configuring-flows/parsing-configs#advanced-configurations-with-omegaconf" rel="noopener ugc nofollow" target="_blank"><strong class="my gv">manage a hierarchy of cascading configuration files</strong></a>. You can also use a domain-specific tool for generating Configs, such as Netflix’s <em class="nv">Metaboost</em> which we cover below.</li><li id="ca32" class="mw mx gu my b mz pv nb nc nd pw nf ng nh px nj nk nl py nn no np pz nr ns nt ps pt pu bk">You can also <a class="af nu" href="https://docs.metaflow.org/metaflow/configuring-flows/custom-parsers#generating-configs-programmatically" rel="noopener ugc nofollow" target="_blank"><strong class="my gv">generate configurations on the fly</strong></a>, e.g. fetch Configs from an external service, or inspect the execution environment, such as the current GIT branch, and include it as an extra piece of context in runs.</li></ul><p id="ac3f" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">A major benefit of Config over previous more hacky solutions for configuring flows is that they work seamlessly with other features of Metaflow: you can run steps remotely and deploy flows to production, even when relying on custom parsers, without having to worry about packaging Configs or parsers manually or keeping Configs consistent across tasks. Configs also work with the <a class="af nu" href="https://docs.metaflow.org/metaflow/managing-flows/runner" rel="noopener ugc nofollow" target="_blank">Runner</a> and <a class="af nu" href="https://docs.metaflow.org/metaflow/managing-flows/deployer" rel="noopener ugc nofollow" target="_blank">Deployer</a>.</p><h1 id="8ae8" class="oo op gu bf oq or os ot ou ov ow ox oy oz pa pb pc pd pe pf pg ph pi pj pk pl bk">The Hollywood principle: don’t call us, we’ll call you</h1><p id="84f1" class="pw-post-body-paragraph mw mx gu my b mz pm nb nc nd pn nf ng nh po nj nk nl pp nn no np pq nr ns nt gn bk">When used in conjunction with a configuration manager like <a class="af nu" href="https://hydra.cc" rel="noopener ugc nofollow" target="_blank">Hydra</a>, Configs enable a pattern that is highly relevant for ML and AI use cases: orchestrating experiments over multiple configurations or sweeping over parameter spaces. While Metaflow has always supported <a class="af nu" href="https://docs.outerbounds.com/grid-search-with-metaflow/" rel="noopener ugc nofollow" target="_blank">sweeping over parameter grids</a> easily using foreaches, it hasn’t been easily possible to alter the flow itself, e.g. to change <a class="af nu" href="https://docs.metaflow.org/api/step-decorators/resources" rel="noopener ugc nofollow" target="_blank">@resources</a> or <a class="af nu" href="https://docs.metaflow.org/api/step-decorators/conda" rel="noopener ugc nofollow" target="_blank">@pypi/@conda</a> dependencies for every experiment.</p><p id="f7ef" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">In a typical case, you trigger a Metaflow flow that consumes a configuration file, changing <em class="nv">how</em> a run behaves. With Hydra, you can <a class="af nu" href="https://en.wikipedia.org/wiki/Inversion_of_control" rel="noopener ugc nofollow" target="_blank">invert the control</a>: it is Hydra that decides <em class="nv">what</em> gets run based on a configuration file. Thanks to Metaflow’s new <a class="af nu" href="https://docs.metaflow.org/metaflow/managing-flows/runner" rel="noopener ugc nofollow" target="_blank">Runner</a> and <a class="af nu" href="https://docs.metaflow.org/metaflow/managing-flows/deployer" rel="noopener ugc nofollow" target="_blank">Deployer</a> APIs, you can create a Hydra app that operates Metaflow programmatically — for instance, to deploy and execute hundreds of variants of a flow in a large-scale experiment.</p><p id="1fda" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><a class="af nu" href="https://docs.metaflow.org/metaflow/configuring-flows/config-driven-experimentation" rel="noopener ugc nofollow" target="_blank">Take a look at two interesting examples of this pattern</a> in the documentation. As a teaser, this video shows Hydra orchestrating deployment of tens of Metaflow flows, each of which benchmarks PyTorch using a varying number of CPU cores and tensor sizes, updating a visualization of the results in real-time as the experiment progresses:</p><figure class="nz oa ob oc od oe"><div class="qk jf l fj"><figcaption class="ok ff ol nw nx om on bf b bg z du">Example using Hydra with Metaflow</figcaption></div></figure><h1 id="88a6" class="oo op gu bf oq or os ot ou ov ow ox oy oz pa pb pc pd pe pf pg ph pi pj pk pl bk">Metaboosting Metaflow — based on a true story</h1><p id="4282" class="pw-post-body-paragraph mw mx gu my b mz pm nb nc nd pn nf ng nh po nj nk nl pp nn no np pq nr ns nt gn bk">To give a motivating example of what configurations look like at Netflix in practice, let’s consider <em class="nv">Metaboost</em>, an internal Netflix CLI tool that helps ML practitioners manage, develop and execute their cross-platform projects, somewhat similar to the open-source Hydra discussed above but with specific integrations to the Netflix ecosystem. Metaboost is an example of an opinionated framework developed by a team already using Metaflow. In fact, a part of the inspiration for introducing Configs in Metaflow came from this very use case.</p><p id="9322" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Metaboost serves as a single interface to three different internal platforms at Netflix that manage ETL/Workflows (<a class="af nu" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/maestro-netflixs-workflow-orchestrator-ee13a06f9c78"><em class="nv">Maestro</em></a>), Machine Learning Pipelines (<a class="af nu" href="https://docs.metaflow.org" rel="noopener ugc nofollow" target="_blank"><em class="nv">Metaflow</em></a>) and Data Warehouse Tables (<em class="nv">Kragle</em>). In this context, having a single configuration system to manage a ML project holistically gives users increased project coherence and decreased project risk.</p><h2 id="fbf6" class="qn op gu bf oq qo qp dy ou qq qr ea oy nh qs qt qu nl qv qw qx np qy qz ra rb bk">Configuration in Metaboost</h2><p id="a6d0" class="pw-post-body-paragraph mw mx gu my b mz pm nb nc nd pn nf ng nh po nj nk nl pp nn no np pq nr ns nt gn bk">Ease of configuration and templatizing are core values of Metaboost. Templatizing in Metaboost is achieved through the concept of <em class="nv">bindings</em>, wherein we can <em class="nv">bind</em> a Metaflow pipeline to an arbitrary label, and then create a corresponding bespoke configuration for that label. The binding-connected configuration is then merged into a global set of configurations containing such information as GIT repository, branch, etc. Binding a Metaflow, will also signal to Metaboost that it should instantiate the Metaflow flow once per binding into our orchestration cluster.</p><p id="747b" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Imagine a ML practitioner on the Netflix Content ML team, sourcing features from hundreds of columns in our data warehouse, and creating a multitude of models against a <em class="nv">growing</em> suite of metrics. When a brand new content metric comes along, with Metaboost, the first version of the metric’s predictive model can easily be created by simply swapping the target column against which the model is trained.</p><p id="dd0d" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Subsequent versions of the model will result from experimenting with hyper parameters, tweaking feature engineering, or conducting feature diets. Metaboost’s bindings, and their integration with Metaflow Configs, can be leveraged to scale the number of experiments as fast as a scientist can create experiment based configurations.</p><h2 id="edb4" class="qn op gu bf oq qo qp dy ou qq qr ea oy nh qs qt qu nl qv qw qx np qy qz ra rb bk">Scaling experiments with Metaboost bindings — backed by Metaflow Config</h2><p id="9ad4" class="pw-post-body-paragraph mw mx gu my b mz pm nb nc nd pn nf ng nh po nj nk nl pp nn no np pq nr ns nt gn bk">Consider a Metaboost ML project named `demo` that creates and loads data to custom tables (ETL managed by Maestro), and then trains a simple model on this data (ML Pipeline managed by Metaflow). The project structure of this repository might look like the following:</p><pre class="nz oa ob oc od qb qc qd bp qe bb bk">├── metaflows<br />│   ├── custom                               -&gt; custom python code, used by<br />|   |   |                                       Metaflow<br />│   │   ├── data.py<br />│   │   └── model.py<br />│   └── training.py                          -&gt; defines our Metaflow pipeline<br />├── schemas<br />│   ├── demo_features_f.tbl.yaml             -&gt; table DDL, stores our ETL<br />|   |                                           output, Metaflow input<br />│   └── demo_predictions_f.tbl.yaml          -&gt; table DDL,<br />|                                               stores our Metaflow output<br />├── settings<br />│   ├── settings.configuration.EXP_01.yaml   -&gt; defines the additive<br />|   |                                           config for Experiment 1<br />│   ├── settings.configuration.EXP_02.yaml   -&gt; defines the additive<br />|   |                                           config for Experiment 2<br />│   ├── settings.configuration.yaml          -&gt; defines our global<br />|   |                                           configuration<br />│   └── settings.environment.yaml            -&gt; defines parameters based on<br />|                                               git branch (e.g. READ_DB)<br />├── tests<br />├── workflows<br />│   ├── sql<br />│   ├── demo.demo_features_f.sch.yaml        -&gt; Maestro workflow, defines ETL<br />│   └── demo.main.sch.yaml                   -&gt; Maestro workflow, orchestrates<br />|                                               ETLs and Metaflow<br />└── metaboost.yaml                           -&gt; defines our project for<br />                                                Metaboost</pre><p id="f679" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">The configuration files in the settings directory above contain the following YAML files:</p><pre class="nz oa ob oc od qb qc qd bp qe bb bk"># settings.configuration.yaml (global configuration)<br />model:<br />  fit_intercept: True<br />conda:<br />  numpy: '1.22.4'<br />  "scikit-learn": '1.4.0'</pre><pre class="rc qb qc qd bp qe bb bk"># settings.configuration.EXP_01.yaml<br />target_column: metricA<br />features:<br />  - runtime<br />  - content_type<br />  - top_billed_talent</pre><pre class="rc qb qc qd bp qe bb bk"># settings.configuration.EXP_02.yaml<br />target_column: metricA<br />features:<br />  - runtime<br />  - director<br />  - box_office</pre><p id="069e" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Metaboost will merge each experiment configuration (<em class="nv">*.EXP*.yaml</em>) into the global configuration (settings.configuration.yaml) <em class="nv">individually</em> at Metaboost command initialization. Let’s take a look at how Metaboost combines these configurations with a Metaboost command:</p><pre class="nz oa ob oc od qb qc qd bp qe bb bk">(venv-demo) ~/projects/metaboost-demo [branch=demoX] <br />$ metaboost metaflow settings show --yaml-path=configurationbinding=EXP_01:<br />model:                     -&gt; defined in setting.configuration.yaml (global)<br />  fit_intercept: true<br />conda:                     -&gt; defined in setting.configuration.yaml (global)<br />  numpy: 1.22.4<br />  "scikit-learn": 1.4.0<br />target_column: metricA     -&gt; defined in setting.configuration.EXP_01.yaml<br />features:                  -&gt; defined in setting.configuration.EXP_01.yaml<br />- runtime<br />- content_type<br />- top_billed_talentbinding=EXP_02:<br />model:                     -&gt; defined in setting.configuration.yaml (global)<br />  fit_intercept: true<br />conda:                     -&gt; defined in setting.configuration.yaml (global)<br />  numpy: 1.22.4<br />  "scikit-learn": 1.4.0<br />target_column: metricA     -&gt; defined in setting.configuration.EXP_02.yaml<br />features:                  -&gt; defined in setting.configuration.EXP_02.yaml<br />- runtime<br />- director<br />- box_office</pre><p id="a78e" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Metaboost understands it should deploy/run two independent instances of training.py — one for the EXP_01 binding and one for the EXP_02 binding. You can also see that Metaboost is aware that the tables and ETL workflows are <em class="nv">not bound</em>, and should only be deployed once. These details of which artifacts to bind and which to leave unbound are encoded in the project’s top-level metaboost.yaml file.</p><pre class="nz oa ob oc od qb qc qd bp qe bb bk">(venv-demo) ~/projects/metaboost-demo [branch=demoX] <br />$ metaboost project listTables (metaboost table list):<br />schemas/demo_predictions_f.tbl.yaml (binding=default):<br />    table_path=prodhive/demo_db/demo_predictions_f<br />schemas/demo_features_f.tbl.yaml (binding=default):<br />    table_path=prodhive/demo_db/demo_features_fWorkflows (metaboost workflow list):<br />workflows/demo.demo_features_f.sch.yaml (binding=default):<br />    cluster=sandbox, workflow.id=demo.branch_demox.demo_features_f<br />workflows/demo.main.sch.yaml (binding=default):<br />    cluster=sandbox, workflow.id=demo.branch_demox.mainMetaflows (metaboost metaflow list):<br />metaflows/training.py (binding=EXP_01): -&gt; EXP_01 instance of training.py<br />    cluster=sandbox, workflow.id=demo.branch_demox.EXP_01.training   <br />metaflows/training.py (binding=EXP_02): -&gt; EXP_02 instance of training.py<br />    cluster=sandbox, workflow.id=demo.branch_demox.EXP_02.training</pre><p id="c673" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Below is a simple Metaflow pipeline that fetches data, executes feature engineering, and trains a LinearRegression model. The work to integrate Metaboost Settings into a user’s Metaflow pipeline (implemented using Metaflow Configs) is as easy as adding a single mix-in to the FlowSpec definition:</p><pre class="nz oa ob oc od qb qc qd bp qe bb bk">from metaflow import FlowSpec, Parameter, conda_base, step<br />from custom.data import feature_engineer, get_data<br />from metaflow.metaboost import MetaboostSettings@conda_base(<br />    libraries=MetaboostSettings.get_deploy_time_settings("configuration.conda")<br />)<br />class DemoTraining(FlowSpec, MetaboostSettings):<br />    prediction_date = Parameter("prediction_date", type=int, default=-1)@step<br />    def start(self):<br />        # get show_settings() for free with the mixin<br />        # and get convenient debugging info<br />        self.show_settings(exclude_patterns=["artifact*", "system*"])self.next(self.get_features)@step<br />    def get_features(self):<br />        # feature engineers on our extracted data<br />        self.fe_df = feature_engineer(<br />            # loads data from our ETL pipeline<br />            data=get_data(prediction_date=self.prediction_date),<br />            features=self.settings.configuration.features +<br />                [self.settings.configuration.target_column]<br />        )self.next(self.train)@step<br />    def train(self):<br />        from sklearn.linear_model import LinearRegression# trains our model<br />        self.model = LinearRegression(<br />            fit_intercept=self.settings.configuration.model.fit_intercept<br />        ).fit(<br />            X=self.fe_df[self.settings.configuration.features],<br />            y=self.fe_df[self.settings.configuration.target_column]<br />        )<br />        print(f"Fit slope: {self.model.coef_[0]}")<br />        print(f"Fit intercept: {self.model.intercept_}")self.next(self.end)@step<br />    def end(self):<br />        passif __name__ == "__main__":<br />    DemoTraining()</pre><p id="1fa8" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">The Metaflow Config is added to the FlowSpec by mixing in the MetaboostSettings class. Referencing a configuration value is as easy as using the dot syntax to drill into whichever parameter you’d like.</p><p id="5872" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Finally let’s take a look at the output from our sample Metaflow above. We execute experiment EXP_01 with</p><pre class="nz oa ob oc od qb qc qd bp qe bb bk">metaboost metaflow run --binding=EXP_01</pre><p id="ea6c" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">which upon execution will merge the configurations into a single <em class="nv">settings</em> file (shown previously) and serialize it as a yaml file to the <em class="nv">.metaboost/settings/compiled/</em> directory.</p><p id="0e34" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">You can see the actual command and args that were sub-processed in the <em class="nv">Metaboost Execution</em> section below. Please note the <strong class="my gv">–config</strong> argument pointing to the serialized yaml file, and then subsequently accessible via <strong class="my gv">self.settings</strong>. Also note the convenient printing of configuration values to stdout during the start step using a mixed in function named <strong class="my gv">show_settings()</strong>.</p><pre class="nz oa ob oc od qb qc qd bp qe bb bk">(venv-demo) ~/projects/metaboost-demo [branch=demoX] <br />$ metaboost metaflow run --binding=EXP_01Metaboost Execution: <br /> - python3.10 /root/repos/cdm-metaboost-irl/metaflows/training.py<br />   --no-pylint --package-suffixes=.py --environment=conda<br />   --config settings<br />   .metaboost/settings/compiled/settings.branch_demox.EXP_01.training.mP4eIStG.yaml<br />   run --prediction_date20241006Metaflow 2.12.39+nflxfastdata(2.13.5);nflx(2.13.5);metaboost(0.0.27)<br />  executing DemoTraining for user:dcasler<br />Validating your flow...<br />    The graph looks good!<br />Bootstrapping Conda environment... (this could take a few minutes)<br />All packages already cached in s3.<br />All environments already cached in s3.Workflow starting (run-id 50), see it in the UI at<br />https://metaflowui.prod.netflix.net/DemoTraining/50[50/start/251640833] Task is starting.<br />[50/start/251640833] Configuration Values:<br />[50/start/251640833]   settings.configuration.conda.numpy            = 1.22.4<br />[50/start/251640833]   settings.configuration.features.0             = runtime<br />[50/start/251640833]   settings.configuration.features.1             = content_type<br />[50/start/251640833]   settings.configuration.features.2             = top_billed_talent<br />[50/start/251640833]   settings.configuration.model.fit_intercept    = True<br />[50/start/251640833]   settings.configuration.target_column          = metricA<br />[50/start/251640833]   settings.environment.READ_DATABASE            = data_warehouse_prod<br />[50/start/251640833]   settings.environment.TARGET_DATABASE          = demo_dev<br />[50/start/251640833] Task finished successfully.[50/get_features/251640840] Task is starting.<br />[50/get_features/251640840] Task finished successfully.[50/train/251640854] Task is starting.<br />[50/train/251640854] Fit slope: 0.4702672504331096<br />[50/train/251640854] Fit intercept: -6.247919678070083<br />[50/train/251640854] Task finished successfully.[50/end/251640868] Task is starting.<br />[50/end/251640868] Task finished successfully.Done! See the run in the UI at<br />https://metaflowui.prod.netflix.net/DemoTraining/50</pre><h2 id="d6d2" class="qn op gu bf oq qo qp dy ou qq qr ea oy nh qs qt qu nl qv qw qx np qy qz ra rb bk">Takeaways</h2><p id="feb6" class="pw-post-body-paragraph mw mx gu my b mz pm nb nc nd pn nf ng nh po nj nk nl pp nn no np pq nr ns nt gn bk">Metaboost is an integration tool that aims to ease the project development, management and execution burden of ML projects at Netflix. It employs a configuration system that combines git based parameters, global configurations and arbitrarily <em class="nv">bound</em> configuration files for use during execution against internal Netflix platforms.</p><p id="7ec5" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Integrating this configuration system with the new Config in Metaflow is incredibly simple (by design), only requiring users to add a mix-in class to their FlowSpec — <a class="af nu" href="https://docs.metaflow.org/metaflow/configuring-flows/custom-parsers#including-default-configs-in-flows" rel="noopener ugc nofollow" target="_blank">similar to this example in Metaflow documentation</a> — and then reference the configuration values in steps or decorators. The example above templatizes a training Metaflow for the sake of experimentation, but users could just as easily use bindings/configs to templatize their flows across target metrics, business initiatives or any other arbitrary lines of work.</p><h1 id="d730" class="oo op gu bf oq or os ot ou ov ow ox oy oz pa pb pc pd pe pf pg ph pi pj pk pl bk">Try it at home</h1><p id="2f85" class="pw-post-body-paragraph mw mx gu my b mz pm nb nc nd pn nf ng nh po nj nk nl pp nn no np pq nr ns nt gn bk">It couldn’t be easier to get started with Configs! Just</p><pre class="nz oa ob oc od qb qc qd bp qe bb bk">pip install -U metaflow</pre><p id="e0ec" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">to get the latest version and <a class="af nu" href="https://docs.metaflow.org/metaflow/configuring-flows/introduction" rel="noopener ugc nofollow" target="_blank">head to the updated documentation</a> for examples. If you are impatient, you can find and execute <a class="af nu" href="https://github.com/outerbounds/config-examples" rel="noopener ugc nofollow" target="_blank">all config-related examples in this repository</a> as well.</p><p id="53e2" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">If you have any questions or feedback about Config (or other Metaflow features), you can reach out to us at the <a class="af nu" href="http://chat.metaflow.org" rel="noopener ugc nofollow" target="_blank">Metaflow community Slack</a>.</p><h1 id="24a8" class="oo op gu bf oq or os ot ou ov ow ox oy oz pa pb pc pd pe pf pg ph pi pj pk pl bk">Acknowledgments</h1><p id="97f9" class="pw-post-body-paragraph mw mx gu my b mz pm nb nc nd pn nf ng nh po nj nk nl pp nn no np pq nr ns nt gn bk">We would like to thank <a class="af nu" href="https://outerbounds.co" rel="noopener ugc nofollow" target="_blank">Outerbounds</a> for their collaboration on this feature; for rigorously testing it and developing a repository of examples to showcase some of the possibilities offered by this feature.</p></div>]]></description>
      <link>https://netflixtechblog.com/introducing-configurable-metaflow-d2fb8e9ba1c6</link>
      <guid>https://netflixtechblog.com/introducing-configurable-metaflow-d2fb8e9ba1c6</guid>
      <pubDate>Fri, 20 Dec 2024 08:11:00 +0100</pubDate>
    </item>
    <item>
      <title><![CDATA[Part 1: A Survey of Analytics Engineering Work at Netflix]]></title>
      <description><![CDATA[<div class="gn go gp gq gr"><div class="ab cb"><div class="ci bh fz ga gb gc"><div><div></div><p id="9766" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><em class="nu">This article is the first in a multi-part series sharing a breadth of Analytics Engineering work at Netflix, recently presented as part of our annual internal Analytics Engineering conference. We kick off with a few topics focused on how we’re empowering Netflix to efficiently produce and effectively deliver high quality, actionable analytic insights across the company. Subsequent posts will detail examples of exciting analytic engineering domain applications and aspects of the technical craft.</em></p><p id="1e6c" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">At Netflix, we seek to entertain the world by ensuring our members find the shows and movies that will thrill them. Analytics at Netflix powers everything from understanding what content will excite and bring members back for more to how we should produce and distribute a content slate that maximizes member joy. Analytics Engineers deliver these insights by establishing deep business and product partnerships; translating business challenges into solutions that unblock critical decisions; and designing, building, and maintaining end-to-end analytical systems.</p><p id="9e33" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Each year, we bring the Analytics Engineering community together for an Analytics Summit — a 3-day internal conference to share analytical deliverables across Netflix, discuss analytic practice, and build relationships within the community. We covered a broad array of exciting topics and wanted to spotlight a few to give you a taste of what we’re working on across Analytics Engineering at Netflix!</p><h1 id="a458" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">DataJunction: Unifying Experimentation and Analytics</h1><p id="6f96" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk"><a class="af oy" href="https://www.linkedin.com/in/shyiann/" rel="noopener ugc nofollow" target="_blank">Yian Shang</a>, <a class="af oy" href="https://www.linkedin.com/in/anhqle/" rel="noopener ugc nofollow" target="_blank">Anh Le</a></p><p id="2200" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">At Netflix, like in many organizations, creating and using metrics is often more complex than it should be. Metric definitions are often scattered across various databases, documentation sites, and code repositories, making it difficult for analysts and data scientists to find reliable information quickly. This fragmentation leads to inconsistencies and wastes valuable time as teams end up reinventing metrics or seeking clarification on definitions that should be standardized and readily accessible.</p><p id="53c3" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Enter <a class="af oy" href="https://datajunction.io/" rel="noopener ugc nofollow" target="_blank">DataJunction</a> (DJ). DJ acts as a central store where metric definitions can live and evolve. Once a metric owner has registered a metric into DJ, metric consumers throughout the organization can apply that same metric definition to a set of filtered records and aggregate to any dimensional grain.</p><p id="10ac" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">As an example, imagine an analyst wanting to create a “Total Streaming Hours” metric. To add this metric to DJ, they need to provide two pieces of information:</p><ul class=""><li id="01ce" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt oz pa pb bk">The fact table that the metric comes from:</li></ul><p id="c29e" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">SELECT<br /> account_id, country_iso_code, streaming_hours<br />FROM streaming_fact_table</p><ul class=""><li id="3e6e" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt oz pa pb bk">The metric expression:</li></ul><p id="aef3" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">`SUM(streaming_hours)`</p><p id="8781" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Then metric consumers throughout the organization can call DJ to request either the SQL or the resulting data. For example,</p><ul class=""><li id="b9e6" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt oz pa pb bk">total_streaming_hours of each account:</li></ul><p id="ca2b" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">dj.sql(metrics=[“total_streaming_hours”], dimensions=[“account_id”]))</p><ul class=""><li id="6782" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt oz pa pb bk">total_streaming_hours of each country:</li></ul><p id="d234" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">dj.sql(metrics=[“total_streaming_hours”], dimensions=[“country_iso_code”]))</p><ul class=""><li id="bda2" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt oz pa pb bk">total_streaming_hours of each account in the US:</li></ul><p id="2dfd" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">dj.sql(metrics=[“total_streaming_hours”], dimensions=[“country_iso_code”], filters=[“country_iso_code = ‘US’”]))</p><p id="425c" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">The key here is that DJ can perform the dimensional join on users’ behalf. If country_iso_code doesn’t already exist in the fact table, the metric owner only needs to tell DJ that account_id is the foreign key to an `users_dimension_table` (we call this process “<a class="af oy" href="https://datajunction.io/docs/0.1.0/data-modeling/dimension-links/" rel="noopener ugc nofollow" target="_blank">dimension linking</a>”). DJ then can perform the joins to bring in any requested dimensions from `users_dimension_table`.</p><p id="123b" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">The Netflix Experimentation Platform heavily leverages this feature today by treating cell assignment as just another dimension that it asks DJ to bring in. For example, to compare the average streaming hours in cell A vs cell B, the Experimentation Platform relies on DJ to bring in “cell_assignment” as a user’s dimension (no different from country_iso_code). A metric can therefore be defined once in DJ and be made available across analytics dashboards and experimentation analysis.</p><p id="eb33" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">DJ has a strong pedigree–there are several prior <a class="af oy" href="https://benn.substack.com/p/bi-by-another-name" rel="noopener ugc nofollow" target="_blank">semantic layers</a> in the industry (e.g. <a class="af oy" href="https://medium.com/airbnb-engineering/how-airbnb-achieved-metric-consistency-at-scale-f23cc53dea70" rel="noopener">Minerva</a> at Airbnb; dbt Transform, Looker, and AtScale as paid solutions). DJ stands out as an <a class="af oy" href="https://github.com/DataJunction/dj" rel="noopener ugc nofollow" target="_blank">open source</a> solution that is actively developed and stress-tested at Netflix. We’d love to see DJ easing <em class="nu">your</em> metric creation and consumption pain points!</p><h1 id="443f" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">LORE: How we’re democratizing analytics at Netflix</h1><p id="4cc6" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk"><a class="af oy" href="https://www.linkedin.com/in/apurvakansara/" rel="noopener ugc nofollow" target="_blank">Apurva Kansara</a></p><p id="19be" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">At Netflix, we rely on data and analytics to inform critical business decisions. Over time, this has resulted in large numbers of dashboard products. While such analytics products are tremendously useful, we noticed a few trends:</p><ol class=""><li id="6777" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt pc pa pb bk">A large portion of such products have less than 5 MAU (monthly active users)</li><li id="abdb" class="mw mx gu my b mz pd nb nc nd pe nf ng nh pf nj nk nl pg nn no np ph nr ns nt pc pa pb bk">We spend a tremendous amount of time building and maintaining business metrics and dimensions</li><li id="6936" class="mw mx gu my b mz pd nb nc nd pe nf ng nh pf nj nk nl pg nn no np ph nr ns nt pc pa pb bk">We see inconsistencies in how a particular metric is calculated, presented, and maintained across the Data &amp; Insights organization.</li><li id="2ab5" class="mw mx gu my b mz pd nb nc nd pe nf ng nh pf nj nk nl pg nn no np ph nr ns nt pc pa pb bk">It is challenging to scale such bespoke solutions to ever-changing and increasingly complex business needs.</li></ol><p id="6960" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Analytics Enablement is a collection of initiatives across Data &amp; Insights all focused on empowering Netflix analytic practitioners to efficiently produce and effectively deliver high-quality, actionable insights.</p><p id="cae8" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Specifically, these initiatives are focused on enabling analytics rather than on the activities that produce analytics (e.g., dashboarding, analysis, research, etc.).</p><figure class="pl pm pn po pp pq pi pj paragraph-image"><div class="pi pj pk"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*gUgNHuu6yqKdfbgg%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*gUgNHuu6yqKdfbgg%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*gUgNHuu6yqKdfbgg%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*gUgNHuu6yqKdfbgg%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*gUgNHuu6yqKdfbgg%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*gUgNHuu6yqKdfbgg%201100w,%20https://miro.medium.com/v2/resize:fit:1250/format:webp/0*gUgNHuu6yqKdfbgg%201250w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 625px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*gUgNHuu6yqKdfbgg 640w, https://miro.medium.com/v2/resize:fit:720/0*gUgNHuu6yqKdfbgg 720w, https://miro.medium.com/v2/resize:fit:750/0*gUgNHuu6yqKdfbgg 750w, https://miro.medium.com/v2/resize:fit:786/0*gUgNHuu6yqKdfbgg 786w, https://miro.medium.com/v2/resize:fit:828/0*gUgNHuu6yqKdfbgg 828w, https://miro.medium.com/v2/resize:fit:1100/0*gUgNHuu6yqKdfbgg 1100w, https://miro.medium.com/v2/resize:fit:1250/0*gUgNHuu6yqKdfbgg 1250w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 625px" /><img alt="" class="bh md pr c" width="625" height="423" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></figure><p id="87d0" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">As part of broad analytics enablement across all business domains, we invested in a chatbot to provide real insights to our end users using the power of LLM. One reason LLMs are well suited for such problems is that they tie the versatility of natural language with the power of data query to enable our business users to query data that would otherwise require sophisticated knowledge of underlying data models.</p><p id="9c03" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Besides providing the end user with an instant answer in a preferred data visualization, LORE instantly learns from the user’s feedback. This allows us to teach LLM a context-rich understanding of internal business metrics that were previously locked in custom code for each of the dashboard products.</p><figure class="pl pm pn po pp pq pi pj paragraph-image"><div role="button" tabindex="0" class="pt pu fj pv bh pw"><div class="pi pj ps"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*onXkeBFPL44KYBQB%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*onXkeBFPL44KYBQB%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*onXkeBFPL44KYBQB%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*onXkeBFPL44KYBQB%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*onXkeBFPL44KYBQB%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*onXkeBFPL44KYBQB%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*onXkeBFPL44KYBQB%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*onXkeBFPL44KYBQB 640w, https://miro.medium.com/v2/resize:fit:720/0*onXkeBFPL44KYBQB 720w, https://miro.medium.com/v2/resize:fit:750/0*onXkeBFPL44KYBQB 750w, https://miro.medium.com/v2/resize:fit:786/0*onXkeBFPL44KYBQB 786w, https://miro.medium.com/v2/resize:fit:828/0*onXkeBFPL44KYBQB 828w, https://miro.medium.com/v2/resize:fit:1100/0*onXkeBFPL44KYBQB 1100w, https://miro.medium.com/v2/resize:fit:1400/0*onXkeBFPL44KYBQB 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pr c" width="700" height="191" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="5bc7" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Some of the challenges we run into:</p><ul class=""><li id="7232" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt oz pa pb bk">Gaining user trust: To gain our end users’ trust, we focused on our model’s explainability. For example, LORE provides human-readable reasoning on how it arrived at the answer that users can cross-verify. LORE also provides a confidence score to our end users based on its grounding in the domain space.</li><li id="1cdb" class="mw mx gu my b mz pd nb nc nd pe nf ng nh pf nj nk nl pg nn no np ph nr ns nt oz pa pb bk">Training: We created easy-to-provide feedback using 👍 and 👎 with a fully integrated fine-tuning loop to allow end-users to teach new domains and questions around it effectively. This allowed us to bootstrap LORE across several domains within Netflix.</li></ul><p id="4d54" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Democratizing analytics can unlock the tremendous potential of data for everyone within the company. With Analytics enablement and LORE, we’ve enabled our business users to truly have a conversation with the data.</p><h1 id="cfe1" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Leveraging Foundational Platform Data to enable Cloud Efficiency Analytics</h1><p id="50f8" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk"><a class="af oy" href="https://www.linkedin.com/in/jhan-104105/?utm_source=share&amp;utm_campaign=share_via&amp;utm_content=profile" rel="noopener ugc nofollow" target="_blank">J Han</a>, <a class="af oy" href="https://www.linkedin.com/in/pallavi-phadnis-75280b20/" rel="noopener ugc nofollow" target="_blank">Pallavi Phadnis</a></p><p id="888c" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">At Netflix, we use Amazon Web Services (AWS) for our cloud infrastructure needs, such as compute, storage, and networking to build and run the streaming platform that we love. Our ecosystem enables engineering teams to run applications and services at scale, utilizing a mix of open-source and proprietary solutions. In order to understand how efficiently we operate in this diverse technological landscape, the Data &amp; Insights organization partners closely with our engineering teams to share key efficiency metrics, empowering internal stakeholders to make informed business decisions.</p><p id="c749" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">This is where our team, Platform DSE (Data Science Engineering), comes in to enable our engineering partners to understand what resources they’re using, how effectively they utilize those resources, and the cost associated with their resource usage. By creating curated datasets and democratizing access via a custom insights app and various integration points, downstream users can gain granular insights essential for making data-driven, cost-effective decisions for the business.</p><p id="6327" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">To address the numerous analytic needs in a scalable way, we’ve developed a two-component solution:</p><ol class=""><li id="3423" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt pc pa pb bk">Foundational Platform Data (FPD): This component provides a centralized data layer for all platform data, featuring a consistent data model and standardized data processing methodology. We work with different platform data providers to get <em class="nu">inventory</em>, <em class="nu">ownership</em>, and <em class="nu">usage</em> data for the respective platforms they own.</li><li id="e86e" class="mw mx gu my b mz pd nb nc nd pe nf ng nh pf nj nk nl pg nn no np ph nr ns nt pc pa pb bk">Cloud Efficiency Analytics (CEA): Built on top of FPD, this component offers an analytics data layer that provides time series efficiency metrics across various business use cases. Once the foundational data is ready, CEA consumes inventory, ownership, and usage data and applies the appropriate <em class="nu">business logic</em> to produce <em class="nu">cost</em> and <em class="nu">ownership attribution</em> at various granularities.</li></ol><p id="f9d1" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">As the source of truth for efficiency metrics, our team’s tenants are to provide accurate, reliable, and accessible data, comprehensive documentation to navigate the complexity of the efficiency space, and well-defined Service Level Agreements (SLAs) to set expectations with downstream consumers during delays, outages, or changes.</p><p id="fba8" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Looking ahead, we aim to continue onboarding platforms, striving for nearly complete cost insight coverage. We’re also exploring new use cases, such as tailored reports for platforms, predictive analytics for optimizing usage and detecting anomalies in cost, and a root cause analysis tool using LLMs.</p><p id="b0a0" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Ultimately, our goal is to enable our engineering organization to make efficiency-conscious decisions when building and maintaining the myriad of services that allows us to enjoy Netflix as a streaming service. For more detail on our modeling approach and principles, check out <a class="af oy" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/cloud-efficiency-at-netflix-f2a142955f83">this post</a>!</p></div></div></div><div class="ab cb px py pz qa" role="separator"><div class="gn go gp gq gr"><div class="ab cb"><div class="ci bh fz ga gb gc"><p id="effc" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Analytics Engineering is a key contributor to building our deep data culture at Netflix, and we are proud to have a large group of stunning colleagues that are not only applying but advancing our analytical capabilities at Netflix. The 2024 Analytics Summit continued to be a wonderful way to give visibility to one another on work across business verticals, celebrate our collective impact, and highlight what’s to come in analytics practice at Netflix.</p><p id="faa3" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">To learn more, follow the <a class="af oy" href="https://research.netflix.com/research-area/analytics" rel="noopener ugc nofollow" target="_blank">Netflix Research Site</a>, and if you are also interested in entertaining the world, have a look at <a class="af oy" href="https://explore.jobs.netflix.net/careers" rel="noopener ugc nofollow" target="_blank">our open roles</a>!</p></div></div></div></div></div>]]></description>
      <link>https://netflixtechblog.com/part-1-a-survey-of-analytics-engineering-work-at-netflix-d761cfd551ee</link>
      <guid>https://netflixtechblog.com/part-1-a-survey-of-analytics-engineering-work-at-netflix-d761cfd551ee</guid>
      <pubDate>Wed, 18 Dec 2024 00:26:00 +0100</pubDate>
    </item>
    <item>
      <title><![CDATA[Cloud Efficiency at Netflix]]></title>
      <description><![CDATA[<div><div></div><p id="c997" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">By</strong> <a class="af nu" href="https://www.linkedin.com/in/jhan-104105?utm_source=share&amp;utm_campaign=share_via&amp;utm_content=profile" rel="noopener ugc nofollow" target="_blank">J Han</a>, <a class="af nu" href="https://www.linkedin.com/in/pallavi-phadnis-75280b20/" rel="noopener ugc nofollow" target="_blank">Pallavi Phadnis</a></p><h1 id="16b4" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk"><strong class="al">Context</strong></h1><p id="f453" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">At Netflix, we use Amazon Web Services (AWS) for our cloud infrastructure needs, such as compute, storage, and networking to build and run the streaming platform that we love. Our ecosystem enables engineering teams to run applications and services at scale, utilizing a mix of open-source and proprietary solutions. In turn, our self-serve platforms allow teams to create and deploy, sometimes custom, workloads more efficiently. This diverse technological landscape generates extensive and rich data from various infrastructure entities, from which, data engineers and analysts collaborate to provide actionable insights to the engineering organization in a continuous feedback loop that ultimately enhances the business.</p><p id="1ac0" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">One crucial way in which we do this is through the democratization of highly curated data sources that sunshine usage and cost patterns across Netflix’s services and teams. The Data &amp; Insights organization partners closely with our engineering teams to share key efficiency metrics, empowering internal stakeholders to make informed business decisions.</p><h1 id="a744" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk"><strong class="al">Data is Key</strong></h1><p id="a1ec" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">This is where our team, Platform DSE (Data Science Engineering), comes in to enable our engineering partners to understand what resources they’re using, how effectively and efficiently they use those resources, and the cost associated with their resource usage. We want our downstream consumers to make cost conscious decisions using our datasets.</p><p id="76e1" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">To address these numerous analytic needs in a scalable way, we’ve developed a two-component solution:</p><ol class=""><li id="c691" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt oy oz pa bk">Foundational Platform Data (FPD): This component provides a centralized data layer for all platform data, featuring a consistent data model and standardized data processing methodology.</li><li id="5b27" class="mw mx gu my b mz pb nb nc nd pc nf ng nh pd nj nk nl pe nn no np pf nr ns nt oy oz pa bk">Cloud Efficiency Analytics (CEA): Built on top of FPD, this component offers an analytics data layer that provides time series efficiency metrics across various business use cases.</li></ol><figure class="pj pk pl pm pn po pg ph paragraph-image"><div role="button" tabindex="0" class="pp pq fj pr bh ps"><div class="pg ph pi"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*vDQJiJUttlRSpVBo%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*vDQJiJUttlRSpVBo%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*vDQJiJUttlRSpVBo%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*vDQJiJUttlRSpVBo%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*vDQJiJUttlRSpVBo%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*vDQJiJUttlRSpVBo%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*vDQJiJUttlRSpVBo%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*vDQJiJUttlRSpVBo 640w, https://miro.medium.com/v2/resize:fit:720/0*vDQJiJUttlRSpVBo 720w, https://miro.medium.com/v2/resize:fit:750/0*vDQJiJUttlRSpVBo 750w, https://miro.medium.com/v2/resize:fit:786/0*vDQJiJUttlRSpVBo 786w, https://miro.medium.com/v2/resize:fit:828/0*vDQJiJUttlRSpVBo 828w, https://miro.medium.com/v2/resize:fit:1100/0*vDQJiJUttlRSpVBo 1100w, https://miro.medium.com/v2/resize:fit:1400/0*vDQJiJUttlRSpVBo 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pt c" width="700" height="593" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="7a9d" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Foundational Platform Data (FPD)</strong></p><p id="8685" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">We work with different platform data providers to get <em class="pu">inventory</em>, <em class="pu">ownership</em>, and <em class="pu">usage</em> data for the respective platforms they own. Below is an example of how this framework applies to the <a class="af nu" href="https://spark.apache.org/" rel="noopener ugc nofollow" target="_blank">Spark</a> platform. FPD establishes<em class="pu"> data contracts</em> with producers to ensure data quality and reliability; these contracts allow the team to leverage a common data model for ownership. The standardized data model and processing promotes scalability and consistency.</p><figure class="pj pk pl pm pn po pg ph paragraph-image"><div role="button" tabindex="0" class="pp pq fj pr bh ps"><div class="pg ph pv"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*cln5xplS7lpdE0KOh0LE1Q.jpeg" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*cln5xplS7lpdE0KOh0LE1Q.jpeg" /><img alt="" class="bh md pt c" width="700" height="135" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="a35c" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Cloud Efficiency Analytics (CEA Data)</strong></p><p id="0ba0" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Once the foundational data is ready, CEA consumes inventory, ownership, and usage data and applies the appropriate <em class="pu">business logic</em> to produce <em class="pu">cost</em> and <em class="pu">ownership attribution</em> at various granularities. The data model approach in CEA is to compartmentalize and be <em class="pu">transparent</em>; we want downstream consumers to understand why they’re seeing resources show up under their name/org and how those costs are calculated. Another benefit to this approach is the ability to pivot quickly as new or changes in business logic is/are introduced.</p><figure class="pj pk pl pm pn po pg ph paragraph-image"><div role="button" tabindex="0" class="pp pq fj pr bh ps"><div class="pg ph pw"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*bvD7xqAO9T9m4s4G%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*bvD7xqAO9T9m4s4G%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*bvD7xqAO9T9m4s4G%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*bvD7xqAO9T9m4s4G%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*bvD7xqAO9T9m4s4G%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*bvD7xqAO9T9m4s4G%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*bvD7xqAO9T9m4s4G%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*bvD7xqAO9T9m4s4G 640w, https://miro.medium.com/v2/resize:fit:720/0*bvD7xqAO9T9m4s4G 720w, https://miro.medium.com/v2/resize:fit:750/0*bvD7xqAO9T9m4s4G 750w, https://miro.medium.com/v2/resize:fit:786/0*bvD7xqAO9T9m4s4G 786w, https://miro.medium.com/v2/resize:fit:828/0*bvD7xqAO9T9m4s4G 828w, https://miro.medium.com/v2/resize:fit:1100/0*bvD7xqAO9T9m4s4G 1100w, https://miro.medium.com/v2/resize:fit:1400/0*bvD7xqAO9T9m4s4G 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pt c" width="700" height="351" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="fcca" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">* For cost accounting purposes, we resolve assets to a single owner, or distribute costs when assets are multi-tenant. However, we do also provide usage and cost at different aggregations for different consumers.</p><h1 id="3fc1" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk"><strong class="al">Data Principles</strong></h1><p id="6a99" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">As the source of truth for efficiency metrics, our team’s tenants are to provide accurate, reliable, and accessible data, comprehensive documentation to navigate the complexity of the efficiency space, and well-defined Service Level Agreements (SLAs) to set expectations with downstream consumers during delays, outages or changes.</p><p id="d22d" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">While ownership and cost may seem straightforward, the complexity of the datasets is considerably high due to the breadth and scope of the business infrastructure and platform specific features. Services can have multiple owners, cost heuristics are unique to each platform, and the scale of infra data is large. As we work on expanding infrastructure coverage to all verticals of the business, we face a unique set of challenges:</p><p id="50e2" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">A Few Sizes to Fit the Majority</strong></p><p id="e6ac" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Despite data contracts and a standardized data model on transforming upstream platform data into FPD and CEA, there is usually some degree of customization that is unique to that particular platform. As the centralized source of truth, we feel the constant tension of where to place the processing burden. Decision-making involves ongoing transparent conversations with both our data producers and consumers, frequent prioritization checks, and alignment with business needs as <a class="af nu" href="https://jobs.netflix.com/culture" rel="noopener ugc nofollow" target="_blank">informed captains</a> in this space.</p><p id="fa84" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Data Guarantees</strong></p><p id="5d38" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">For data correctness and trust, it’s crucial that we have audits and visibility into health metrics at each layer in the pipeline in order to investigate issues and root cause anomalies quickly. Maintaining data completeness while ensuring correctness becomes challenging due to upstream latency and required transformations to have the data ready for consumption. We continuously iterate our audits and incorporate feedback to refine and meet our SLAs.</p><p id="c365" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Abstraction Layers</strong></p><p id="e5ca" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">We value <a class="af nu" href="https://jobs.netflix.com/culture" rel="noopener ugc nofollow" target="_blank">people over process</a>, and it is not uncommon for engineering teams to build custom SaaS solutions for other parts of the organization. Although this fosters innovation and improves development velocity, it can create a bit of a conundrum when it comes to understanding and interpreting usage patterns and attributing cost in a way that makes sense to the business and end consumer. With clear inventory, ownership, and usage data from FPD, and precise attribution in the analytical layer, we aim to provide metrics to downstream users regardless of whether they utilize and build on top of internal platforms or on AWS resources directly.</p><h1 id="acbb" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk"><strong class="al">Future Forward</strong></h1><p id="1f2b" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">Looking ahead, we aim to continue onboarding platforms to FPD and CEA, striving for nearly complete cost insight coverage in the upcoming year. Longer term, we plan to extend FPD to other areas of the business such as security and availability. We aim to move towards proactive approaches via predictive analytics and ML for optimizing usage and detecting anomalies in cost.</p><p id="e9c5" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Ultimately, our goal is to enable our engineering organization to make efficiency-conscious decisions when building and maintaining the myriad of services that allow us to enjoy Netflix as a streaming service.</p><h1 id="78a3" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Acknowledgments</h1><p id="2782" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">The FPD and CEA work would not have been possible without the cross functional input of many outstanding colleagues and our dedicated team building these important data assets.</p><p id="1be1" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">—</p><p id="fdd6" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">A bit about the authors:</p><p id="4e87" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><em class="pu">JHan enjoys nature, reading fantasy, and finding the best chocolate chip cookies and cinnamon rolls. She is adamant about writing the SQL select statement with leading commas.</em></p><p id="0922" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><em class="pu">Pallavi enjoys music, travel and watching astrophysics documentaries. With 15+ years working with data, she knows everything’s better with a dash of analytics and a cup of coffee!</em></p></div>]]></description>
      <link>https://netflixtechblog.com/cloud-efficiency-at-netflix-f2a142955f83</link>
      <guid>https://netflixtechblog.com/cloud-efficiency-at-netflix-f2a142955f83</guid>
      <pubDate>Tue, 17 Dec 2024 23:17:00 +0100</pubDate>
    </item>
    <item>
      <title><![CDATA[Title Launch Observability at Netflix Scale]]></title>
      <description><![CDATA[<div><div><h2 id="088d" class="pw-subtitle-paragraph hr gt gu bf b hs ht hu hv hw hx hy hz ia ib ic id ie if ig cq du">Part 1: Understanding The Challenges</h2><div></div><p id="51e3" class="pw-post-body-paragraph ni nj gu nk b hs nl nm nn hv no np nq nr ns nt nu nv nw nx ny nz oa ob oc od gn bk"><strong class="nk gv">By:</strong> <a class="af oe" href="https://www.linkedin.com/in/varun-khaitan/" rel="noopener ugc nofollow" target="_blank">Varun Khaitan</a></p><p id="72bf" class="pw-post-body-paragraph ni nj gu nk b hs nl nm nn hv no np nq nr ns nt nu nv nw nx ny nz oa ob oc od gn bk">With special thanks to my stunning colleagues: <a class="af oe" href="https://www.linkedin.com/in/mallikarao/" rel="noopener ugc nofollow" target="_blank">Mallika Rao</a>, <a class="af oe" href="https://www.linkedin.com/in/esmir-mesic/" rel="noopener ugc nofollow" target="_blank">Esmir Mesic</a>, <a class="af oe" href="https://www.linkedin.com/in/hugodesmarques/" rel="noopener ugc nofollow" target="_blank">Hugo Marques</a></p><h1 id="c3f4" class="of og gu bf oh oi oj hu ok ol om hx on oo op oq or os ot ou ov ow ox oy oz pa bk">Introduction</h1><p id="0a98" class="pw-post-body-paragraph ni nj gu nk b hs pb nm nn hv pc np nq nr pd nt nu nv pe nx ny nz pf ob oc od gn bk">At Netflix, we manage over a thousand global content launches each month, backed by billions of dollars in annual investment. Ensuring the success and discoverability of each title across our platform is a top priority, as we aim to connect every story with the right audience to delight our members. To achieve this, we are committed to building robust systems that deliver comprehensive observability, enabling us to take full accountability for every title on our service.</p><h1 id="b58e" class="of og gu bf oh oi oj hu ok ol om hx on oo op oq or os ot ou ov ow ox oy oz pa bk">The Challenge of Title Launch Observability</h1><p id="e5ac" class="pw-post-body-paragraph ni nj gu nk b hs pb nm nn hv pc np nq nr pd nt nu nv pe nx ny nz pf ob oc od gn bk">As engineers, we’re wired to track system metrics like error rates, latencies, and CPU utilization — but what about metrics that matter to a title’s success?</p><p id="e27e" class="pw-post-body-paragraph ni nj gu nk b hs nl nm nn hv no np nq nr ns nt nu nv nw nx ny nz oa ob oc od gn bk">Consider the following example of two different Netflix Homepages:</p><figure class="pj pk pl pm pn po pg ph paragraph-image"><div role="button" tabindex="0" class="pp pq fj pr bh ps"><div class="pg ph pi"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*B4iyOBZJZEo7eW-p%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*B4iyOBZJZEo7eW-p%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*B4iyOBZJZEo7eW-p%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*B4iyOBZJZEo7eW-p%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*B4iyOBZJZEo7eW-p%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*B4iyOBZJZEo7eW-p%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*B4iyOBZJZEo7eW-p%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*B4iyOBZJZEo7eW-p 640w, https://miro.medium.com/v2/resize:fit:720/0*B4iyOBZJZEo7eW-p 720w, https://miro.medium.com/v2/resize:fit:750/0*B4iyOBZJZEo7eW-p 750w, https://miro.medium.com/v2/resize:fit:786/0*B4iyOBZJZEo7eW-p 786w, https://miro.medium.com/v2/resize:fit:828/0*B4iyOBZJZEo7eW-p 828w, https://miro.medium.com/v2/resize:fit:1100/0*B4iyOBZJZEo7eW-p 1100w, https://miro.medium.com/v2/resize:fit:1400/0*B4iyOBZJZEo7eW-p 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh mp pt c" width="700" height="382" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="pu ff pv pg ph pw px bf b bg z du">Sample Homepage A</figcaption></figure><figure class="pj pk pl pm pn po pg ph paragraph-image"><div role="button" tabindex="0" class="pp pq fj pr bh ps"><div class="pg ph pi"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*5F9ATQbyOp99jMwJ%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*5F9ATQbyOp99jMwJ%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*5F9ATQbyOp99jMwJ%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*5F9ATQbyOp99jMwJ%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*5F9ATQbyOp99jMwJ%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*5F9ATQbyOp99jMwJ%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*5F9ATQbyOp99jMwJ%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*5F9ATQbyOp99jMwJ 640w, https://miro.medium.com/v2/resize:fit:720/0*5F9ATQbyOp99jMwJ 720w, https://miro.medium.com/v2/resize:fit:750/0*5F9ATQbyOp99jMwJ 750w, https://miro.medium.com/v2/resize:fit:786/0*5F9ATQbyOp99jMwJ 786w, https://miro.medium.com/v2/resize:fit:828/0*5F9ATQbyOp99jMwJ 828w, https://miro.medium.com/v2/resize:fit:1100/0*5F9ATQbyOp99jMwJ 1100w, https://miro.medium.com/v2/resize:fit:1400/0*5F9ATQbyOp99jMwJ 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh mp pt c" width="700" height="386" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="pu ff pv pg ph pw px bf b bg z du">Sample Homepage B</figcaption></figure><p id="f8a6" class="pw-post-body-paragraph ni nj gu nk b hs nl nm nn hv no np nq nr ns nt nu nv nw nx ny nz oa ob oc od gn bk">To a basic recommendation system, the two sample pages might appear equivalent as long as the viewer watches the top title. Yet, these pages couldn’t be more different. Each title represents countless hours of effort and creativity, and our systems need to honor that uniqueness.</p><p id="ea38" class="pw-post-body-paragraph ni nj gu nk b hs nl nm nn hv no np nq nr ns nt nu nv nw nx ny nz oa ob oc od gn bk">How do we bridge this gap? How can we design systems that recognize these nuances and empower every title to shine and bring joy to our members?</p><h1 id="8bf0" class="of og gu bf oh oi oj hu ok ol om hx on oo op oq or os ot ou ov ow ox oy oz pa bk">The operational needs of a personalization system</h1><p id="931f" class="pw-post-body-paragraph ni nj gu nk b hs pb nm nn hv pc np nq nr pd nt nu nv pe nx ny nz pf ob oc od gn bk">In the early days of Netflix Originals, our launch team would huddle together at midnight, manually verifying that titles appeared in all the right places. While this hands-on approach worked for a handful of titles, it quickly became clear that it couldn’t scale. As Netflix expanded globally and the volume of title launches skyrocketed, the operational challenges of maintaining this manual process became undeniable.</p><p id="d7bd" class="pw-post-body-paragraph ni nj gu nk b hs nl nm nn hv no np nq nr ns nt nu nv nw nx ny nz oa ob oc od gn bk">Operating a personalization system for a global streaming service involves addressing numerous inquiries about why certain titles appear or fail to appear at specific times and places. <br />Some examples:</p><ul class=""><li id="8d25" class="ni nj gu nk b hs nl nm nn hv no np nq nr ns nt nu nv nw nx ny nz oa ob oc od py pz qa bk">Why is title X not showing on the Coming Soon row for a particular member?</li><li id="a0cc" class="ni nj gu nk b hs qb nm nn hv qc np nq nr qd nt nu nv qe nx ny nz qf ob oc od py pz qa bk">Why is title Y missing from the search page in Brazil?</li><li id="f5d7" class="ni nj gu nk b hs qb nm nn hv qc np nq nr qd nt nu nv qe nx ny nz qf ob oc od py pz qa bk">Is title Z being displayed correctly in all product experiences as intended?</li></ul><p id="83a3" class="pw-post-body-paragraph ni nj gu nk b hs nl nm nn hv no np nq nr ns nt nu nv nw nx ny nz oa ob oc od gn bk">As Netflix scaled, we faced the mounting challenge of providing accurate, timely answers to increasingly complex queries about title performance and discoverability. This led to a suite of fragmented scripts, runbooks, and ad hoc solutions scattered across teams — an approach that was neither sustainable nor efficient.</p><p id="b860" class="pw-post-body-paragraph ni nj gu nk b hs nl nm nn hv no np nq nr ns nt nu nv nw nx ny nz oa ob oc od gn bk">The stakes are even higher when ensuring every title launches flawlessly. Metadata and assets must be correctly configured, data must flow seamlessly, microservices must process titles without error, and algorithms must function as intended. The complexity of these operational demands underscored the urgent need for a scalable solution.</p><h1 id="fc77" class="of og gu bf oh oi oj hu ok ol om hx on oo op oq or os ot ou ov ow ox oy oz pa bk">Automating the Operations</h1><p id="2817" class="pw-post-body-paragraph ni nj gu nk b hs pb nm nn hv pc np nq nr pd nt nu nv pe nx ny nz pf ob oc od gn bk">It becomes evident over time that we need to automate our operations to scale with the business. As we thought more about this problem and possible solutions, two clear options emerged.</p><h1 id="26ce" class="of og gu bf oh oi oj hu ok ol om hx on oo op oq or os ot ou ov ow ox oy oz pa bk">Option 1: Log Processing</h1><p id="be78" class="pw-post-body-paragraph ni nj gu nk b hs pb nm nn hv pc np nq nr pd nt nu nv pe nx ny nz pf ob oc od gn bk">Log processing offers a straightforward solution for monitoring and analyzing title launches. By logging all titles as they are displayed, we can process these logs to identify anomalies and gain insights into system performance. This approach provides a few advantages:</p><ol class=""><li id="c8ff" class="ni nj gu nk b hs nl nm nn hv no np nq nr ns nt nu nv nw nx ny nz oa ob oc od qg pz qa bk"><strong class="nk gv">Low burden on existing systems:</strong> Log processing imposes minimal changes to existing infrastructure. By leveraging logs, which are already generated during regular operations, we can scale observability without significant system modifications. This allows us to focus on data analysis and problem-solving rather than managing complex system changes.</li><li id="9430" class="ni nj gu nk b hs qb nm nn hv qc np nq nr qd nt nu nv qe nx ny nz qf ob oc od qg pz qa bk"><strong class="nk gv">Using the source of truth:</strong> Logs serve as a reliable “source of truth” by providing a comprehensive record of system events. They allow us to verify whether titles are presented as intended and investigate any discrepancies. This capability is crucial for ensuring our recommendation systems and user interfaces function correctly, supporting successful title launches.</li></ol><p id="6c94" class="pw-post-body-paragraph ni nj gu nk b hs nl nm nn hv no np nq nr ns nt nu nv nw nx ny nz oa ob oc od gn bk">However, taking this approach also presents several challenges:</p><ol class=""><li id="ad3d" class="ni nj gu nk b hs nl nm nn hv no np nq nr ns nt nu nv nw nx ny nz oa ob oc od qg pz qa bk"><strong class="nk gv">Catching Issues Ahead of Time:</strong> Logging primarily addresses post-launch scenarios, as logs are generated only after titles are shown to members. To detect issues proactively, we need to simulate traffic and predict system behavior in advance. Once artificial traffic is generated, discarding the response object and relying solely on logs becomes inefficient.</li><li id="3d21" class="ni nj gu nk b hs qb nm nn hv qc np nq nr qd nt nu nv qe nx ny nz qf ob oc od qg pz qa bk"><strong class="nk gv">Appropriate Accuracy:</strong> Comprehensive logging requires services to log both included and excluded titles, along with reasons for exclusion. This could lead to an exponential increase in logged data. Utilizing probabilistic logging methods could compromise accuracy, making it difficult to ascertain whether a title’s absence in logs is due to exclusion or random chance.</li><li id="289c" class="ni nj gu nk b hs qb nm nn hv qc np nq nr qd nt nu nv qe nx ny nz qf ob oc od qg pz qa bk"><strong class="nk gv">SLA and Cost Considerations:</strong> Our existing online logging systems do not natively support logging at the title granularity level. While reengineering these systems to accommodate this additional axis is possible, it would entail increased costs. Additionally, the time-sensitive nature of these investigations precludes the use of cold storage, which cannot meet the stringent SLAs required.</li></ol><h1 id="aac6" class="of og gu bf oh oi oj hu ok ol om hx on oo op oq or os ot ou ov ow ox oy oz pa bk">Option 2: Observability Endpoints in Our Personalization Systems</h1><p id="7199" class="pw-post-body-paragraph ni nj gu nk b hs pb nm nn hv pc np nq nr pd nt nu nv pe nx ny nz pf ob oc od gn bk">To prioritize title launch observability, we could adopt a centralized approach. By introducing observability endpoints across all systems, we can enable real-time data flow into a dedicated microservice for title launch observability. This approach embeds observability directly into the very fabric of services managing title launches and personalization, ensuring seamless monitoring and insights. Key benefits and strategies include:</p><ol class=""><li id="6ecb" class="ni nj gu nk b hs nl nm nn hv no np nq nr ns nt nu nv nw nx ny nz oa ob oc od qg pz qa bk"><strong class="nk gv">Real-Time Monitoring: </strong>Observability endpoints enable real-time monitoring of system performance and title placements, allowing us to detect and address issues as they arise.</li><li id="f705" class="ni nj gu nk b hs qb nm nn hv qc np nq nr qd nt nu nv qe nx ny nz qf ob oc od qg pz qa bk"><strong class="nk gv">Proactive Issue Detection: </strong>By simulating future traffic(an aspect we call “time travel”) and capturing system responses ahead of time, we can preemptively identify potential issues before they impact our members or the business.</li><li id="f16a" class="ni nj gu nk b hs qb nm nn hv qc np nq nr qd nt nu nv qe nx ny nz qf ob oc od qg pz qa bk"><strong class="nk gv">Enhanced Accuracy:</strong> Observability endpoints provide precise data on title inclusions and exclusions, allowing us to make accurate assertions about system behavior and title visibility. It also provides us with advanced debugability information needed to fix identified issues.</li><li id="e717" class="ni nj gu nk b hs qb nm nn hv qc np nq nr qd nt nu nv qe nx ny nz qf ob oc od qg pz qa bk"><strong class="nk gv">Scalability and Cost Efficiency:</strong> While initial implementation required some investment, this approach ultimately offers a scalable and cost-effective solution to managing title launches at Netflix scale.</li></ol><p id="cfd1" class="pw-post-body-paragraph ni nj gu nk b hs nl nm nn hv no np nq nr ns nt nu nv nw nx ny nz oa ob oc od gn bk">Choosing this option also comes with some tradeoffs:</p><ol class=""><li id="b2fe" class="ni nj gu nk b hs nl nm nn hv no np nq nr ns nt nu nv nw nx ny nz oa ob oc od qg pz qa bk"><strong class="nk gv">Significant Initial Investment: </strong>Several systems would need to create new endpoints and refactor their codebases to adopt this new method of prioritizing launches.</li><li id="1a2c" class="ni nj gu nk b hs qb nm nn hv qc np nq nr qd nt nu nv qe nx ny nz qf ob oc od qg pz qa bk"><strong class="nk gv">Synchronization Risk: </strong>There would be a potential risk that these new endpoints may not accurately represent production behavior, thus necessitating conscious efforts to ensure all endpoints remain synchronized.</li></ol><h1 id="9e33" class="of og gu bf oh oi oj hu ok ol om hx on oo op oq or os ot ou ov ow ox oy oz pa bk">Up Next</h1><p id="1466" class="pw-post-body-paragraph ni nj gu nk b hs pb nm nn hv pc np nq nr pd nt nu nv pe nx ny nz pf ob oc od gn bk">By adopting a comprehensive observability strategy that includes real-time monitoring, proactive issue detection, and source of truth reconciliation, we’ve significantly enhanced our ability to ensure the successful launch and discovery of titles across Netflix, enriching the global viewing experience for our members. In the next part of this series, we’ll dive into how we achieved this, sharing key technical insights and details.</p><p id="3bce" class="pw-post-body-paragraph ni nj gu nk b hs nl nm nn hv no np nq nr ns nt nu nv nw nx ny nz oa ob oc od gn bk">Stay tuned for a closer look at the innovation behind the scenes!</p></div></div>]]></description>
      <link>https://netflixtechblog.com/title-launch-observability-at-netflix-scale-c88c586629eb</link>
      <guid>https://netflixtechblog.com/title-launch-observability-at-netflix-scale-c88c586629eb</guid>
      <pubDate>Tue, 17 Dec 2024 22:54:00 +0100</pubDate>
    </item>
    <item>
      <title><![CDATA[Netflix’s Distributed Counter Abstraction]]></title>
      <description><![CDATA[<div><div></div><p id="f7ea" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">By: <a class="af nu" href="https://www.linkedin.com/in/rajiv-shringi/" rel="noopener ugc nofollow" target="_blank">Rajiv Shringi</a>, <a class="af nu" href="https://www.linkedin.com/in/oleksii-tkachuk-98b47375/" rel="noopener ugc nofollow" target="_blank">Oleksii Tkachuk</a>, <a class="af nu" href="https://www.linkedin.com/in/kartik894/" rel="noopener ugc nofollow" target="_blank">Kartik Sathyanarayanan</a></p><h1 id="0da9" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Introduction</h1><p id="41fb" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">In our previous blog post, we introduced <a class="af nu" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/introducing-netflix-timeseries-data-abstraction-layer-31552f6326f8">Netflix’s TimeSeries Abstraction</a>, a distributed service designed to store and query large volumes of temporal event data with low millisecond latencies. Today, we’re excited to present the <strong class="my gv">Distributed Counter Abstraction</strong>. This counting service, built on top of the TimeSeries Abstraction, enables distributed counting at scale while maintaining similar low latency performance. As with all our abstractions, we use our <a class="af nu" href="https://netflixtechblog.medium.com/data-gateway-a-platform-for-growing-and-protecting-the-data-tier-f1ed8db8f5c6" rel="noopener">Data Gateway Control Plane</a> to shard, configure, and deploy this service globally.</p><p id="aebc" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Distributed counting is a challenging problem in computer science. In this blog post, we’ll explore the diverse counting requirements at Netflix, the challenges of achieving accurate counts in near real-time, and the rationale behind our chosen approach, including the necessary trade-offs.</p><p id="fb3c" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Note</strong>: <em class="oy">When it comes to distributed counters, terms such as ‘accurate’ or ‘precise’ should be taken with a grain of salt. In this context, they refer to a count very close to accurate, presented with minimal delays.</em></p><h1 id="21f6" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Use Cases and Requirements</h1><p id="1e4f" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">At Netflix, our counting use cases include tracking millions of user interactions, monitoring how often specific features or experiences are shown to users, and counting multiple facets of data during <a class="af nu" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/its-all-a-bout-testing-the-netflix-experimentation-platform-4e1ca458c15">A/B test experiments</a>, among others.</p><p id="e35a" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">At Netflix, these use cases can be classified into two broad categories:</p><ol class=""><li id="fc33" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt oz pa pb bk"><strong class="my gv">Best-Effort</strong>: For this category, the count doesn’t have to be very accurate or durable. However, this category requires near-immediate access to the current count at low latencies, all while keeping infrastructure costs to a minimum.</li><li id="d9a3" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk"><strong class="my gv">Eventually Consistent</strong>: This category needs accurate and durable counts, and is willing to tolerate a slight delay in accuracy and a slightly higher infrastructure cost as a trade-off.</li></ol><p id="7d8e" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Both categories share common requirements, such as high throughput and high availability. The table below provides a detailed overview of the diverse requirements across these two categories.</p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi pj"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*_Mx2WRBWOfASpK_e2xgoVw.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*_Mx2WRBWOfASpK_e2xgoVw.png" /><img alt="" class="bh md pu c" width="700" height="494" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><h1 id="db7c" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Distributed Counter Abstraction</h1><p id="16d7" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">To meet the outlined requirements, the Counter Abstraction was designed to be highly configurable. It allows users to choose between different counting modes, such as <strong class="my gv">Best-Effort</strong> or <strong class="my gv">Eventually Consistent</strong>, while considering the documented trade-offs of each option. After selecting a mode, users can interact with APIs without needing to worry about the underlying storage mechanisms and counting methods.</p><p id="0799" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Let’s take a closer look at the structure and functionality of the API.</p><h1 id="4626" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">API</h1><p id="0433" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">Counters are organized into separate namespaces that users set up for each of their specific use cases. Each namespace can be configured with different parameters, such as Type of Counter, Time-To-Live (TTL), and Counter Cardinality, using the service’s Control Plane.</p><p id="cc02" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">The Counter Abstraction API resembles Java’s <a class="af nu" href="https://docs.oracle.com/en/java/javase/22/docs/api/java.base/java/util/concurrent/atomic/AtomicInteger.html" rel="noopener ugc nofollow" target="_blank">AtomicInteger</a> interface:</p><p id="d3f4" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">AddCount/AddAndGetCount</strong>: Adjusts the count for the specified counter by the given delta value within a dataset. The delta value can be positive or negative. The <em class="oy">AddAndGetCount</em> counterpart also returns the count after performing the add operation.</p><pre class="pk pl pm pn po pv pw px bp py bb bk">{<br />  "namespace": "my_dataset",<br />  "counter_name": "counter123",<br />  "delta": 2,<br />  "idempotency_token": { <br />    "token": "some_event_id",<br />    "generation_time": "2024-10-05T14:48:00Z"<br />  }<br />}</pre><p id="2e22" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">The idempotency token can be used for counter types that support them. Clients can use this token to safely retry or <a class="af nu" href="https://research.google/pubs/the-tail-at-scale/" rel="noopener ugc nofollow" target="_blank">hedge</a> their requests. Failures in a distributed system are a given, and having the ability to safely retry requests enhances the reliability of the service.</p><p id="b098" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">GetCount</strong>: Retrieves the count value of the specified counter within a dataset.</p><pre class="pk pl pm pn po pv pw px bp py bb bk">{<br />  "namespace": "my_dataset",<br />  "counter_name": "counter123"<br />}</pre><p id="a50d" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">ClearCount</strong>: Effectively resets the count to 0 for the specified counter within a dataset.</p><pre class="pk pl pm pn po pv pw px bp py bb bk">{<br />  "namespace": "my_dataset",<br />  "counter_name": "counter456",<br />  "idempotency_token": {...}<br />}</pre><p id="560d" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Now, let’s look at the different types of counters supported within the Abstraction.</p><h1 id="3afc" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Types of Counters</h1><p id="0ea6" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">The service primarily supports two types of counters: <strong class="my gv">Best-Effort</strong> and <strong class="my gv">Eventually Consistent</strong>, along with a third experimental type: <strong class="my gv">Accurate</strong>. In the following sections, we’ll describe the different approaches for these types of counters and the trade-offs associated with each.</p><h1 id="1042" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Best Effort Regional Counter</h1><p id="1497" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">This type of counter is powered by <a class="af nu" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/announcing-evcache-distributed-in-memory-datastore-for-cloud-c26a698c27f7">EVCache</a>, Netflix’s distributed caching solution built on the widely popular <a class="af nu" href="https://memcached.org/" rel="noopener ugc nofollow" target="_blank">Memcached</a>. It is suitable for use cases like A/B experiments, where many concurrent experiments are run for relatively short durations and an approximate count is sufficient. Setting aside the complexities of provisioning, resource allocation, and control plane management, the core of this solution is remarkably straightforward:</p><pre class="pk pl pm pn po pv pw px bp py bb bk">// counter cache key<br />counterCacheKey = &lt;namespace&gt;:&lt;counter_name&gt;// add operation<br />return delta &gt; 0<br />    ? cache.incr(counterCacheKey, delta, TTL)<br />    : cache.decr(counterCacheKey, Math.abs(delta), TTL);// get operation<br />cache.get(counterCacheKey);// clear counts from all replicas<br />cache.delete(counterCacheKey, ReplicaPolicy.ALL);</pre><p id="70af" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">EVCache delivers extremely high throughput at low millisecond latency or better within a single region, enabling a multi-tenant setup within a shared cluster, saving infrastructure costs. However, there are some trade-offs: it lacks cross-region replication for the <em class="oy">increment</em> operation and does not provide <a class="af nu" href="https://netflix.github.io/EVCache/features/#consistency" rel="noopener ugc nofollow" target="_blank">consistency guarantees</a>, which may be necessary for an accurate count. Additionally, idempotency is not natively supported, making it unsafe to retry or hedge requests.</p><h1 id="1746" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Eventually Consistent Global Counter</h1><p id="3c43" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">While some users may accept the limitations of a Best-Effort counter, others opt for precise counts, durability and global availability. In the following sections, we’ll explore various strategies for achieving durable and accurate counts. Our objective is to highlight the challenges inherent in global distributed counting and explain the reasoning behind our chosen approach.</p><p id="e5ff" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Approach 1: Storing a Single Row per Counter</strong></p><p id="f787" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Let’s start simple by using a single row per counter key within a table in a globally replicated datastore.</p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi qe"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*X6k4-4N36IQ5yEPe%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*X6k4-4N36IQ5yEPe%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*X6k4-4N36IQ5yEPe%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*X6k4-4N36IQ5yEPe%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*X6k4-4N36IQ5yEPe%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*X6k4-4N36IQ5yEPe%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*X6k4-4N36IQ5yEPe%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*X6k4-4N36IQ5yEPe 640w, https://miro.medium.com/v2/resize:fit:720/0*X6k4-4N36IQ5yEPe 720w, https://miro.medium.com/v2/resize:fit:750/0*X6k4-4N36IQ5yEPe 750w, https://miro.medium.com/v2/resize:fit:786/0*X6k4-4N36IQ5yEPe 786w, https://miro.medium.com/v2/resize:fit:828/0*X6k4-4N36IQ5yEPe 828w, https://miro.medium.com/v2/resize:fit:1100/0*X6k4-4N36IQ5yEPe 1100w, https://miro.medium.com/v2/resize:fit:1400/0*X6k4-4N36IQ5yEPe 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="578" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="ca59" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Let’s examine some of the drawbacks of this approach:</p><ul class=""><li id="61e8" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt qf pa pb bk"><strong class="my gv">Lack of Idempotency</strong>: There is no idempotency key baked into the storage data-model preventing users from safely retrying requests. Implementing idempotency would likely require using an external system for such keys, which can further degrade performance or cause race conditions.</li><li id="2b44" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt qf pa pb bk"><strong class="my gv">Heavy Contention</strong>: To update counts reliably, every writer must perform a Compare-And-Swap operation for a given counter using locks or transactions. Depending on the throughput and concurrency of operations, this can lead to significant contention, heavily impacting performance.</li></ul><p id="a373" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Secondary Keys</strong>: One way to reduce contention in this approach would be to use a secondary key, such as a <em class="oy">bucket_id</em>, which allows for distributing writes by splitting a given counter into <em class="oy">buckets</em>, while enabling reads to aggregate across buckets. The challenge lies in determining the appropriate number of buckets. A static number may still lead to contention with <em class="oy">hot keys</em>, while dynamically assigning the number of buckets per counter across millions of counters presents a more complex problem.</p><p id="121f" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Let’s see if we can iterate on our solution to overcome these drawbacks.</p><p id="875d" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Approach 2: Per Instance Aggregation</strong></p><p id="ca5b" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">To address issues of hot keys and contention from writing to the same row in real-time, we could implement a strategy where each instance aggregates the counts in memory and then flushes them to disk at regular intervals. Introducing sufficient jitter to the flush process can further reduce contention.</p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi qg"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*6iUKbxJ093jJTiYL%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*6iUKbxJ093jJTiYL%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*6iUKbxJ093jJTiYL%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*6iUKbxJ093jJTiYL%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*6iUKbxJ093jJTiYL%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*6iUKbxJ093jJTiYL%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*6iUKbxJ093jJTiYL%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*6iUKbxJ093jJTiYL 640w, https://miro.medium.com/v2/resize:fit:720/0*6iUKbxJ093jJTiYL 720w, https://miro.medium.com/v2/resize:fit:750/0*6iUKbxJ093jJTiYL 750w, https://miro.medium.com/v2/resize:fit:786/0*6iUKbxJ093jJTiYL 786w, https://miro.medium.com/v2/resize:fit:828/0*6iUKbxJ093jJTiYL 828w, https://miro.medium.com/v2/resize:fit:1100/0*6iUKbxJ093jJTiYL 1100w, https://miro.medium.com/v2/resize:fit:1400/0*6iUKbxJ093jJTiYL 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="336" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="24b1" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">However, this solution presents a new set of issues:</p><ul class=""><li id="dba3" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt qf pa pb bk"><strong class="my gv">Vulnerability to Data Loss</strong>: The solution is vulnerable to data loss for all in-memory data during instance failures, restarts, or deployments.</li><li id="c41b" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt qf pa pb bk"><strong class="my gv">Inability to Reliably Reset Counts</strong>: Due to counting requests being distributed across multiple machines, it is challenging to establish consensus on the exact point in time when a counter reset occurred.</li><li id="2535" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt qf pa pb bk"><strong class="my gv">Lack of Idempotency: </strong>Just like the previous approach, this approach does not natively guarantee idempotency. One way to ensure idempotency is to consistently route the same set of counters to the same instance. However, such an approach may introduce additional complexity, and potential challenges with availability and latency in the write path.</li></ul><p id="038c" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">That said, this approach may still be suitable in scenarios where these trade-offs are acceptable. However, let’s see if we can address some of these issues with a different event-based approach.</p><p id="f599" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Approach 3: Using Durable Queues</strong></p><p id="7eb6" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">In this approach, we log counter events into a durable queuing system like <a class="af nu" href="https://kafka.apache.org/" rel="noopener ugc nofollow" target="_blank">Apache Kafka</a> to prevent any potential data loss. By creating multiple topic partitions and hashing the counter key to a specific partition, we ensure that the same set of counters are processed by the same set of consumers. This setup simplifies facilitating idempotency checks and resetting counts. Furthermore, by leveraging additional stream processing frameworks such as <a class="af nu" href="https://kafka.apache.org/documentation/streams/" rel="noopener ugc nofollow" target="_blank">Kafka Streams</a> or <a class="af nu" href="https://flink.apache.org/" rel="noopener ugc nofollow" target="_blank">Apache Flink</a>, we can implement windowed aggregations.</p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi qg"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*mQikuGyuzZ_lT7Y4%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*mQikuGyuzZ_lT7Y4%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*mQikuGyuzZ_lT7Y4%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*mQikuGyuzZ_lT7Y4%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*mQikuGyuzZ_lT7Y4%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*mQikuGyuzZ_lT7Y4%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*mQikuGyuzZ_lT7Y4%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*mQikuGyuzZ_lT7Y4 640w, https://miro.medium.com/v2/resize:fit:720/0*mQikuGyuzZ_lT7Y4 720w, https://miro.medium.com/v2/resize:fit:750/0*mQikuGyuzZ_lT7Y4 750w, https://miro.medium.com/v2/resize:fit:786/0*mQikuGyuzZ_lT7Y4 786w, https://miro.medium.com/v2/resize:fit:828/0*mQikuGyuzZ_lT7Y4 828w, https://miro.medium.com/v2/resize:fit:1100/0*mQikuGyuzZ_lT7Y4 1100w, https://miro.medium.com/v2/resize:fit:1400/0*mQikuGyuzZ_lT7Y4 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="437" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="25bf" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">However, this approach comes with some challenges:</p><ul class=""><li id="708e" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt qf pa pb bk"><strong class="my gv">Potential Delays</strong>: Having the same consumer process all the counts from a given partition can lead to backups and delays, resulting in stale counts.</li><li id="f448" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt qf pa pb bk"><strong class="my gv">Rebalancing Partitions</strong>: This approach requires auto-scaling and rebalancing of topic partitions as the cardinality of counters and throughput increases.</li></ul><p id="5f3d" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Furthermore, all approaches that pre-aggregate counts make it challenging to support two of our requirements for accurate counters:</p><ul class=""><li id="5818" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt qf pa pb bk"><strong class="my gv">Auditing of Counts</strong>: Auditing involves extracting data to an offline system for analysis to ensure that increments were applied correctly to reach the final value. This process can also be used to track the provenance of increments. However, auditing becomes infeasible when counts are aggregated without storing the individual increments.</li><li id="51be" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt qf pa pb bk"><strong class="my gv">Potential Recounting</strong>: Similar to auditing, if adjustments to increments are necessary and recounting of events within a time window is required, pre-aggregating counts makes this infeasible.</li></ul><p id="ccda" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Barring those few requirements, this approach can still be effective if we determine the right way to scale our queue partitions and consumers while maintaining idempotency. However, let’s explore how we can adjust this approach to meet the auditing and recounting requirements.</p><p id="83bd" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Approach 4: Event Log of Individual Increments</strong></p><p id="57ab" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">In this approach, we log each individual counter increment along with its <strong class="my gv">event_time</strong> and <strong class="my gv">event_id</strong>. The event_id can include the source information of where the increment originated. The combination of event_time and event_id can also serve as the idempotency key for the write.</p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi qh"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*0wKFK7xyTHnEKIhO%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*0wKFK7xyTHnEKIhO%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*0wKFK7xyTHnEKIhO%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*0wKFK7xyTHnEKIhO%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*0wKFK7xyTHnEKIhO%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*0wKFK7xyTHnEKIhO%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*0wKFK7xyTHnEKIhO%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*0wKFK7xyTHnEKIhO 640w, https://miro.medium.com/v2/resize:fit:720/0*0wKFK7xyTHnEKIhO 720w, https://miro.medium.com/v2/resize:fit:750/0*0wKFK7xyTHnEKIhO 750w, https://miro.medium.com/v2/resize:fit:786/0*0wKFK7xyTHnEKIhO 786w, https://miro.medium.com/v2/resize:fit:828/0*0wKFK7xyTHnEKIhO 828w, https://miro.medium.com/v2/resize:fit:1100/0*0wKFK7xyTHnEKIhO 1100w, https://miro.medium.com/v2/resize:fit:1400/0*0wKFK7xyTHnEKIhO 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="532" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="e421" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">However, <em class="oy">in its simplest form</em>, this approach has several drawbacks:</p><ul class=""><li id="1932" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt qf pa pb bk"><strong class="my gv">Read Latency</strong>: Each read request requires scanning all increments for a given counter potentially degrading performance.</li><li id="5891" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt qf pa pb bk"><strong class="my gv">Duplicate Work</strong>: Multiple threads might duplicate the effort of aggregating the same set of counters during read operations, leading to wasted effort and subpar resource utilization.</li><li id="973d" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt qf pa pb bk"><strong class="my gv">Wide Partitions</strong>: If using a datastore like <a class="af nu" href="https://cassandra.apache.org/_/index.html" rel="noopener ugc nofollow" target="_blank">Apache Cassandra</a>, storing many increments for the same counter could lead to a <a class="af nu" href="https://thelastpickle.com/blog/2019/01/11/wide-partitions-cassandra-3-11.html" rel="noopener ugc nofollow" target="_blank">wide partition</a>, affecting read performance.</li><li id="21ef" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt qf pa pb bk"><strong class="my gv">Large Data Footprint</strong>: Storing each increment individually could also result in a substantial data footprint over time. Without an efficient data retention strategy, this approach may struggle to scale effectively.</li></ul><p id="e879" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">The combined impact of these issues can lead to increased infrastructure costs that may be difficult to justify. However, adopting an event-driven approach seems to be a significant step forward in addressing some of the challenges we’ve encountered and meeting our requirements.</p><p id="04e4" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">How can we improve this solution further?</p><h1 id="08e8" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Netflix’s Approach</h1><p id="0918" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">We use a combination of the previous approaches, where we log each counting activity as an event, and continuously aggregate these events in the background using queues and a sliding time window. Additionally, we employ a bucketing strategy to prevent wide partitions. In the following sections, we’ll explore how this approach addresses the previously mentioned drawbacks and meets all our requirements.</p><p id="ff08" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Note</strong>: <em class="oy">From here on, we will use the words “</em><strong class="my gv"><em class="oy">rollup</em></strong><em class="oy">” and “</em><strong class="my gv"><em class="oy">aggregate</em></strong><em class="oy">” interchangeably. They essentially mean the same thing, i.e., collecting individual counter increments/decrements and arriving at the final value.</em></p><p id="68cd" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">TimeSeries Event Store:</strong></p><p id="aa41" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">We chose the <a class="af nu" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/introducing-netflix-timeseries-data-abstraction-layer-31552f6326f8">TimeSeries Data Abstraction</a> as our event store, where counter mutations are ingested as event records. Some of the benefits of storing events in TimeSeries include:</p><p id="11da" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">High-Performance</strong>: The TimeSeries abstraction already addresses many of our requirements, including high availability and throughput, reliable and fast performance, and more.</p><p id="b03a" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Reducing Code Complexity</strong>: We reduce a lot of code complexity in Counter Abstraction by delegating a major portion of the functionality to an existing service.</p><p id="6c3a" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">TimeSeries Abstraction uses Cassandra as the underlying event store, but it can be configured to work with any persistent store. Here is what it looks like:</p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi qg"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*ge4X7ywSmtizcNE5%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*ge4X7ywSmtizcNE5%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*ge4X7ywSmtizcNE5%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*ge4X7ywSmtizcNE5%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*ge4X7ywSmtizcNE5%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*ge4X7ywSmtizcNE5%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*ge4X7ywSmtizcNE5%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*ge4X7ywSmtizcNE5 640w, https://miro.medium.com/v2/resize:fit:720/0*ge4X7ywSmtizcNE5 720w, https://miro.medium.com/v2/resize:fit:750/0*ge4X7ywSmtizcNE5 750w, https://miro.medium.com/v2/resize:fit:786/0*ge4X7ywSmtizcNE5 786w, https://miro.medium.com/v2/resize:fit:828/0*ge4X7ywSmtizcNE5 828w, https://miro.medium.com/v2/resize:fit:1100/0*ge4X7ywSmtizcNE5 1100w, https://miro.medium.com/v2/resize:fit:1400/0*ge4X7ywSmtizcNE5 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="334" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="2b96" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Handling Wide Partitions</strong>: The <em class="oy">time_bucket</em> and <em class="oy">event_bucket</em> columns play a crucial role in breaking up a wide partition, preventing high-throughput counter events from overwhelming a given partition. <em class="oy">For more information regarding this, refer to our previous </em><a class="af nu" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/introducing-netflix-timeseries-data-abstraction-layer-31552f6326f8"><em class="oy">blog</em></a>.</p><p id="3dc8" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">No Over-Counting</strong>: The <em class="oy">event_time</em>, <em class="oy">event_id</em> and <em class="oy">event_item_key</em> columns form the idempotency key for the events for a given counter, enabling clients to retry safely without the risk of over-counting.</p><p id="43a9" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Event Ordering</strong>: TimeSeries orders all events in descending order of time allowing us to leverage this property for events like count resets.</p><p id="278b" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Event Retention</strong>: The TimeSeries Abstraction includes retention policies to ensure that events are not stored indefinitely, saving disk space and reducing infrastructure costs. Once events have been aggregated and moved to a more cost-effective store for audits, there’s no need to retain them in the primary storage.</p><p id="f647" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Now, let’s see how these events are aggregated for a given counter.</p><p id="5a6c" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Aggregating Count Events:</strong></p><p id="80b9" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">As mentioned earlier, collecting all individual increments for every read request would be cost-prohibitive in terms of read performance. Therefore, a background aggregation process is necessary to continually converge counts and ensure optimal read performance.</p><p id="2ed6" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><em class="oy">But how can we safely aggregate count events amidst ongoing write operations?</em></p><p id="0a22" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">This is where the concept of <em class="oy">Eventually Consistent </em>counts becomes crucial. <em class="oy">By intentionally lagging behind the current time by a safe margin</em>, we ensure that aggregation always occurs within an immutable window.</p><p id="5460" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Lets see what that looks like:</p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi qg"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*EOpW-VnA_YZF7KOP%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*EOpW-VnA_YZF7KOP%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*EOpW-VnA_YZF7KOP%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*EOpW-VnA_YZF7KOP%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*EOpW-VnA_YZF7KOP%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*EOpW-VnA_YZF7KOP%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*EOpW-VnA_YZF7KOP%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*EOpW-VnA_YZF7KOP 640w, https://miro.medium.com/v2/resize:fit:720/0*EOpW-VnA_YZF7KOP 720w, https://miro.medium.com/v2/resize:fit:750/0*EOpW-VnA_YZF7KOP 750w, https://miro.medium.com/v2/resize:fit:786/0*EOpW-VnA_YZF7KOP 786w, https://miro.medium.com/v2/resize:fit:828/0*EOpW-VnA_YZF7KOP 828w, https://miro.medium.com/v2/resize:fit:1100/0*EOpW-VnA_YZF7KOP 1100w, https://miro.medium.com/v2/resize:fit:1400/0*EOpW-VnA_YZF7KOP 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="470" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="a980" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Let’s break this down:</p><ul class=""><li id="c8ce" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt qf pa pb bk"><strong class="my gv">lastRollupTs</strong>: This represents the most recent time when the counter value was last aggregated. For a counter being operated for the first time, this timestamp defaults to a reasonable time in the past.</li><li id="881b" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt qf pa pb bk"><strong class="my gv">Immutable Window and Lag</strong>: Aggregation can only occur safely within an immutable window that is no longer receiving counter events. The “acceptLimit” parameter of the TimeSeries Abstraction plays a crucial role here, as it rejects incoming events with timestamps beyond this limit. During aggregations, this window is pushed slightly further back to account for clock skews.</li></ul><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi qi"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*DbtPCHPWoaauUkDr%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*DbtPCHPWoaauUkDr%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*DbtPCHPWoaauUkDr%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*DbtPCHPWoaauUkDr%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*DbtPCHPWoaauUkDr%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*DbtPCHPWoaauUkDr%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*DbtPCHPWoaauUkDr%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*DbtPCHPWoaauUkDr 640w, https://miro.medium.com/v2/resize:fit:720/0*DbtPCHPWoaauUkDr 720w, https://miro.medium.com/v2/resize:fit:750/0*DbtPCHPWoaauUkDr 750w, https://miro.medium.com/v2/resize:fit:786/0*DbtPCHPWoaauUkDr 786w, https://miro.medium.com/v2/resize:fit:828/0*DbtPCHPWoaauUkDr 828w, https://miro.medium.com/v2/resize:fit:1100/0*DbtPCHPWoaauUkDr 1100w, https://miro.medium.com/v2/resize:fit:1400/0*DbtPCHPWoaauUkDr 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="153" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="1fee" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">This does mean that the counter value will lag behind its most recent update by some margin (typically in the order of seconds). <em class="oy">This approach does leave the door open for missed events due to cross-region replication issues. See “Future Work” section at the end.</em></p><ul class=""><li id="f593" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt qf pa pb bk"><strong class="my gv">Aggregation Process</strong>: The rollup process aggregates all events in the aggregation window <em class="oy">since the last rollup </em>to arrive at the new value.</li></ul><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi qj"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*oSHneX5BOi5VNGYM%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*oSHneX5BOi5VNGYM%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*oSHneX5BOi5VNGYM%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*oSHneX5BOi5VNGYM%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*oSHneX5BOi5VNGYM%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*oSHneX5BOi5VNGYM%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*oSHneX5BOi5VNGYM%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*oSHneX5BOi5VNGYM 640w, https://miro.medium.com/v2/resize:fit:720/0*oSHneX5BOi5VNGYM 720w, https://miro.medium.com/v2/resize:fit:750/0*oSHneX5BOi5VNGYM 750w, https://miro.medium.com/v2/resize:fit:786/0*oSHneX5BOi5VNGYM 786w, https://miro.medium.com/v2/resize:fit:828/0*oSHneX5BOi5VNGYM 828w, https://miro.medium.com/v2/resize:fit:1100/0*oSHneX5BOi5VNGYM 1100w, https://miro.medium.com/v2/resize:fit:1400/0*oSHneX5BOi5VNGYM 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="129" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="19c8" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Rollup Store:</strong></p><p id="48a1" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">We save the results of this aggregation in a persistent store. The next aggregation will simply continue from this checkpoint.</p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi qg"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*93S_a1YJ6zacuBnn%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*93S_a1YJ6zacuBnn%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*93S_a1YJ6zacuBnn%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*93S_a1YJ6zacuBnn%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*93S_a1YJ6zacuBnn%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*93S_a1YJ6zacuBnn%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*93S_a1YJ6zacuBnn%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*93S_a1YJ6zacuBnn 640w, https://miro.medium.com/v2/resize:fit:720/0*93S_a1YJ6zacuBnn 720w, https://miro.medium.com/v2/resize:fit:750/0*93S_a1YJ6zacuBnn 750w, https://miro.medium.com/v2/resize:fit:786/0*93S_a1YJ6zacuBnn 786w, https://miro.medium.com/v2/resize:fit:828/0*93S_a1YJ6zacuBnn 828w, https://miro.medium.com/v2/resize:fit:1100/0*93S_a1YJ6zacuBnn 1100w, https://miro.medium.com/v2/resize:fit:1400/0*93S_a1YJ6zacuBnn 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="318" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="586a" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">We create one such Rollup table <em class="oy">per dataset</em> and use Cassandra as our persistent store. However, as you will soon see in the Control Plane section, the Counter service can be configured to work with any persistent store.</p><p id="18db" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">LastWriteTs</strong>: Every time a given counter receives a write, we also log a <strong class="my gv">last-write-timestamp</strong> as a columnar update in this table. This is done using Cassandra’s <a class="af nu" href="https://docs.datastax.com/en/cql-oss/3.x/cql/cql_reference/cqlInsert.html#cqlInsert__timestamp-value" rel="noopener ugc nofollow" target="_blank">USING TIMESTAMP</a> feature to predictably apply the Last-Write-Win (LWW) semantics. This timestamp is the same as the <em class="oy">event_time</em> for the event. In the subsequent sections, we’ll see how this timestamp is used to keep some counters in active rollup circulation until they have caught up to their latest value.</p><p id="336a" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Rollup Cache</strong></p><p id="25be" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">To optimize read performance, these values are cached in EVCache for each counter. We combine the <strong class="my gv">lastRollupCount</strong> and <strong class="my gv">lastRollupTs</strong> <em class="oy">into a single cached value per counter</em> to prevent potential mismatches between the count and its corresponding checkpoint timestamp.</p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi qk"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*giCU1AtWUYMXHZcI%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*giCU1AtWUYMXHZcI%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*giCU1AtWUYMXHZcI%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*giCU1AtWUYMXHZcI%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*giCU1AtWUYMXHZcI%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*giCU1AtWUYMXHZcI%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*giCU1AtWUYMXHZcI%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*giCU1AtWUYMXHZcI 640w, https://miro.medium.com/v2/resize:fit:720/0*giCU1AtWUYMXHZcI 720w, https://miro.medium.com/v2/resize:fit:750/0*giCU1AtWUYMXHZcI 750w, https://miro.medium.com/v2/resize:fit:786/0*giCU1AtWUYMXHZcI 786w, https://miro.medium.com/v2/resize:fit:828/0*giCU1AtWUYMXHZcI 828w, https://miro.medium.com/v2/resize:fit:1100/0*giCU1AtWUYMXHZcI 1100w, https://miro.medium.com/v2/resize:fit:1400/0*giCU1AtWUYMXHZcI 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="496" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="1bbf" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">But, how do we know which counters to trigger rollups for? Let’s explore our Write and Read path to understand this better.</p><p id="77ab" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Add/Clear Count:</strong></p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi qg"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*wsxgnWH1yR0gHAEL%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*wsxgnWH1yR0gHAEL%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*wsxgnWH1yR0gHAEL%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*wsxgnWH1yR0gHAEL%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*wsxgnWH1yR0gHAEL%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*wsxgnWH1yR0gHAEL%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*wsxgnWH1yR0gHAEL%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*wsxgnWH1yR0gHAEL 640w, https://miro.medium.com/v2/resize:fit:720/0*wsxgnWH1yR0gHAEL 720w, https://miro.medium.com/v2/resize:fit:750/0*wsxgnWH1yR0gHAEL 750w, https://miro.medium.com/v2/resize:fit:786/0*wsxgnWH1yR0gHAEL 786w, https://miro.medium.com/v2/resize:fit:828/0*wsxgnWH1yR0gHAEL 828w, https://miro.medium.com/v2/resize:fit:1100/0*wsxgnWH1yR0gHAEL 1100w, https://miro.medium.com/v2/resize:fit:1400/0*wsxgnWH1yR0gHAEL 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="359" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="6b09" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">An <em class="oy">add</em> or <em class="oy">clear</em> count request writes durably to the TimeSeries Abstraction and updates the last-write-timestamp in the Rollup store. If the durability acknowledgement fails, clients can retry their requests with the same idempotency token without the risk of overcounting.Upon durability, we send a <em class="oy">fire-and-forget </em>request to trigger the rollup for the request counter.</p><p id="6a87" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">GetCount:</strong></p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi qg"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*76pQR6OISx9yuRmi%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*76pQR6OISx9yuRmi%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*76pQR6OISx9yuRmi%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*76pQR6OISx9yuRmi%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*76pQR6OISx9yuRmi%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*76pQR6OISx9yuRmi%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*76pQR6OISx9yuRmi%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*76pQR6OISx9yuRmi 640w, https://miro.medium.com/v2/resize:fit:720/0*76pQR6OISx9yuRmi 720w, https://miro.medium.com/v2/resize:fit:750/0*76pQR6OISx9yuRmi 750w, https://miro.medium.com/v2/resize:fit:786/0*76pQR6OISx9yuRmi 786w, https://miro.medium.com/v2/resize:fit:828/0*76pQR6OISx9yuRmi 828w, https://miro.medium.com/v2/resize:fit:1100/0*76pQR6OISx9yuRmi 1100w, https://miro.medium.com/v2/resize:fit:1400/0*76pQR6OISx9yuRmi 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="359" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="23ce" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">We return the last rolled-up count as<em class="oy"> a quick point-read operation</em>, accepting the trade-off of potentially delivering a slightly stale count. We also trigger a rollup during the read operation to advance the last-rollup-timestamp, enhancing the performance of <em class="oy">subsequent</em> aggregations. This process also <em class="oy">self-remediates </em>a stale count if any previous rollups had failed.</p><p id="2bdc" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">With this approach, the counts<em class="oy"> continually converge</em> to their latest value. Now, let’s see how we scale this approach to millions of counters and thousands of concurrent operations using our Rollup Pipeline.</p><p id="9ab5" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Rollup Pipeline:</strong></p><p id="c974" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Each <strong class="my gv">Counter-Rollup</strong> server operates a rollup pipeline to efficiently aggregate counts across millions of counters. This is where most of the complexity in Counter Abstraction comes in. In the following sections, we will share key details on how efficient aggregations are achieved.</p><p id="d23a" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Light-Weight Roll-Up Event: </strong>As seen in our Write and Read paths above, every operation on a counter sends a light-weight event to the Rollup server:</p><pre class="pk pl pm pn po pv pw px bp py bb bk">rollupEvent: {<br />  "namespace": "my_dataset",<br />  "counter": "counter123"<br />}</pre><p id="8d93" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Note that this event does not include the increment. This is only an indication to the Rollup server that this counter has been accessed and now needs to be aggregated. Knowing exactly which specific counters need to be aggregated prevents scanning the entire event dataset for the purpose of aggregations.</p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi qg"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*Yusg6kC9Jj9ayjbi%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*Yusg6kC9Jj9ayjbi%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*Yusg6kC9Jj9ayjbi%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*Yusg6kC9Jj9ayjbi%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*Yusg6kC9Jj9ayjbi%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*Yusg6kC9Jj9ayjbi%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*Yusg6kC9Jj9ayjbi%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*Yusg6kC9Jj9ayjbi 640w, https://miro.medium.com/v2/resize:fit:720/0*Yusg6kC9Jj9ayjbi 720w, https://miro.medium.com/v2/resize:fit:750/0*Yusg6kC9Jj9ayjbi 750w, https://miro.medium.com/v2/resize:fit:786/0*Yusg6kC9Jj9ayjbi 786w, https://miro.medium.com/v2/resize:fit:828/0*Yusg6kC9Jj9ayjbi 828w, https://miro.medium.com/v2/resize:fit:1100/0*Yusg6kC9Jj9ayjbi 1100w, https://miro.medium.com/v2/resize:fit:1400/0*Yusg6kC9Jj9ayjbi 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="284" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="0e5d" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">In-Memory Rollup Queues:</strong> A given Rollup server instance runs a set of <em class="oy">in-memory</em> queues to receive rollup events and parallelize aggregations. In the first version of this service, we settled on using in-memory queues to reduce provisioning complexity, save on infrastructure costs, and make rebalancing the number of queues fairly straightforward. However, this comes with the trade-off of potentially missing rollup events in case of an instance crash. For more details, see the “Stale Counts” section in “Future Work.”</p><p id="0d6d" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Minimize Duplicate Effort</strong>: We use a fast non-cryptographic hash like <a class="af nu" href="https://xxhash.com/" rel="noopener ugc nofollow" target="_blank">XXHash</a> to ensure that the same set of counters end up on the same queue. Further, we try to minimize the amount of duplicate aggregation work by having a separate rollup stack that chooses to run <em class="oy">fewer</em> <em class="oy">beefier</em> instances.</p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi ql"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*u3p0kGfuwvK5mP_j%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*u3p0kGfuwvK5mP_j%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*u3p0kGfuwvK5mP_j%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*u3p0kGfuwvK5mP_j%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*u3p0kGfuwvK5mP_j%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*u3p0kGfuwvK5mP_j%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*u3p0kGfuwvK5mP_j%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*u3p0kGfuwvK5mP_j 640w, https://miro.medium.com/v2/resize:fit:720/0*u3p0kGfuwvK5mP_j 720w, https://miro.medium.com/v2/resize:fit:750/0*u3p0kGfuwvK5mP_j 750w, https://miro.medium.com/v2/resize:fit:786/0*u3p0kGfuwvK5mP_j 786w, https://miro.medium.com/v2/resize:fit:828/0*u3p0kGfuwvK5mP_j 828w, https://miro.medium.com/v2/resize:fit:1100/0*u3p0kGfuwvK5mP_j 1100w, https://miro.medium.com/v2/resize:fit:1400/0*u3p0kGfuwvK5mP_j 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="440" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="223c" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Availability and Race Conditions: </strong>Having a single Rollup server instance can minimize duplicate aggregation work but may create availability challenges for triggering rollups. <em class="oy">If</em> we choose to horizontally scale the Rollup servers, we allow threads to overwrite rollup values while avoiding any form of distributed locking mechanisms to maintain high availability and performance. This approach remains safe because aggregation occurs within an immutable window. Although the concept of <em class="oy">now()</em> may differ between threads, causing rollup values to sometimes fluctuate, the counts will eventually converge to an accurate value within each immutable aggregation window.</p><p id="acf2" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Rebalancing Queues</strong>: If we need to scale the number of queues, a simple Control Plane configuration update followed by a re-deploy is enough to rebalance the number of queues.</p><pre class="pk pl pm pn po pv pw px bp py bb bk">      "eventual_counter_config": {             <br />          "queue_config": {                    <br />            "num_queues" : 8,  // change to 16 and re-deploy<br />...</pre><p id="3a8b" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Handling Deployments</strong>: During deployments, these queues shut down gracefully, draining all existing events first, while the new Rollup server instance starts up with potentially new queue configurations. There may be a brief period when both the old and new Rollup servers are active, but as mentioned before, this race condition is managed since aggregations occur within immutable windows.</p><p id="a67a" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Minimize Rollup Effort</strong>: Receiving multiple events for the same counter doesn’t mean rolling it up multiple times. We drain these rollup events into a Set, ensuring <em class="oy">a given counter is rolled up only once</em> <em class="oy">during a rollup window</em>.</p><p id="c500" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Efficient Aggregation: </strong>Each rollup consumer processes a batch of counters simultaneously. Within each batch, it queries the underlying TimeSeries abstraction in parallel to aggregate events within specified time boundaries. The TimeSeries abstraction optimizes these range scans to achieve low millisecond latencies.</p><p id="fea5" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Dynamic Batching</strong>: The Rollup server dynamically adjusts the number of time partitions that need to be scanned based on cardinality of counters in order to prevent overwhelming the underlying store with many parallel read requests.</p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi qm"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*hoPpSmQeScn87q0U%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*hoPpSmQeScn87q0U%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*hoPpSmQeScn87q0U%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*hoPpSmQeScn87q0U%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*hoPpSmQeScn87q0U%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*hoPpSmQeScn87q0U%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*hoPpSmQeScn87q0U%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*hoPpSmQeScn87q0U 640w, https://miro.medium.com/v2/resize:fit:720/0*hoPpSmQeScn87q0U 720w, https://miro.medium.com/v2/resize:fit:750/0*hoPpSmQeScn87q0U 750w, https://miro.medium.com/v2/resize:fit:786/0*hoPpSmQeScn87q0U 786w, https://miro.medium.com/v2/resize:fit:828/0*hoPpSmQeScn87q0U 828w, https://miro.medium.com/v2/resize:fit:1100/0*hoPpSmQeScn87q0U 1100w, https://miro.medium.com/v2/resize:fit:1400/0*hoPpSmQeScn87q0U 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="557" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="9446" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Adaptive Back-Pressure</strong>: Each consumer waits for one batch to complete before issuing the rollups for the next batch. It adjusts the wait time between batches based on the performance of the previous batch. This approach provides back-pressure during rollups to prevent overwhelming the underlying TimeSeries store.</p><p id="2693" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Handling Convergence</strong>:</p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi qn"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*-hlw324cMUaC6pQJ%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*-hlw324cMUaC6pQJ%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*-hlw324cMUaC6pQJ%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*-hlw324cMUaC6pQJ%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*-hlw324cMUaC6pQJ%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*-hlw324cMUaC6pQJ%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*-hlw324cMUaC6pQJ%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*-hlw324cMUaC6pQJ 640w, https://miro.medium.com/v2/resize:fit:720/0*-hlw324cMUaC6pQJ 720w, https://miro.medium.com/v2/resize:fit:750/0*-hlw324cMUaC6pQJ 750w, https://miro.medium.com/v2/resize:fit:786/0*-hlw324cMUaC6pQJ 786w, https://miro.medium.com/v2/resize:fit:828/0*-hlw324cMUaC6pQJ 828w, https://miro.medium.com/v2/resize:fit:1100/0*-hlw324cMUaC6pQJ 1100w, https://miro.medium.com/v2/resize:fit:1400/0*-hlw324cMUaC6pQJ 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="447" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="da2c" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">In order to prevent <strong class="my gv">low-cardinality</strong> counters from lagging behind too much and subsequently scanning too many time partitions, they are kept in constant rollup circulation. For <strong class="my gv">high-cardinality</strong> counters, continuously circulating them would consume excessive memory in our Rollup queues. This is where the <strong class="my gv">last-write-timestamp</strong> mentioned previously plays a crucial role. The Rollup server inspects this timestamp to determine if a given counter needs to be re-queued, ensuring that we continue aggregating until it has fully caught up with the writes.</p><p id="af9d" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Now, let’s see how we leverage this counter type to provide an up-to-date current count in near-realtime.</p><h1 id="7747" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Experimental: Accurate Global Counter</h1><p id="8e14" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">We are experimenting with a slightly modified version of the Eventually Consistent counter. Again, take the term ‘Accurate’ with a grain of salt. The key difference between this type of counter and its counterpart is that the <em class="oy">delta</em>, representing the counts since the last-rolled-up timestamp, is computed in real-time.</p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi qg"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*FVOlMO0VgrQoVBBi%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*FVOlMO0VgrQoVBBi%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*FVOlMO0VgrQoVBBi%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*FVOlMO0VgrQoVBBi%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*FVOlMO0VgrQoVBBi%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*FVOlMO0VgrQoVBBi%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*FVOlMO0VgrQoVBBi%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*FVOlMO0VgrQoVBBi 640w, https://miro.medium.com/v2/resize:fit:720/0*FVOlMO0VgrQoVBBi 720w, https://miro.medium.com/v2/resize:fit:750/0*FVOlMO0VgrQoVBBi 750w, https://miro.medium.com/v2/resize:fit:786/0*FVOlMO0VgrQoVBBi 786w, https://miro.medium.com/v2/resize:fit:828/0*FVOlMO0VgrQoVBBi 828w, https://miro.medium.com/v2/resize:fit:1100/0*FVOlMO0VgrQoVBBi 1100w, https://miro.medium.com/v2/resize:fit:1400/0*FVOlMO0VgrQoVBBi 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="290" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="a925" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Aggregating this delta in real-time can impact the performance of this operation, depending on the number of events and partitions that need to be scanned to retrieve this delta. The same principle of rolling up in batches applies here to prevent scanning too many partitions in parallel.</p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi qo"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*M3dbSof98dTfeuNe%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*M3dbSof98dTfeuNe%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*M3dbSof98dTfeuNe%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*M3dbSof98dTfeuNe%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*M3dbSof98dTfeuNe%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*M3dbSof98dTfeuNe%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*M3dbSof98dTfeuNe%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*M3dbSof98dTfeuNe 640w, https://miro.medium.com/v2/resize:fit:720/0*M3dbSof98dTfeuNe 720w, https://miro.medium.com/v2/resize:fit:750/0*M3dbSof98dTfeuNe 750w, https://miro.medium.com/v2/resize:fit:786/0*M3dbSof98dTfeuNe 786w, https://miro.medium.com/v2/resize:fit:828/0*M3dbSof98dTfeuNe 828w, https://miro.medium.com/v2/resize:fit:1100/0*M3dbSof98dTfeuNe 1100w, https://miro.medium.com/v2/resize:fit:1400/0*M3dbSof98dTfeuNe 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="239" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="3b68" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Conversely, if the counters in this dataset areaccessedfrequently, the time gap for the delta remains narrow, making this approach of fetching current counts quite effective.</p><p id="49a2" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Now, let’s see how all this complexity is managed by having a unified Control Plane configuration.</p><h1 id="98ba" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Control Plane</h1><p id="47f1" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">The <a class="af nu" href="https://netflixtechblog.medium.com/data-gateway-a-platform-for-growing-and-protecting-the-data-tier-f1ed8db8f5c6" rel="noopener">Data Gateway Platform Control Plane</a> manages control settings for all abstractions and namespaces, including the Counter Abstraction. Below, is an example of a control plane configuration for a namespace that supports eventually consistent counters with low cardinality:</p><pre class="pk pl pm pn po pv pw px bp py bb bk">"persistence_configuration": [<br />  {<br />    "id": "CACHE",                             // Counter cache config<br />    "scope": "dal=counter",                                                   <br />    "physical_storage": {<br />      "type": "EVCACHE",                       // type of cache storage<br />      "cluster": "evcache_dgw_counter_tier1"   // Shared EVCache cluster<br />    }<br />  },<br />  {<br />    "id": "COUNTER_ROLLUP",<br />    "scope": "dal=counter",                    // Counter abstraction config<br />    "physical_storage": {                     <br />      "type": "CASSANDRA",                     // type of Rollup store<br />      "cluster": "cass_dgw_counter_uc1",       // physical cluster name<br />      "dataset": "my_dataset_1"                // namespace/dataset   <br />    },<br />    "counter_cardinality": "LOW",              // supported counter cardinality<br />    "config": {<br />      "counter_type": "EVENTUAL",              // Type of counter<br />      "eventual_counter_config": {             // eventual counter type<br />        "internal_config": {                  <br />          "queue_config": {                    // adjust w.r.t cardinality<br />            "num_queues" : 8,                  // Rollup queues per instance<br />            "coalesce_ms": 10000,              // coalesce duration for rollups<br />            "capacity_bytes": 16777216         // allocated memory per queue<br />          },<br />          "rollup_batch_count": 32             // parallelization factor<br />        }<br />      }<br />    }<br />  },<br />  {<br />    "id": "EVENT_STORAGE",<br />    "scope": "dal=ts",                         // TimeSeries Event store<br />    "physical_storage": {<br />      "type": "CASSANDRA",                     // persistent store type<br />      "cluster": "cass_dgw_counter_uc1",       // physical cluster name<br />      "dataset": "my_dataset_1",               // keyspace name<br />    },<br />    "config": {                              <br />      "time_partition": {                      // time-partitioning for events<br />        "buckets_per_id": 4,                   // event buckets within<br />        "seconds_per_bucket": "600",           // smaller width for LOW card<br />        "seconds_per_slice": "86400",          // width of a time slice table<br />      },<br />      "accept_limit": "5s",                    // boundary for immutability<br />    },<br />    "lifecycleConfigs": {<br />      "lifecycleConfig": [<br />        {<br />          "type": "retention",                 // Event retention<br />          "config": {<br />            "close_after": "518400s",<br />            "delete_after": "604800s"          // 7 day count event retention<br />          }<br />        }<br />      ]<br />    }<br />  }<br />]</pre><p id="9fd9" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Using such a control plane configuration, we compose multiple abstraction layers using containers deployed on the same host, with each container fetching configuration specific to its scope.</p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi qp"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*4MdrlEjWg2MXU9S3%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*4MdrlEjWg2MXU9S3%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*4MdrlEjWg2MXU9S3%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*4MdrlEjWg2MXU9S3%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*4MdrlEjWg2MXU9S3%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*4MdrlEjWg2MXU9S3%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*4MdrlEjWg2MXU9S3%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*4MdrlEjWg2MXU9S3 640w, https://miro.medium.com/v2/resize:fit:720/0*4MdrlEjWg2MXU9S3 720w, https://miro.medium.com/v2/resize:fit:750/0*4MdrlEjWg2MXU9S3 750w, https://miro.medium.com/v2/resize:fit:786/0*4MdrlEjWg2MXU9S3 786w, https://miro.medium.com/v2/resize:fit:828/0*4MdrlEjWg2MXU9S3 828w, https://miro.medium.com/v2/resize:fit:1100/0*4MdrlEjWg2MXU9S3 1100w, https://miro.medium.com/v2/resize:fit:1400/0*4MdrlEjWg2MXU9S3 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="737" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><h1 id="6176" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Provisioning</h1><p id="ec90" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">As with the TimeSeries abstraction, our automation uses a bunch of user inputs regarding their workload and cardinalities to arrive at the right set of infrastructure and related control plane configuration. You can learn more about this process in a talk given by one of our stunning colleagues, <a class="af nu" href="https://www.linkedin.com/in/joseph-lynch-9976a431/" rel="noopener ugc nofollow" target="_blank">Joey Lynch</a> : <a class="af nu" href="https://www.youtube.com/watch?v=Lf6B1PxIvAs" rel="noopener ugc nofollow" target="_blank">How Netflix optimally provisions infrastructure in the cloud</a>.</p><h1 id="7e07" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Performance</h1><p id="f469" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">At the time of writing this blog, this service was processing close to <strong class="my gv">75K count requests/second</strong><em class="oy"> globally</em> across the different API endpoints and datasets:</p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi qg"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*1h_af4Kk3YrZrqlc%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*1h_af4Kk3YrZrqlc%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*1h_af4Kk3YrZrqlc%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*1h_af4Kk3YrZrqlc%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*1h_af4Kk3YrZrqlc%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*1h_af4Kk3YrZrqlc%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*1h_af4Kk3YrZrqlc%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*1h_af4Kk3YrZrqlc 640w, https://miro.medium.com/v2/resize:fit:720/0*1h_af4Kk3YrZrqlc 720w, https://miro.medium.com/v2/resize:fit:750/0*1h_af4Kk3YrZrqlc 750w, https://miro.medium.com/v2/resize:fit:786/0*1h_af4Kk3YrZrqlc 786w, https://miro.medium.com/v2/resize:fit:828/0*1h_af4Kk3YrZrqlc 828w, https://miro.medium.com/v2/resize:fit:1100/0*1h_af4Kk3YrZrqlc 1100w, https://miro.medium.com/v2/resize:fit:1400/0*1h_af4Kk3YrZrqlc 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="357" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="92ef" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">while providing<strong class="my gv"> single-digit millisecond</strong> latencies for all its endpoints:</p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi qg"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*UnI7eore6gvuqrrF%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*UnI7eore6gvuqrrF%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*UnI7eore6gvuqrrF%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*UnI7eore6gvuqrrF%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*UnI7eore6gvuqrrF%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*UnI7eore6gvuqrrF%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*UnI7eore6gvuqrrF%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*UnI7eore6gvuqrrF 640w, https://miro.medium.com/v2/resize:fit:720/0*UnI7eore6gvuqrrF 720w, https://miro.medium.com/v2/resize:fit:750/0*UnI7eore6gvuqrrF 750w, https://miro.medium.com/v2/resize:fit:786/0*UnI7eore6gvuqrrF 786w, https://miro.medium.com/v2/resize:fit:828/0*UnI7eore6gvuqrrF 828w, https://miro.medium.com/v2/resize:fit:1100/0*UnI7eore6gvuqrrF 1100w, https://miro.medium.com/v2/resize:fit:1400/0*UnI7eore6gvuqrrF 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="366" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><h1 id="3772" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Future Work</h1><p id="19ec" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">While our system is robust, we still have work to do in making it more reliable and enhancing its features. Some of that work includes:</p><ul class=""><li id="bafb" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt qf pa pb bk"><strong class="my gv">Regional Rollups: </strong>Cross-region replication issues can result in missed events from other regions. An alternate strategy involves establishing a rollup table for each region, and then tallying them in a global rollup table. A key challenge in this design would be effectively communicating the clearing of the counter across regions.</li><li id="7818" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt qf pa pb bk"><strong class="my gv">Error Detection and Stale Counts</strong>: Excessively stale counts can occur if rollup events are lost or if a rollups fails and isn’t retried. This isn’t an issue for frequently accessed counters, as they remain in rollup circulation. This issue is more pronounced for counters that aren’t accessed frequently. Typically, the initial read for such a counter will trigger a rollup,<em class="oy"> self-remediating </em>the issue. However, for use cases that cannot accept potentially stale initial reads, we plan to implement improved error detection and utilize durable queues for resilient retries.</li></ul><h1 id="18c4" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Conclusion</h1><p id="64c0" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">Distributed counting remains a challenging problem in computer science. In this blog, we explored multiple approaches to implement and deploy a Counting service at scale. While there may be other methods for distributed counting, our goal has been to deliver blazing fast performance at low infrastructure costs while maintaining high availability and providing idempotency guarantees. Along the way, we make various trade-offs to meet the diverse counting requirements at Netflix. We hope you found this blog post insightful.</p><p id="a883" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Stay tuned for <strong class="my gv">Part 3 </strong>of Composite Abstractions at Netflix, where we’ll introduce our <strong class="my gv">Graph Abstraction</strong>, a new service being built on top of the <a class="af nu" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/introducing-netflixs-key-value-data-abstraction-layer-1ea8a0a11b30">Key-Value Abstraction</a> <em class="oy">and</em> the <a class="af nu" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/introducing-netflix-timeseries-data-abstraction-layer-31552f6326f8">TimeSeries Abstraction</a> to handle high-throughput, low-latency graphs.</p><h1 id="71bd" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Acknowledgments</h1><p id="78b5" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">Special thanks to our stunning colleagues who contributed to the Counter Abstraction’s success: <a class="af nu" href="https://www.linkedin.com/in/joseph-lynch-9976a431/" rel="noopener ugc nofollow" target="_blank">Joey Lynch</a>, <a class="af nu" href="https://www.linkedin.com/in/vinaychella/" rel="noopener ugc nofollow" target="_blank">Vinay Chella</a>, <a class="af nu" href="https://www.linkedin.com/in/kaidanfullerton/" rel="noopener ugc nofollow" target="_blank">Kaidan Fullerton</a>, <a class="af nu" href="https://www.linkedin.com/in/tomdevoe/" rel="noopener ugc nofollow" target="_blank">Tom DeVoe</a>, <a class="af nu" href="https://www.linkedin.com/in/mengqingwang/" rel="noopener ugc nofollow" target="_blank">Mengqing Wang</a></p></div>]]></description>
      <link>https://netflixtechblog.com/netflixs-distributed-counter-abstraction-8d0c45eb66b2</link>
      <guid>https://netflixtechblog.com/netflixs-distributed-counter-abstraction-8d0c45eb66b2</guid>
      <pubDate>Tue, 12 Nov 2024 21:45:00 +0100</pubDate>
    </item>
    <item>
      <title><![CDATA[Investigation of a Workbench UI Latency Issue]]></title>
      <description><![CDATA[<div><div></div><p id="c66a" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">By: <a class="af nu" href="https://www.linkedin.com/in/hechaoli/" rel="noopener ugc nofollow" target="_blank">Hechao Li</a> and <a class="af nu" href="https://www.linkedin.com/in/mayworm/" rel="noopener ugc nofollow" target="_blank">Marcelo Mayworm</a></p><p id="2b76" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">With special thanks to our stunning colleagues <a class="af nu" href="https://www.linkedin.com/in/amer-ather-9071181/" rel="noopener ugc nofollow" target="_blank">Amer Ather</a>, <a class="af nu" href="https://www.linkedin.com/in/itaydafna" rel="noopener ugc nofollow" target="_blank">Itay Dafna</a>, <a class="af nu" href="https://www.linkedin.com/in/lucaepozzi/" rel="noopener ugc nofollow" target="_blank">Luca Pozzi</a>, <a class="af nu" href="https://www.linkedin.com/in/matheusdeoleao/" rel="noopener ugc nofollow" target="_blank">Matheus Leão</a>, and <a class="af nu" href="https://www.linkedin.com/in/yeji682/" rel="noopener ugc nofollow" target="_blank">Ye Ji</a>.</p><h1 id="c1c3" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Overview</h1><p id="072a" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">At Netflix, the Analytics and Developer Experience organization, part of the Data Platform, offers a product called Workbench. Workbench is a remote development workspace based on<a class="af nu" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/titus-the-netflix-container-management-platform-is-now-open-source-f868c9fb5436"> Titus</a> that allows data practitioners to work with big data and machine learning use cases at scale. A common use case for Workbench is running<a class="af nu" href="https://jupyterlab.readthedocs.io/en/latest/" rel="noopener ugc nofollow" target="_blank"> JupyterLab</a> Notebooks.</p><p id="d5f7" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Recently, several users reported that their JupyterLab UI becomes slow and unresponsive when running certain notebooks. This document details the intriguing process of debugging this issue, all the way from the UI down to the Linux kernel.</p><h1 id="1fae" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Symptom</h1><p id="6f03" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">Machine Learning engineer <a class="af nu" href="https://www.linkedin.com/in/lucaepozzi/" rel="noopener ugc nofollow" target="_blank">Luca Pozzi</a> reported to our Data Platform team that their <strong class="my gv">JupyterLab UI on their workbench becomes slow and unresponsive when running some of their Notebooks.</strong> Restarting the <em class="oy">ipykernel</em> process, which runs the Notebook, might temporarily alleviate the problem, but the frustration persists as more notebooks are run.</p><h1 id="c9b9" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Quantify the Slowness</h1><p id="35ea" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">While we observed the issue firsthand, the term “UI being slow” is subjective and difficult to measure. To investigate this issue, <strong class="my gv">we needed a quantitative analysis of the slowness</strong>.</p><p id="2efa" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><a class="af nu" href="https://www.linkedin.com/in/itaydafna" rel="noopener ugc nofollow" target="_blank">Itay Dafna</a> devised an effective and simple method to quantify the UI slowness. Specifically, we opened a terminal via JupyterLab and held down a key (e.g., “j”) for 15 seconds while running the user’s notebook. The input to stdin is sent to the backend (i.e., JupyterLab) via a WebSocket, and the output to stdout is sent back from the backend and displayed on the UI. We then exported the <em class="oy">.har </em>file recording all communications from the browser and loaded it into a Notebook for analysis.</p><figure class="pc pd pe pf pg ph oz pa paragraph-image"><div role="button" tabindex="0" class="pi pj fj pk bh pl"><div class="oz pa pb"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*ltV3CYtNjLCzolXD%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*ltV3CYtNjLCzolXD%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*ltV3CYtNjLCzolXD%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*ltV3CYtNjLCzolXD%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*ltV3CYtNjLCzolXD%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*ltV3CYtNjLCzolXD%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*ltV3CYtNjLCzolXD%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*ltV3CYtNjLCzolXD 640w, https://miro.medium.com/v2/resize:fit:720/0*ltV3CYtNjLCzolXD 720w, https://miro.medium.com/v2/resize:fit:750/0*ltV3CYtNjLCzolXD 750w, https://miro.medium.com/v2/resize:fit:786/0*ltV3CYtNjLCzolXD 786w, https://miro.medium.com/v2/resize:fit:828/0*ltV3CYtNjLCzolXD 828w, https://miro.medium.com/v2/resize:fit:1100/0*ltV3CYtNjLCzolXD 1100w, https://miro.medium.com/v2/resize:fit:1400/0*ltV3CYtNjLCzolXD 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pm c" width="700" height="252" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="e91b" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Using this approach, we observed latencies ranging from 1 to 10 seconds, averaging 7.4 seconds.</p><figure class="pc pd pe pf pg ph oz pa paragraph-image"><div role="button" tabindex="0" class="pi pj fj pk bh pl"><div class="oz pa pn"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*H7KW62J0jZKPTjQH%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*H7KW62J0jZKPTjQH%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*H7KW62J0jZKPTjQH%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*H7KW62J0jZKPTjQH%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*H7KW62J0jZKPTjQH%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*H7KW62J0jZKPTjQH%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*H7KW62J0jZKPTjQH%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*H7KW62J0jZKPTjQH 640w, https://miro.medium.com/v2/resize:fit:720/0*H7KW62J0jZKPTjQH 720w, https://miro.medium.com/v2/resize:fit:750/0*H7KW62J0jZKPTjQH 750w, https://miro.medium.com/v2/resize:fit:786/0*H7KW62J0jZKPTjQH 786w, https://miro.medium.com/v2/resize:fit:828/0*H7KW62J0jZKPTjQH 828w, https://miro.medium.com/v2/resize:fit:1100/0*H7KW62J0jZKPTjQH 1100w, https://miro.medium.com/v2/resize:fit:1400/0*H7KW62J0jZKPTjQH 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pm c" width="700" height="176" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><h1 id="6cd5" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Blame The Notebook</h1><p id="ef5b" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">Now that we have an objective metric for the slowness, let’s officially start our investigation. If you have read the symptom carefully, you must have noticed that the slowness only occurs when the user runs <strong class="my gv">certain</strong> notebooks but not others.</p><p id="c042" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Therefore, the first step is scrutinizing the specific Notebook experiencing the issue. Why does the UI always slow down after running this particular Notebook? Naturally, you would think that there must be something wrong with the code running in it.</p><p id="cf81" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Upon closely examining the user’s Notebook, we noticed a library called <em class="oy">pystan</em> , which provides Python bindings to a native C++ library called stan, looked suspicious. Specifically, <em class="oy">pystan</em> uses <em class="oy">asyncio</em>. However, <strong class="my gv">because there is already an existing <em class="oy">asyncio</em> event loop running in the Notebook process and <em class="oy">asyncio</em> cannot be nested by design, in order for <em class="oy">pystan</em> to work, the authors of <em class="oy">pystan</em> </strong><a class="af nu" href="https://pystan.readthedocs.io/en/latest/faq.html#how-can-i-use-pystan-with-jupyter-notebook-or-jupyterlab" rel="noopener ugc nofollow" target="_blank"><strong class="my gv">recommend</strong></a><strong class="my gv"> injecting <em class="oy">pystan</em> into the existing event loop by using a package called </strong><a class="af nu" href="https://pypi.org/project/nest-asyncio/" rel="noopener ugc nofollow" target="_blank"><strong class="my gv"><em class="oy">nest_asyncio</em></strong></a>, a library that became unmaintained because <a class="af nu" href="https://github.com/erdewit/ib_insync/commit/ef5ea29e44e0c40bbadbc16c2281b3ac58aa4a40" rel="noopener ugc nofollow" target="_blank">the author unfortunately passed away</a>.</p><p id="de21" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Given this seemingly hacky usage, we naturally suspected that the events injected by <em class="oy">pystan</em> into the event loop were blocking the handling of the WebSocket messages used to communicate with the JupyterLab UI. This reasoning sounds very plausible. However, <strong class="my gv">the user claimed that there were cases when a Notebook not using <em class="oy">pystan</em> runs, the UI also became slow</strong>.</p><p id="ca77" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Moreover, after several rounds of discussion with ChatGPT, we learned more about the architecture and realized that, in theory, <strong class="my gv">the usage of <em class="oy">pystan</em> and <em class="oy">nest_asyncio</em> should not cause the slowness in handling the UI WebSocket</strong> for the following reasons:</p><p id="17ba" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Even though <em class="oy">pystan</em> uses <em class="oy">nest_asyncio</em> to inject itself into the main event loop, <strong class="my gv">the Notebook runs on a child process (i.e.</strong>,<strong class="my gv"> the <em class="oy">ipykernel</em> process) of the <em class="oy">jupyter-lab</em> server process</strong>, which means the main event loop being injected by <em class="oy">pystan</em> is that of the <em class="oy">ipykernel</em> process, not the <em class="oy">jupyter-server</em> process. Therefore, even if <em class="oy">pystan</em> blocks the event loop, it shouldn’t impact the <em class="oy">jupyter-lab</em> main event loop that is used for UI websocket communication. See the diagram below:</p><figure class="pc pd pe pf pg ph oz pa paragraph-image"><div role="button" tabindex="0" class="pi pj fj pk bh pl"><div class="oz pa po"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*DsQuZV5qnRXp5mVw%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*DsQuZV5qnRXp5mVw%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*DsQuZV5qnRXp5mVw%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*DsQuZV5qnRXp5mVw%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*DsQuZV5qnRXp5mVw%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*DsQuZV5qnRXp5mVw%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*DsQuZV5qnRXp5mVw%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*DsQuZV5qnRXp5mVw 640w, https://miro.medium.com/v2/resize:fit:720/0*DsQuZV5qnRXp5mVw 720w, https://miro.medium.com/v2/resize:fit:750/0*DsQuZV5qnRXp5mVw 750w, https://miro.medium.com/v2/resize:fit:786/0*DsQuZV5qnRXp5mVw 786w, https://miro.medium.com/v2/resize:fit:828/0*DsQuZV5qnRXp5mVw 828w, https://miro.medium.com/v2/resize:fit:1100/0*DsQuZV5qnRXp5mVw 1100w, https://miro.medium.com/v2/resize:fit:1400/0*DsQuZV5qnRXp5mVw 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pm c" width="700" height="591" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="c601" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">In other words, <strong class="my gv"><em class="oy">pystan</em> events are injected to the event loop B in this diagram instead of event loop A</strong>. So, it shouldn’t block the UI WebSocket events.</p><p id="1d8c" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">You might also think that because event loop A handles both the WebSocket events from the UI and the ZeroMQ socket events from the <em class="oy">ipykernel</em> process, a high volume of ZeroMQ events generated by the notebook could block the WebSocket. However, <strong class="my gv">when we captured packets on the ZeroMQ socket while reproducing the issue, we didn’t observe heavy traffic on this socket that could cause such blocking</strong>.</p><p id="f5d9" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">A stronger piece of evidence to rule out <em class="oy">pystan</em> was that we were ultimately able to reproduce the issue even without it, which I’ll dive into later.</p><h1 id="ccc2" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Blame Noisy Neighbors</h1><p id="87ad" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">The Workbench instance runs as a <a class="af nu" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/titus-the-netflix-container-management-platform-is-now-open-source-f868c9fb5436">Titus container</a>. To efficiently utilize our compute resources, <strong class="my gv">Titus employs a CPU oversubscription feature</strong>, meaning the combined virtual CPUs allocated to containers exceed the number of available physical CPUs on a Titus agent. <strong class="my gv">If a container is unfortunate enough to be scheduled alongside other “noisy” containers — those that consume a lot of CPU resources — it could suffer from CPU deficiency.</strong></p><p id="d99f" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">However, after examining the CPU utilization of neighboring containers on the same Titus agent as the Workbench instance, as well as the overall CPU utilization of the Titus agent, we quickly ruled out this hypothesis. Using the top command on the Workbench, we observed that when running the Notebook, <strong class="my gv">the Workbench instance uses only 4 out of the 64 CPUs allocated to it</strong>. Simply put, <strong class="my gv">this workload is not CPU-bound.</strong></p><figure class="pc pd pe pf pg ph oz pa paragraph-image"><div role="button" tabindex="0" class="pi pj fj pk bh pl"><div class="oz pa pp"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*YXsntKLiontnkNhf%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*YXsntKLiontnkNhf%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*YXsntKLiontnkNhf%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*YXsntKLiontnkNhf%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*YXsntKLiontnkNhf%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*YXsntKLiontnkNhf%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*YXsntKLiontnkNhf%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*YXsntKLiontnkNhf 640w, https://miro.medium.com/v2/resize:fit:720/0*YXsntKLiontnkNhf 720w, https://miro.medium.com/v2/resize:fit:750/0*YXsntKLiontnkNhf 750w, https://miro.medium.com/v2/resize:fit:786/0*YXsntKLiontnkNhf 786w, https://miro.medium.com/v2/resize:fit:828/0*YXsntKLiontnkNhf 828w, https://miro.medium.com/v2/resize:fit:1100/0*YXsntKLiontnkNhf 1100w, https://miro.medium.com/v2/resize:fit:1400/0*YXsntKLiontnkNhf 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pm c" width="700" height="252" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><h1 id="ac15" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Blame The Network</h1><p id="8e12" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">The next theory was that the network between the web browser UI (on the laptop) and the JupyterLab server was slow. To investigate, we <strong class="my gv">captured all the packets between the laptop and the server</strong> while running the Notebook and continuously pressing ‘j’ in the terminal.</p><p id="0018" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">When the UI experienced delays, we observed a 5-second pause in packet transmission from server port 8888 to the laptop. Meanwhile,<strong class="my gv"> traffic from other ports, such as port 22 for SSH, remained unaffected</strong>. This led us to conclude that the pause was caused by the application running on port 8888 (i.e., the JupyterLab process) rather than the network.</p><figure class="pc pd pe pf pg ph oz pa paragraph-image"><div role="button" tabindex="0" class="pi pj fj pk bh pl"><div class="oz pa pq"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*c660xBwF4XuCA8KN%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*c660xBwF4XuCA8KN%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*c660xBwF4XuCA8KN%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*c660xBwF4XuCA8KN%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*c660xBwF4XuCA8KN%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*c660xBwF4XuCA8KN%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*c660xBwF4XuCA8KN%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*c660xBwF4XuCA8KN 640w, https://miro.medium.com/v2/resize:fit:720/0*c660xBwF4XuCA8KN 720w, https://miro.medium.com/v2/resize:fit:750/0*c660xBwF4XuCA8KN 750w, https://miro.medium.com/v2/resize:fit:786/0*c660xBwF4XuCA8KN 786w, https://miro.medium.com/v2/resize:fit:828/0*c660xBwF4XuCA8KN 828w, https://miro.medium.com/v2/resize:fit:1100/0*c660xBwF4XuCA8KN 1100w, https://miro.medium.com/v2/resize:fit:1400/0*c660xBwF4XuCA8KN 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pm c" width="700" height="115" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><h1 id="ff04" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">The Minimal Reproduction</h1><p id="b5d7" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">As previously mentioned, another strong piece of evidence proving the innocence of pystan was that we could reproduce the issue without it. By gradually stripping down the “bad” Notebook, we eventually arrived at a minimal snippet of code that reproduces the issue without any third-party dependencies or complex logic:</p><pre class="pc pd pe pf pg pr ps pt bp pu bb bk">import time<br />import os<br />from multiprocessing import ProcessN = os.cpu_count()def launch_worker(worker_id):<br />  time.sleep(60)if __name__ == '__main__':<br />  with open('/root/2GB_file', 'r') as file:<br />    data = file.read()<br />    processes = []<br />    for i in range(N):<br />      p = Process(target=launch_worker, args=(i,))<br />      processes.append(p)<br />      p.start()for p in processes:<br />      p.join()</pre><p id="04dc" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">The code does only two things:</p><ol class=""><li id="2d46" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt qa qb qc bk">Read a 2GB file into memory (the Workbench instance has 480G memory in total so this memory usage is almost negligible).</li><li id="9c35" class="mw mx gu my b mz qd nb nc nd qe nf ng nh qf nj nk nl qg nn no np qh nr ns nt qa qb qc bk">Start N processes where N is the number of CPUs. The N processes do nothing but sleep.</li></ol><p id="9ea8" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">There is no doubt that this is the most silly piece of code I’ve ever written. It is neither CPU bound nor memory bound. Yet <strong class="my gv">it can cause the JupyterLab UI to stall for as many as 10 seconds!</strong></p><h1 id="2dca" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Questions</h1><p id="24d9" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">There are a couple of interesting observations that raise several questions:</p><ul class=""><li id="bba4" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt qi qb qc bk">We noticed that <strong class="my gv">both steps are required in order to reproduce the issue</strong>. If you don’t read the 2GB file (that is not even used!), the issue is not reproducible. <strong class="my gv">Why using 2GB out of 480GB memory could impact the performance?</strong></li><li id="3e91" class="mw mx gu my b mz qd nb nc nd qe nf ng nh qf nj nk nl qg nn no np qh nr ns nt qi qb qc bk"><strong class="my gv">When the UI delay occurs, the <em class="oy">jupyter-lab</em> process CPU utilization spikes to 100%</strong>, hinting at contention on the single-threaded event loop in this process (event loop A in the diagram before). <strong class="my gv">What does the <em class="oy">jupyter-lab</em> process need the CPU for, given that it is not the process that runs the Notebook?</strong></li><li id="72c2" class="mw mx gu my b mz qd nb nc nd qe nf ng nh qf nj nk nl qg nn no np qh nr ns nt qi qb qc bk">The code runs in a Notebook, which means it runs in the <em class="oy">ipykernel</em> process, that is a child process of the <em class="oy">jupyter-lab</em> process. <strong class="my gv">How can anything that happens in a child process cause the parent process to have CPU contention?</strong></li><li id="101d" class="mw mx gu my b mz qd nb nc nd qe nf ng nh qf nj nk nl qg nn no np qh nr ns nt qi qb qc bk">The workbench has 64CPUs. But when we printed <em class="oy">os.cpu_count()</em>, the output was 96. That means <strong class="my gv">the code starts more processes than the number of CPUs</strong>. <strong class="my gv">Why is that?</strong></li></ul><p id="9d59" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Let’s answer the last question first. In fact, if you run <em class="oy">lscpu</em> and <em class="oy">nproc</em> commands inside a Titus container, you will also see different results — the former gives you 96, which is the number of physical CPUs on the Titus agent, whereas the latter gives you 64, which is the number of virtual CPUs allocated to the container. This discrepancy is due to the lack of a “CPU namespace” in the Linux kernel, causing the number of physical CPUs to be leaked to the container when calling certain functions to get the CPU count. The assumption here is that Python <strong class="my gv"><em class="oy">os.cpu_count()</em> uses the same function as the <em class="oy">lscpu</em> command, causing it to get the CPU count of the host instead of the container</strong>. Python 3.13 has <a class="af nu" href="https://docs.python.org/3.13/library/os.html#os.process_cpu_count" rel="noopener ugc nofollow" target="_blank">a new call that can be used to get the accurate CPU count</a>, but it’s not GA’ed yet.</p><p id="ea87" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">It will be proven later that this inaccurate number of CPUs can be a contributing factor to the slowness.</p><h1 id="cdd1" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">More Clues</h1><p id="b480" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">Next, we used <em class="oy">py-spy</em> to do a profiling of the <em class="oy">jupyter-lab</em> process. Note that we profiled the parent <em class="oy">jupyter-lab </em>process, <strong class="my gv">not</strong> the <em class="oy">ipykernel</em> child process that runs the reproduction code. The profiling result is as follows:</p><figure class="pc pd pe pf pg ph oz pa paragraph-image"><div role="button" tabindex="0" class="pi pj fj pk bh pl"><div class="oz pa pq"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*ho2C4015Disa8aFv%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*ho2C4015Disa8aFv%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*ho2C4015Disa8aFv%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*ho2C4015Disa8aFv%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*ho2C4015Disa8aFv%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*ho2C4015Disa8aFv%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*ho2C4015Disa8aFv%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*ho2C4015Disa8aFv 640w, https://miro.medium.com/v2/resize:fit:720/0*ho2C4015Disa8aFv 720w, https://miro.medium.com/v2/resize:fit:750/0*ho2C4015Disa8aFv 750w, https://miro.medium.com/v2/resize:fit:786/0*ho2C4015Disa8aFv 786w, https://miro.medium.com/v2/resize:fit:828/0*ho2C4015Disa8aFv 828w, https://miro.medium.com/v2/resize:fit:1100/0*ho2C4015Disa8aFv 1100w, https://miro.medium.com/v2/resize:fit:1400/0*ho2C4015Disa8aFv 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pm c" width="700" height="433" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="55b0" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">As one can see, <strong class="my gv">a lot of CPU time (89%!!) is spent on a function called <em class="oy">__parse_smaps_rollup</em></strong>. In comparison, the terminal handler used only 0.47% CPU time. From the stack trace, we see that <strong class="my gv">this function is inside the event loop A</strong>,<strong class="my gv"> so it can definitely cause the UI WebSocket events to be delayed</strong>.</p><p id="fd28" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">The stack trace also shows that this function is ultimately called by a function used by a Jupyter lab extension called <em class="oy">jupyter_resource_usage</em>. <strong class="my gv">We then disabled this extension and restarted the <em class="oy">jupyter-lab</em> process. As you may have guessed, we could no longer reproduce the slowness!</strong></p><p id="5e8f" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">But our puzzle is not solved yet. Why does this extension cause the UI to slow down? Let’s keep digging.</p><h1 id="ff5d" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Root Cause Analysis</h1><p id="694f" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">From the name of the extension and the names of the other functions it calls, we can infer that this extension is used to get resources such as CPU and memory usage information. Examining the code, we see that this function call stack is triggered when an API endpoint <em class="oy">/metrics/v1</em> is called from the UI. <strong class="my gv">The UI apparently calls this function periodically</strong>, according to the network traffic tab in Chrome’s Developer Tools.</p><p id="5465" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Now let’s look at the implementation starting from the call <em class="oy">get(jupter_resource_usage/api.py:42)</em> . The full code is <a class="af nu" href="https://github.com/jupyter-server/jupyter-resource-usage/blob/6f15ef91d5c7e50853516b90b5e53b3913d2ed34/jupyter_resource_usage/api.py#L28" rel="noopener ugc nofollow" target="_blank">here</a> and the key lines are shown below:</p><pre class="pc pd pe pf pg pr ps pt bp pu bb bk">cur_process = psutil.Process()<br />all_processes = [cur_process] + cur_process.children(recursive=True)for p in all_processes:<br />  info = p.memory_full_info()</pre><p id="1f1a" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Basically, it gets all children processes of the <em class="oy">jupyter-lab</em> process recursively, including both the <em class="oy">ipykernel</em> Notebook process and all processes created by the Notebook. Obviously, <strong class="my gv">the cost of this function is linear to the number of all children processes</strong>. In the reproduction code, we create 96 processes. So here we will have at least 96 (sleep processes) + 1 (<em class="oy">ipykernel</em> process) + 1 (<em class="oy">jupyter-lab</em> process) = 98 processes when it should actually be 64 (allocated CPUs) + 1 (<em class="oy">ipykernel</em> process) + 1 <em class="oy">(jupyter-lab</em> process) = 66 processes, because the number of CPUs allocated to the container is, in fact, 64.</p><figure class="pc pd pe pf pg ph oz pa paragraph-image"><div role="button" tabindex="0" class="pi pj fj pk bh pl"><div class="oz pa qj"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*sHTjycVMUk1yVAsk%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*sHTjycVMUk1yVAsk%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*sHTjycVMUk1yVAsk%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*sHTjycVMUk1yVAsk%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*sHTjycVMUk1yVAsk%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*sHTjycVMUk1yVAsk%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*sHTjycVMUk1yVAsk%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*sHTjycVMUk1yVAsk 640w, https://miro.medium.com/v2/resize:fit:720/0*sHTjycVMUk1yVAsk 720w, https://miro.medium.com/v2/resize:fit:750/0*sHTjycVMUk1yVAsk 750w, https://miro.medium.com/v2/resize:fit:786/0*sHTjycVMUk1yVAsk 786w, https://miro.medium.com/v2/resize:fit:828/0*sHTjycVMUk1yVAsk 828w, https://miro.medium.com/v2/resize:fit:1100/0*sHTjycVMUk1yVAsk 1100w, https://miro.medium.com/v2/resize:fit:1400/0*sHTjycVMUk1yVAsk 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pm c" width="700" height="326" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="6210" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">This is truly ironic. <strong class="my gv">The more CPUs we have, the slower we are!</strong></p><p id="98c2" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">At this point, we have answered one question: <strong class="my gv">Why does starting many grandchildren processes in the child process cause the parent process to be slow? </strong>Because the parent process runs a function that’s linear to the number all children process recursively.</p><p id="1f1f" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">However, this solves only half of the puzzle. If you remember the previous analysis, <strong class="my gv">starting many child processes ALONE doesn’t reproduce the issue</strong>. If we don’t read the 2GB file, even if we create 2x more processes, we can’t reproduce the slowness.</p><p id="147b" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">So now we must answer the next question: <strong class="my gv">Why does reading a 2GB file in the child process affect the parent process performance, </strong>especially when the workbench has as much as 480GB memory in total?</p><p id="49ac" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">To answer this question, let’s look closely at the function <em class="oy">__parse_smaps_rollup</em>. As the name implies, <a class="af nu" href="https://github.com/giampaolo/psutil/blob/c034e6692cf736b5e87d14418a8153bb03f6cf42/psutil/_pslinux.py#L1978" rel="noopener ugc nofollow" target="_blank">this function</a> parses the file <em class="oy">/proc/&lt;pid&gt;/smaps_rollup</em>.</p><pre class="pc pd pe pf pg pr ps pt bp pu bb bk">def _parse_smaps_rollup(self):<br />  uss = pss = swap = 0<br />  with open_binary("{}/{}/smaps_rollup".format(self._procfs_path, self.pid)) as f:<br />  for line in f:<br />    if line.startswith(b”Private_”):<br />    # Private_Clean, Private_Dirty, Private_Hugetlb<br />      s uss += int(line.split()[1]) * 1024<br />    elif line.startswith(b”Pss:”):<br />      pss = int(line.split()[1]) * 1024<br />    elif line.startswith(b”Swap:”):<br />      swap = int(line.split()[1]) * 1024<br />return (uss, pss, swap)</pre><p id="6952" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Naturally, you might think that when memory usage increases, this file becomes larger in size, causing the function to take longer to parse. Unfortunately, this is not the answer because:</p><ul class=""><li id="2f67" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt qi qb qc bk">First, <a class="af nu" href="https://www.kernel.org/doc/Documentation/ABI/testing/procfs-smaps_rollup" rel="noopener ugc nofollow" target="_blank"><strong class="my gv">the number of lines in this file is constant</strong></a><strong class="my gv"> for all processes</strong>.</li><li id="173a" class="mw mx gu my b mz qd nb nc nd qe nf ng nh qf nj nk nl qg nn no np qh nr ns nt qi qb qc bk">Second, <strong class="my gv">this is a special file in the /proc filesystem, which should be seen as a kernel interface</strong> instead of a regular file on disk. In other words, <strong class="my gv">I/O operations of this file are handled by the kernel rather than disk</strong>.</li></ul><p id="5700" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">This file was introduced in <a class="af nu" href="https://github.com/torvalds/linux/commit/493b0e9d945fa9dfe96be93ae41b4ca4b6fdb317#diff-cb79e2d6ea6f9627ff68d1342a219f800e04ff6c6fa7b90c7e66bb391b2dd3ee" rel="noopener ugc nofollow" target="_blank">this commit</a> in 2017, with the purpose of improving the performance of user programs that determine aggregate memory statistics. Let’s first focus on <a class="af nu" href="https://elixir.bootlin.com/linux/v6.5.13/source/fs/proc/task_mmu.c#L1025" rel="noopener ugc nofollow" target="_blank">the handler of <em class="oy">open</em> syscall</a> on this <em class="oy">/proc/&lt;pid&gt;/smaps_rollup</em>.</p><figure class="pc pd pe pf pg ph oz pa paragraph-image"><div role="button" tabindex="0" class="pi pj fj pk bh pl"><div class="oz pa qk"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*vGOD79Tleii7X22B%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*vGOD79Tleii7X22B%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*vGOD79Tleii7X22B%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*vGOD79Tleii7X22B%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*vGOD79Tleii7X22B%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*vGOD79Tleii7X22B%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*vGOD79Tleii7X22B%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*vGOD79Tleii7X22B 640w, https://miro.medium.com/v2/resize:fit:720/0*vGOD79Tleii7X22B 720w, https://miro.medium.com/v2/resize:fit:750/0*vGOD79Tleii7X22B 750w, https://miro.medium.com/v2/resize:fit:786/0*vGOD79Tleii7X22B 786w, https://miro.medium.com/v2/resize:fit:828/0*vGOD79Tleii7X22B 828w, https://miro.medium.com/v2/resize:fit:1100/0*vGOD79Tleii7X22B 1100w, https://miro.medium.com/v2/resize:fit:1400/0*vGOD79Tleii7X22B 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pm c" width="700" height="579" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="52cf" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Following through the <em class="oy">single_open</em> <a class="af nu" href="https://elixir.bootlin.com/linux/v6.5.13/source/fs/seq_file.c#L582" rel="noopener ugc nofollow" target="_blank">function</a>, we will find that it uses the function <em class="oy">show_smaps_rollup</em> for the show operation, which can translate to the <em class="oy">read</em> system call on the file. Next, we look at the <em class="oy">show_smaps_rollup</em> <a class="af nu" href="https://elixir.bootlin.com/linux/v6.5.13/source/fs/proc/task_mmu.c#L916" rel="noopener ugc nofollow" target="_blank">implementation</a>. You will notice <strong class="my gv">a do-while loop that is linear to the virtual memory area</strong>.</p><pre class="pc pd pe pf pg pr ps pt bp pu bb bk">static int show_smaps_rollup(struct seq_file *m, void *v) {<br />  …<br />  vma_start = vma-&gt;vm_start;<br />  do {<br />    smap_gather_stats(vma, &amp;mss, 0);<br />    last_vma_end = vma-&gt;vm_end;<br />    …<br />  } for_each_vma(vmi, vma);<br />  …<br />}</pre><p id="976c" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">This perfectly <strong class="my gv">explains why the function gets slower when a 2GB file is read into memory</strong>. <strong class="my gv">Because the handler of reading the <em class="oy">smaps_rollup</em> file now takes longer to run the while loop</strong>. Basically, even though <strong class="my gv"><em class="oy">smaps_rollup</em></strong> already improved the performance of getting memory information compared to the old method of parsing the <em class="oy">/proc/&lt;pid&gt;/smaps</em> file, <strong class="my gv">it is still linear to the virtual memory used</strong>.</p><h1 id="d903" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">More Quantitative Analysis</h1><p id="3a6e" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">Even though at this point the puzzle is solved, let’s conduct a more quantitative analysis. How much is the time difference when reading the <em class="oy">smaps_rollup</em> file with small versus large virtual memory utilization? Let’s write some simple benchmark code like below:</p><pre class="pc pd pe pf pg pr ps pt bp pu bb bk">import osdef read_smaps_rollup(pid):<br />  with open("/proc/{}/smaps_rollup".format(pid), "rb") as f:<br />    for line in f:<br />      passif __name__ == “__main__”:<br />  pid = os.getpid()read_smaps_rollup(pid)with open(“/root/2G_file”, “rb”) as f:<br />    data = f.read()read_smaps_rollup(pid)</pre><p id="56c3" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">This program performs the following steps:</p><ol class=""><li id="d3b3" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt qa qb qc bk">Reads the <em class="oy">smaps_rollup</em> file of the current process.</li><li id="2032" class="mw mx gu my b mz qd nb nc nd qe nf ng nh qf nj nk nl qg nn no np qh nr ns nt qa qb qc bk">Reads a 2GB file into memory.</li><li id="7966" class="mw mx gu my b mz qd nb nc nd qe nf ng nh qf nj nk nl qg nn no np qh nr ns nt qa qb qc bk">Repeats step 1.</li></ol><p id="12da" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">We then use <em class="oy">strace</em> to find the accurate time of reading the <em class="oy">smaps_rollup</em> file.</p><pre class="pc pd pe pf pg pr ps pt bp pu bb bk">$ sudo strace -T -e trace=openat,read python3 benchmark.py 2&gt;&amp;1 | grep “smaps_rollup” -A 1openat(AT_FDCWD, “/proc/3107492/smaps_rollup”, O_RDONLY|O_CLOEXEC) = 3 &lt;0.000023&gt;<br />read(3, “560b42ed4000–7ffdadcef000 — -p 0”…, 1024) = 670 &lt;0.000259&gt;<br />...<br />openat(AT_FDCWD, “/proc/3107492/smaps_rollup”, O_RDONLY|O_CLOEXEC) = 3 &lt;0.000029&gt;<br />read(3, “560b42ed4000–7ffdadcef000 — -p 0”…, 1024) = 670 &lt;0.027698&gt;</pre><p id="2e29" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">As you can see, both times, the read <em class="oy">syscall</em> returned 670, meaning the file size remained the same at 670 bytes. However, <strong class="my gv">the time it took the second time (i.e.</strong>,<strong class="my gv"> 0.027698 seconds) is 100x the time it took the first time (i.e.</strong>,<strong class="my gv"> 0.000259 seconds)</strong>! This means that if there are 98 processes, the time spent on reading this file alone will be 98 * 0.027698 = 2.7 seconds! Such a delay can significantly affect the UI experience.</p><h1 id="8c5e" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Solution</h1><p id="9ac7" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">This extension is used to display the CPU and memory usage of the notebook process on the bar at the bottom of the Notebook:</p><figure class="pc pd pe pf pg ph oz pa paragraph-image"><div class="oz pa ql"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*bNYMYTc5QQAxLyya%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*bNYMYTc5QQAxLyya%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*bNYMYTc5QQAxLyya%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*bNYMYTc5QQAxLyya%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*bNYMYTc5QQAxLyya%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*bNYMYTc5QQAxLyya%201100w,%20https://miro.medium.com/v2/resize:fit:1048/format:webp/0*bNYMYTc5QQAxLyya%201048w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 524px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*bNYMYTc5QQAxLyya 640w, https://miro.medium.com/v2/resize:fit:720/0*bNYMYTc5QQAxLyya 720w, https://miro.medium.com/v2/resize:fit:750/0*bNYMYTc5QQAxLyya 750w, https://miro.medium.com/v2/resize:fit:786/0*bNYMYTc5QQAxLyya 786w, https://miro.medium.com/v2/resize:fit:828/0*bNYMYTc5QQAxLyya 828w, https://miro.medium.com/v2/resize:fit:1100/0*bNYMYTc5QQAxLyya 1100w, https://miro.medium.com/v2/resize:fit:1048/0*bNYMYTc5QQAxLyya 1048w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 524px" /><img alt="" class="bh md pm c" width="524" height="33" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></figure><p id="0389" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">We confirmed with the user that disabling the <em class="oy">jupyter-resource-usage</em> extension meets their requirements for UI responsiveness, and that this extension is not critical to their use case. Therefore, we provided a way for them to disable the extension.</p><h1 id="2e46" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Summary</h1><p id="5cb4" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">This was such a challenging issue that required debugging from the UI all the way down to the Linux kernel. It is fascinating that the problem is linear to both the number of CPUs and the virtual memory size — two dimensions that are generally viewed separately.</p><p id="dde1" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Overall, we hope you enjoyed the irony of:</p><ol class=""><li id="b7b8" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt qa qb qc bk">The extension used to monitor CPU usage causing CPU contention.</li><li id="93e2" class="mw mx gu my b mz qd nb nc nd qe nf ng nh qf nj nk nl qg nn no np qh nr ns nt qa qb qc bk">An interesting case where the more CPUs you have, the slower you get!</li></ol><p id="dfe6" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">If you’re excited by tackling such technical challenges and have the opportunity to solve complex technical challenges and drive innovation, consider joining our <a class="af nu" href="https://explore.jobs.netflix.net/careers?query=Data+Platform&amp;pid=790298020581&amp;domain=netflix.com&amp;sort_by=relevance" rel="noopener ugc nofollow" target="_blank">Data Platform team</a>s. Be part of shaping the future of Data Security and Infrastructure, Data Developer Experience, Analytics Infrastructure and Enablement, and more. Explore the impact you can make with us!</p></div>]]></description>
      <link>https://netflixtechblog.com/investigation-of-a-workbench-ui-latency-issue-faa017b4653d</link>
      <guid>https://netflixtechblog.com/investigation-of-a-workbench-ui-latency-issue-faa017b4653d</guid>
      <pubDate>Mon, 14 Oct 2024 22:02:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Introducing Netflix TimeSeries Data Abstraction Layer]]></title>
      <description><![CDATA[<div><div></div><p id="b30e" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><a class="af nu" href="https://www.linkedin.com/in/rajiv-shringi" rel="noopener ugc nofollow" target="_blank">Rajiv Shringi</a> <a class="af nu" href="https://www.linkedin.com/in/vinaychella/" rel="noopener ugc nofollow" target="_blank">Vinay Chella</a> <a class="af nu" href="https://www.linkedin.com/in/kaidanfullerton/" rel="noopener ugc nofollow" target="_blank">Kaidan Fullerton</a> <a class="af nu" href="https://www.linkedin.com/in/oleksii-tkachuk-98b47375/" rel="noopener ugc nofollow" target="_blank">Oleksii Tkachuk</a> <a class="af nu" href="https://www.linkedin.com/in/joseph-lynch-9976a431/" rel="noopener ugc nofollow" target="_blank">Joey Lynch</a></p><h1 id="b44d" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk"><strong class="al">Introduction</strong></h1><p id="2bf4" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">As Netflix continues to expand and diversify into various sectors like <strong class="my gv">Video on Demand</strong> and <strong class="my gv">Gaming</strong>, the ability to ingest and store vast amounts of temporal data — often reaching petabytes — with millisecond access latency has become increasingly vital. In previous blog posts, we introduced the <a class="af nu" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/introducing-netflixs-key-value-data-abstraction-layer-1ea8a0a11b30"><strong class="my gv">Key-Value Data Abstraction Layer</strong></a> and the <a class="af nu" href="https://netflixtechblog.medium.com/data-gateway-a-platform-for-growing-and-protecting-the-data-tier-f1ed8db8f5c6" rel="noopener"><strong class="my gv">Data Gateway Platform</strong></a>, both of which are integral to Netflix’s data architecture. The Key-Value Abstraction offers a flexible, scalable solution for storing and accessing structured key-value data, while the Data Gateway Platform provides essential infrastructure for protecting, configuring, and deploying the data tier.</p><p id="f295" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Building on these foundational abstractions, we developed the <strong class="my gv">TimeSeries Abstraction</strong> — a versatile and scalable solution designed to efficiently store and query large volumes of temporal event data with low millisecond latencies, all in a cost-effective manner across various use cases.</p><p id="f9ce" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">In this post, we will delve into the architecture, design principles, and real-world applications of the <strong class="my gv">TimeSeries Abstraction</strong>, demonstrating how it enhances our platform’s ability to manage temporal data at scale.</p><p id="f726" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Note: </strong><em class="oy">Contrary to what the name may suggest, this system is not built as a general-purpose time series database. We do not use it for metrics, histograms, timers, or any such near-real time analytics use case. Those use cases are well served by the Netflix </em><a class="af nu" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/introducing-atlas-netflixs-primary-telemetry-platform-bd31f4d8ed9a"><em class="oy">Atlas</em></a><em class="oy"> telemetry system. Instead, we focus on addressing the challenge of storing and accessing extremely high-throughput, immutable temporal event data in a low-latency and cost-efficient manner.</em></p><h1 id="a578" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Challenges</h1><p id="ce8f" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">At Netflix, temporal data is continuously generated and utilized, whether from user interactions like video-play events, asset impressions, or complex micro-service network activities. Effectively managing this data at scale to extract valuable insights is crucial for ensuring optimal user experiences and system reliability.</p><p id="11e4" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">However, storing and querying such data presents a unique set of challenges:</p><ul class=""><li id="ef3e" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt oz pa pb bk"><strong class="my gv">High Throughput</strong>: Managing up to 10 million writes per second while maintaining high availability.</li><li id="35ee" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk"><strong class="my gv">Efficient Querying in Large Datasets</strong>: Storing petabytes of data while ensuring primary key reads return results within low double-digit milliseconds, and supporting searches and aggregations across multiple secondary attributes.</li><li id="6d28" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk"><strong class="my gv">Global Reads and Writes</strong>: Facilitating read and write operations from anywhere in the world with adjustable consistency models.</li><li id="89d9" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk"><strong class="my gv">Tunable Configuration</strong>: Offering the ability to partition datasets in either a single-tenant or multi-tenant datastore, with options to adjust various dataset aspects such as retention and consistency.</li><li id="78f2" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk"><strong class="my gv">Handling Bursty Traffic</strong>: Managing significant traffic spikes during high-demand events, such as new content launches or regional failovers.</li><li id="7ea6" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk"><strong class="my gv">Cost Efficiency</strong>: Reducing the cost per byte and per operation to optimize long-term retention while minimizing infrastructure expenses, which can amount to millions of dollars for Netflix.</li></ul><h1 id="264f" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">TimeSeries Abstraction</h1><p id="dc2a" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">The TimeSeries Abstraction was developed to meet these requirements, built around the following core design principles:</p><ul class=""><li id="1e7d" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt oz pa pb bk"><strong class="my gv">Partitioned Data</strong>: Data is partitioned using a unique temporal partitioning strategy combined with an event bucketing approach to efficiently manage bursty workloads and streamline queries.</li><li id="caa3" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk"><strong class="my gv">Flexible Storage</strong>: The service is designed to integrate with various storage backends, including <a class="af nu" href="https://cassandra.apache.org/_/index.html" rel="noopener ugc nofollow" target="_blank">Apache Cassandra</a> and <a class="af nu" href="https://www.elastic.co/elasticsearch" rel="noopener ugc nofollow" target="_blank">Elasticsearch</a>, allowing Netflix to customize storage solutions based on specific use case requirements.</li><li id="3db7" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk"><strong class="my gv">Configurability</strong>: TimeSeries offers a range of tunable options for each dataset, providing the flexibility needed to accommodate a wide array of use cases.</li><li id="b931" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk"><strong class="my gv">Scalability</strong>: The architecture supports both horizontal and vertical scaling, enabling the system to handle increasing throughput and data volumes as Netflix expands its user base and services.</li><li id="b382" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk"><strong class="my gv">Sharded Infrastructure</strong>: Leveraging the <strong class="my gv">Data Gateway Platform</strong>, we can deploy single-tenant and/or multi-tenant infrastructure with the necessary access and traffic isolation.</li></ul><p id="2a47" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Let’s dive into the various aspects of this abstraction.</p><h1 id="ea66" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Data Model</h1><p id="fbf3" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">We follow a unique event data model that encapsulates all the data we want to capture for events, while allowing us to query them efficiently.</p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi pj"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*jl30Jl559Fnd29in%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*jl30Jl559Fnd29in%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*jl30Jl559Fnd29in%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*jl30Jl559Fnd29in%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*jl30Jl559Fnd29in%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*jl30Jl559Fnd29in%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*jl30Jl559Fnd29in%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*jl30Jl559Fnd29in 640w, https://miro.medium.com/v2/resize:fit:720/0*jl30Jl559Fnd29in 720w, https://miro.medium.com/v2/resize:fit:750/0*jl30Jl559Fnd29in 750w, https://miro.medium.com/v2/resize:fit:786/0*jl30Jl559Fnd29in 786w, https://miro.medium.com/v2/resize:fit:828/0*jl30Jl559Fnd29in 828w, https://miro.medium.com/v2/resize:fit:1100/0*jl30Jl559Fnd29in 1100w, https://miro.medium.com/v2/resize:fit:1400/0*jl30Jl559Fnd29in 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="342" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="7228" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Let’s start with the smallest unit of data in the abstraction and work our way up.</p><ul class=""><li id="cf78" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt oz pa pb bk"><strong class="my gv">Event Item</strong>: An event item is a key-value pair that users use to store data for a given event. For example: <em class="oy">{“device_type”: “ios”}</em>.</li><li id="55c6" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk"><strong class="my gv">Event</strong>: An event is a structured collection of one or more such event items. An event occurs at a specific point in time and is identified by a client-generated timestamp and an event identifier (such as a UUID). This combination of <strong class="my gv">event_time</strong> and <strong class="my gv">event_id</strong> also forms part of the unique idempotency key for the event, enabling users to safely retry requests.</li><li id="9145" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk"><strong class="my gv">Time Series ID</strong>: A <strong class="my gv">time_series_id</strong> is a collection of one or more such events over the dataset’s retention period. For instance, a <strong class="my gv">device_id</strong> would store all events occurring for a given device over the retention period. All events are immutable, and the TimeSeries service only ever appends events to a given time series ID.</li><li id="ef6e" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk"><strong class="my gv">Namespace</strong>: A namespace is a collection of time series IDs and event data, representing the complete TimeSeries dataset. Users can create one or more namespaces for each of their use cases. The abstraction applies various tunable options at the namespace level, which we will discuss further when we explore the service’s control plane.</li></ul><h1 id="eda3" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">API</h1><p id="a85a" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">The abstraction provides the following APIs to interact with the event data.</p><ul class=""><li id="8d2a" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt oz pa pb bk"><strong class="my gv">WriteEventRecordsSync</strong>: This endpoint writes a batch of events and sends back a durability acknowledgement to the client. This is used in cases where users require a guarantee of durability.</li><li id="9513" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk"><strong class="my gv">WriteEventRecords</strong>: This is the fire-and-forget version of the above endpoint. It enqueues a batch of events without the durability acknowledgement. This is used in cases like logging or tracing, where users care more about throughput and can tolerate a small amount of data loss.</li></ul><pre class="pk pl pm pn po pv pw px bp py bb bk">{<br />  "namespace": "my_dataset",<br />  "events": [<br />    {<br />      "timeSeriesId": "profile100",<br />      "eventTime": "2024-10-03T21:24:23.988Z",<br />      "eventId": "550e8400-e29b-41d4-a716-446655440000",<br />      "eventItems": [<br />        {<br />          "eventItemKey": "ZGV2aWNlVHlwZQ==",  <br />          "eventItemValue": "aW9z"<br />        },<br />        {<br />          "eventItemKey": "ZGV2aWNlTWV0YWRhdGE=",<br />          "eventItemValue": "c29tZSBtZXRhZGF0YQ=="<br />        }<br />      ]<br />    },<br />    {<br />      "timeSeriesId": "profile100",<br />      "eventTime": "2024-10-03T21:23:30.000Z",<br />      "eventId": "123e4567-e89b-12d3-a456-426614174000",<br />      "eventItems": [<br />        {<br />          "eventItemKey": "ZGV2aWNlVHlwZQ==",  <br />          "eventItemValue": "YW5kcm9pZA=="<br />        }<br />      ]<br />    }<br />  ]<br />}</pre><p id="4e72" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">ReadEventRecords</strong>: Given a combination of a namespace, a timeSeriesId, a timeInterval, and optional eventFilters, this endpoint returns all the matching events, sorted descending by event_time, with low millisecond latency.</p><pre class="pk pl pm pn po pv pw px bp py bb bk">{<br />  "namespace": "my_dataset",<br />  "timeSeriesId": "profile100",<br />  "timeInterval": {<br />    "start": "2024-10-02T21:00:00.000Z",<br />    "end":   "2024-10-03T21:00:00.000Z"<br />  },<br />  "eventFilters": [<br />    {<br />      "matchEventItemKey": "ZGV2aWNlVHlwZQ==",<br />      "matchEventItemValue": "aW9z"<br />    }<br />  ],<br />  "pageSize": 100,<br />  "totalRecordLimit": 1000<br />}</pre><p id="d4d9" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">SearchEventRecords</strong>: Given a search criteria and a time interval, this endpoint returns all the matching events. These use cases are fine with eventually consistent reads.</p><pre class="pk pl pm pn po pv pw px bp py bb bk">{<br />  "namespace": "my_dataset",<br />  "timeInterval": {<br />    "start": "2024-10-02T21:00:00.000Z",<br />    "end": "2024-10-03T21:00:00.000Z"<br />  },<br />  "searchQuery": {<br />    "booleanQuery": {<br />      "searchQuery": [<br />        {<br />          "equals": {<br />            "eventItemKey": "deviceType",<br />            "eventItemValue": "aW9z"<br />          }<br />        },<br />        {<br />          "equals": {<br />            "eventItemKey": "deviceType",<br />            "eventItemValue": "YW5kcm9pZA=="<br />          }<br />        }<br />      ],<br />      "operator": "OR"<br />    }<br />  },<br />  "pageSize": 100,<br />  "totalRecordLimit": 1000<br />}</pre><p id="9863" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">AggregateEventRecords</strong>: Given a search criteria and an aggregation mode (e.g. DistinctAggregation) , this endpoint performs the given aggregation within a given time interval. Similar to the Search endpoint, users can tolerate eventual consistency and a potentially higher latency (in seconds).</p><pre class="pk pl pm pn po pv pw px bp py bb bk">{<br />  "namespace": "my_dataset",<br />  "timeInterval": {<br />    "start": "2024-10-02T21:00:00.000Z",<br />    "end": "2024-10-03T21:00:00.000Z"<br />  },<br />  "searchQuery": {...some search criteria...},<br />  "aggregationQuery": {<br />    "distinct": {<br />      "eventItemKey": "deviceType",<br />      "pageSize": 100<br />    }<br />  }<br />}</pre><p id="626b" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">In the subsequent sections, we will talk about how we interact with this data at the storage layer.</p><h1 id="006b" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Storage Layer</h1><p id="39dd" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">The storage layer for TimeSeries comprises a primary data store and an optional index data store. The primary data store ensures data durability during writes and is used for primary read operations, while the index data store is utilized for search and aggregate operations. At Netflix, <strong class="my gv">Apache Cassandra</strong> is the preferred choice for storing durable data in high-throughput scenarios, while <strong class="my gv">Elasticsearch</strong> is the preferred data store for indexing. However, similar to our approach with the API, the storage layer is not tightly coupled to these specific data stores. Instead, we define storage API contracts that must be fulfilled, allowing us the flexibility to replace the underlying data stores as needed.</p><h1 id="6efb" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Primary Datastore</h1><p id="3689" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">In this section, we will talk about how we leverage <strong class="my gv">Apache Cassandra</strong> for TimeSeries use cases.</p><h2 id="9b30" class="qe nw gu bf nx qf qg dy ob qh qi ea of nh qj qk ql nl qm qn qo np qp qq qr qs bk">Partitioning Scheme</h2><p id="6241" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">At Netflix’s scale, the continuous influx of event data can quickly overwhelm traditional databases. Temporal partitioning addresses this challenge by dividing the data into manageable chunks based on time intervals, such as hourly, daily, or monthly windows. This approach enables efficient querying of specific time ranges without the need to scan the entire dataset. It also allows Netflix to archive, compress, or delete older data efficiently, optimizing both storage and query performance. Additionally, this partitioning mitigates the performance issues typically associated with <a class="af nu" href="https://thelastpickle.com/blog/2019/01/11/wide-partitions-cassandra-3-11.html" rel="noopener ugc nofollow" target="_blank">wide partitions</a> in Cassandra. By employing this strategy, we can operate at much higher disk utilization, as it reduces the need to reserve large amounts of disk space for compactions, thereby saving costs.</p><p id="04ec" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Here is what it looks like :</p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi pj"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*MxuEH6_pOVDcAMie%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*MxuEH6_pOVDcAMie%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*MxuEH6_pOVDcAMie%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*MxuEH6_pOVDcAMie%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*MxuEH6_pOVDcAMie%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*MxuEH6_pOVDcAMie%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*MxuEH6_pOVDcAMie%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*MxuEH6_pOVDcAMie 640w, https://miro.medium.com/v2/resize:fit:720/0*MxuEH6_pOVDcAMie 720w, https://miro.medium.com/v2/resize:fit:750/0*MxuEH6_pOVDcAMie 750w, https://miro.medium.com/v2/resize:fit:786/0*MxuEH6_pOVDcAMie 786w, https://miro.medium.com/v2/resize:fit:828/0*MxuEH6_pOVDcAMie 828w, https://miro.medium.com/v2/resize:fit:1100/0*MxuEH6_pOVDcAMie 1100w, https://miro.medium.com/v2/resize:fit:1400/0*MxuEH6_pOVDcAMie 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="335" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="ae71" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Time Slice: </strong>Atime slice is the unit of data retention and maps directly to a Cassandra table. We create multiple such time slices, each covering a specific interval of time. An event lands in one of these slices based on the <strong class="my gv">event_time</strong>. These slices are joined with <em class="oy">no time gaps</em>in between, with operations being <em class="oy">start-inclusive</em> and <em class="oy">end-exclusive</em>, ensuring that all data lands in one of the slices.</p><p id="05ba" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Why not use row-based Time-To-Live (TTL)?</strong></p><p id="bd01" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Using TTL on individual events would generate a significant number of <a class="af nu" href="https://thelastpickle.com/blog/2016/07/27/about-deletes-and-tombstones.html" rel="noopener ugc nofollow" target="_blank">tombstones</a> in Cassandra, degrading performance, especially during range scans. By employing discrete time slices and dropping them, we avoid the tombstone issue entirely. The tradeoff is that data may be retained slightly longer than necessary, as an entire table’s time range must fall outside the retention window before it can be dropped. Additionally, TTLs are difficult to adjust later, whereas TimeSeries can extend the dataset retention instantly with a single control plane operation.</p><p id="fc7e" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Time Buckets</strong>: Within a time slice, data is further partitioned into time buckets. This facilitates effective range scans by allowing us to target specific time buckets for a given query range. The tradeoff is that if a user wants to read the entire range of data over a large time period, we must scan many partitions. We mitigate potential latency by scanning these partitions in parallel and aggregating the data at the end. In most cases, the advantage of targeting smaller data subsets outweighs the read amplification from these scatter-gather operations. Typically, users read a smaller subset of data rather than the entire retention range.</p><p id="17b2" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Event Buckets</strong>: To manage extremely high-throughput write operations, which may result in a burst of writes for a given time series within a short period, we further divide the time bucket into event buckets. This prevents overloading the same partition for a given time range and also reduces partition sizes further, albeit with a slight increase in read amplification.</p><p id="ab2a" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Note</strong>: <em class="oy">With Cassandra 4.x onwards, we notice a substantial improvement in the performance of scanning a range of data in a wide partition. See </em><strong class="my gv"><em class="oy">Future Enhancements</em></strong><em class="oy"> at the end to see the </em><strong class="my gv"><em class="oy">Dynamic Event bucketing</em></strong><em class="oy"> work that aims to take advantage of this.</em></p><h2 id="e7cd" class="qe nw gu bf nx qf qg dy ob qh qi ea of nh qj qk ql nl qm qn qo np qp qq qr qs bk">Storage Tables</h2><p id="baaa" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">We use two kinds of tables</p><ul class=""><li id="1406" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt oz pa pb bk"><strong class="my gv">Data tables</strong>: These are the time slices that store the actual event data.</li><li id="a159" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk"><strong class="my gv">Metadata table</strong>: This table stores information about how each time slice is configured <em class="oy">per namespace</em>.</li></ul><h2 id="51df" class="qe nw gu bf nx qf qg dy ob qh qi ea of nh qj qk ql nl qm qn qo np qp qq qr qs bk">Data tables</h2><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi pj"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*ktuEBzveeK4f1mWH%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*ktuEBzveeK4f1mWH%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*ktuEBzveeK4f1mWH%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*ktuEBzveeK4f1mWH%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*ktuEBzveeK4f1mWH%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*ktuEBzveeK4f1mWH%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*ktuEBzveeK4f1mWH%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*ktuEBzveeK4f1mWH 640w, https://miro.medium.com/v2/resize:fit:720/0*ktuEBzveeK4f1mWH 720w, https://miro.medium.com/v2/resize:fit:750/0*ktuEBzveeK4f1mWH 750w, https://miro.medium.com/v2/resize:fit:786/0*ktuEBzveeK4f1mWH 786w, https://miro.medium.com/v2/resize:fit:828/0*ktuEBzveeK4f1mWH 828w, https://miro.medium.com/v2/resize:fit:1100/0*ktuEBzveeK4f1mWH 1100w, https://miro.medium.com/v2/resize:fit:1400/0*ktuEBzveeK4f1mWH 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="430" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="4c23" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">The partition key enables splitting events for a <strong class="my gv">time_series_id</strong> over a range of <strong class="my gv">time_bucket(s)</strong> and <strong class="my gv">event_bucket(s)</strong>, thus mitigating hot partitions, while the clustering key allows us to keep data sorted on disk in the order we almost always want to read it. The <strong class="my gv">value_metadata</strong> column stores metadata for the <strong class="my gv">event_item_value</strong> such as compression.</p><p id="654d" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Writing to the data table:</strong></p><p id="524c" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">User writes will land in a given time slice, time bucket, and event bucket as a factor of the <strong class="my gv">event_time</strong> attached to the event. This factor is dictated by the control plane configuration of a given namespace.</p><p id="2892" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">For example:</p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi qt"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*P4IThIE_PE9F8KYi%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*P4IThIE_PE9F8KYi%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*P4IThIE_PE9F8KYi%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*P4IThIE_PE9F8KYi%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*P4IThIE_PE9F8KYi%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*P4IThIE_PE9F8KYi%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*P4IThIE_PE9F8KYi%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*P4IThIE_PE9F8KYi 640w, https://miro.medium.com/v2/resize:fit:720/0*P4IThIE_PE9F8KYi 720w, https://miro.medium.com/v2/resize:fit:750/0*P4IThIE_PE9F8KYi 750w, https://miro.medium.com/v2/resize:fit:786/0*P4IThIE_PE9F8KYi 786w, https://miro.medium.com/v2/resize:fit:828/0*P4IThIE_PE9F8KYi 828w, https://miro.medium.com/v2/resize:fit:1100/0*P4IThIE_PE9F8KYi 1100w, https://miro.medium.com/v2/resize:fit:1400/0*P4IThIE_PE9F8KYi 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="48" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="79ad" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Reading from the data table:</strong></p><p id="a48c" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">The below illustration depicts at a high-level on how we scatter-gather the reads from multiple partitions and join the result set at the end to return the final result.</p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi pj"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*a805txbeIDqYP73d%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*a805txbeIDqYP73d%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*a805txbeIDqYP73d%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*a805txbeIDqYP73d%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*a805txbeIDqYP73d%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*a805txbeIDqYP73d%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*a805txbeIDqYP73d%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*a805txbeIDqYP73d 640w, https://miro.medium.com/v2/resize:fit:720/0*a805txbeIDqYP73d 720w, https://miro.medium.com/v2/resize:fit:750/0*a805txbeIDqYP73d 750w, https://miro.medium.com/v2/resize:fit:786/0*a805txbeIDqYP73d 786w, https://miro.medium.com/v2/resize:fit:828/0*a805txbeIDqYP73d 828w, https://miro.medium.com/v2/resize:fit:1100/0*a805txbeIDqYP73d 1100w, https://miro.medium.com/v2/resize:fit:1400/0*a805txbeIDqYP73d 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="494" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><h2 id="0c7b" class="qe nw gu bf nx qf qg dy ob qh qi ea of nh qj qk ql nl qm qn qo np qp qq qr qs bk">Metadata table</h2><p id="20b0" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">This table stores the configuration data about the time slices for a given namespace.</p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi pj"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*asJFOjl1iwlSajJc%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*asJFOjl1iwlSajJc%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*asJFOjl1iwlSajJc%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*asJFOjl1iwlSajJc%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*asJFOjl1iwlSajJc%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*asJFOjl1iwlSajJc%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*asJFOjl1iwlSajJc%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*asJFOjl1iwlSajJc 640w, https://miro.medium.com/v2/resize:fit:720/0*asJFOjl1iwlSajJc 720w, https://miro.medium.com/v2/resize:fit:750/0*asJFOjl1iwlSajJc 750w, https://miro.medium.com/v2/resize:fit:786/0*asJFOjl1iwlSajJc 786w, https://miro.medium.com/v2/resize:fit:828/0*asJFOjl1iwlSajJc 828w, https://miro.medium.com/v2/resize:fit:1100/0*asJFOjl1iwlSajJc 1100w, https://miro.medium.com/v2/resize:fit:1400/0*asJFOjl1iwlSajJc 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="317" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="6fce" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Note the following:</p><ul class=""><li id="217e" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt oz pa pb bk"><strong class="my gv">No Time Gaps</strong>: The end_time of a given time slice overlaps with the start_time of the next time slice, ensuring all events find a home.</li><li id="bac3" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk"><strong class="my gv">Retention</strong>: The status indicates which tables fall inside and outside of the retention window.</li><li id="e410" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk"><strong class="my gv">Flexible</strong>: This metadata can be adjusted per time slice, allowing us to tune the partition settings of future time slices based on observed data patterns in the current time slice.</li></ul><p id="f612" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">There is a lot more information that can be stored into the <strong class="my gv">metadata</strong> column (e.g., compaction settings for the table), but we only show the partition settings here for brevity.</p><h1 id="a523" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Index Datastore</h1><p id="5bab" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">To support secondary access patterns via non-primary key attributes, we index data into Elasticsearch. Users can configure a list of attributes per namespace that they wish to search and/or aggregate data on. The service extracts these fields from events as they stream in, indexing the resultant documents into Elasticsearch. Depending on the throughput, we may use Elasticsearch as a reverse index, retrieving the full data from Cassandra, or we may store the entire source data directly in Elasticsearch.</p><p id="4df3" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv">Note</strong>:<em class="oy"> Again, users are never directly exposed to Elasticsearch, just like they are not directly exposed to Cassandra. Instead, they interact with the Search and Aggregate API endpoints that translate a given query to that needed for the underlying datastore.</em></p><p id="3b37" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">In the next section, we will talk about how we configure these data stores for different datasets.</p><h1 id="1edb" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Control Plane</h1><p id="1f56" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">The data plane is responsible for executing the read and write operations, while the control plane configures every aspect of a namespace’s behavior. The data plane communicates with the TimeSeries control stack, which manages this configuration information. In turn, the TimeSeries control stack interacts with a sharded <strong class="my gv">Data Gateway Platform Control Plane</strong> that oversees control configurations for all abstractions and namespaces.</p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi qu"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*aB6OKXoG-mT65Vh1%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*aB6OKXoG-mT65Vh1%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*aB6OKXoG-mT65Vh1%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*aB6OKXoG-mT65Vh1%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*aB6OKXoG-mT65Vh1%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*aB6OKXoG-mT65Vh1%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*aB6OKXoG-mT65Vh1%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*aB6OKXoG-mT65Vh1 640w, https://miro.medium.com/v2/resize:fit:720/0*aB6OKXoG-mT65Vh1 720w, https://miro.medium.com/v2/resize:fit:750/0*aB6OKXoG-mT65Vh1 750w, https://miro.medium.com/v2/resize:fit:786/0*aB6OKXoG-mT65Vh1 786w, https://miro.medium.com/v2/resize:fit:828/0*aB6OKXoG-mT65Vh1 828w, https://miro.medium.com/v2/resize:fit:1100/0*aB6OKXoG-mT65Vh1 1100w, https://miro.medium.com/v2/resize:fit:1400/0*aB6OKXoG-mT65Vh1 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="499" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="7f9c" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Separating the responsibilities of the data plane and control plane helps maintain the high availability of our data plane, as the control plane takes on tasks that may require some form of schema consensus from the underlying data stores.</p><h1 id="cd98" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Namespace Configuration</h1><p id="7d27" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">The below configuration snippet demonstrates the immense flexibility of the service and how we can tune several things per namespace using our control plane.</p><pre class="pk pl pm pn po pv pw px bp py bb bk">"persistence_configuration": [<br />  {<br />    "id": "PRIMARY_STORAGE",<br />    "physical_storage": {<br />      "type": "CASSANDRA",                  // type of primary storage<br />      "cluster": "cass_dgw_ts_tracing",     // physical cluster name<br />      "dataset": "tracing_default"          // maps to the keyspace<br />    },<br />    "config": {<br />      "timePartition": {<br />        "secondsPerTimeSlice": "129600",    // width of a time slice<br />        "secondPerTimeBucket": "3600",      // width of a time bucket<br />        "eventBuckets": 4                   // how many event buckets within<br />      },<br />      "queueBuffering": {<br />        "coalesce": "1s",                   // how long to coalesce writes<br />        "bufferCapacity": 4194304           // queue capacity in bytes<br />      },<br />      "consistencyScope": "LOCAL",          // single-region/multi-region<br />      "consistencyTarget": "EVENTUAL",      // read/write consistency<br />      "acceptLimit": "129600s"              // how far back writes are allowed<br />    },<br />    "lifecycleConfigs": {<br />      "lifecycleConfig": [                  // Primary store data retention<br />        {<br />          "type": "retention",<br />          "config": {<br />            "close_after": "1296000s",      // close for reads/writes<br />            "delete_after": "1382400s"      // drop time slice<br />          }<br />        }<br />      ]<br />    }<br />  },<br />  {<br />    "id": "INDEX_STORAGE",<br />    "physicalStorage": {<br />      "type": "ELASTICSEARCH",              // type of index storage<br />      "cluster": "es_dgw_ts_tracing",       // ES cluster name<br />      "dataset": "tracing_default_useast1"  // base index name<br />    },<br />    "config": {<br />      "timePartition": {<br />        "secondsPerSlice": "129600"         // width of the index slice<br />      },<br />      "consistencyScope": "LOCAL",<br />      "consistencyTarget": "EVENTUAL",      // how should we read/write data<br />      "acceptLimit": "129600s",             // how far back writes are allowed<br />      "indexConfig": {<br />        "fieldMapping": {                   // fields to extract to index<br />          "tags.nf.app": "KEYWORD",<br />          "tags.duration": "INTEGER",<br />          "tags.enabled": "BOOLEAN"<br />        },<br />        "refreshInterval": "60s"            // Index related settings<br />      }<br />    },<br />    "lifecycleConfigs": {<br />      "lifecycleConfig": [<br />        {<br />          "type": "retention",              // Index retention settings<br />          "config": {<br />            "close_after": "1296000s",<br />            "delete_after": "1382400s"<br />          }<br />        }<br />      ]<br />    }<br />  }<br />]</pre><h1 id="fcef" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Provisioning Infrastructure</h1><p id="85c4" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">With so many different parameters, we need automated provisioning workflows to deduce the best settings for a given workload. When users want to create their namespaces, they specify a list of <em class="oy">workload</em> <em class="oy">desires</em>, which the automation translates into concrete infrastructure and related control plane configuration. We highly encourage you to watch this <a class="af nu" href="https://www.youtube.com/watch?v=2aBVKXi8LKk" rel="noopener ugc nofollow" target="_blank">ApacheCon talk</a>, by one of our stunning colleagues <strong class="my gv">Joey Lynch,</strong> on how we achieve this. We may go into detail on this subject in one of our future blog posts.</p><p id="311d" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Once the system provisions the initial infrastructure, it then scales in response to the user workload. The next section describes how this is achieved.</p><h1 id="5890" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Scalability</h1><p id="561f" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">Our users may operate with limited information at the time of provisioning their namespaces, resulting in best-effort provisioning estimates. Further, evolving use-cases may introduce new throughput requirements over time. Here’s how we manage this:</p><ul class=""><li id="37e7" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt oz pa pb bk"><strong class="my gv">Horizontal scaling</strong>: TimeSeries server instances can auto-scale up and down as per attached scaling policies to meet the traffic demand. The storage server capacity can be recomputed to accommodate changing requirements using our <a class="af nu" href="https://github.com/Netflix-Skunkworks/service-capacity-modeling/tree/main/service_capacity_modeling" rel="noopener ugc nofollow" target="_blank">capacity planner</a>.</li><li id="6f26" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk"><strong class="my gv">Vertical scaling</strong>: We may also choose to vertically scale our TimeSeries server instances or our storage instances to get greater CPU, RAM and/or attached storage capacity.</li><li id="f575" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk"><strong class="my gv">Scaling disk</strong>: We may attach <a class="af nu" href="https://aws.amazon.com/ebs/" rel="noopener ugc nofollow" target="_blank">EBS</a> to store data if the capacity planner prefers infrastructure that offers larger storage at a lower cost rather than SSDs optimized for latency. In such cases, we deploy jobs to scale the EBS volume when the disk storage reaches a certain percentage threshold.</li><li id="e2bd" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk"><strong class="my gv">Re-partitioning data</strong>: Inaccurate workload estimates can lead to over or under-partitioning of our datasets. TimeSeries control-plane can adjust the partitioning configuration for upcoming time slices, once we realize the nature of data in the wild (via partition histograms). In the future we plan to support re-partitioning of older data and dynamic partitioning of current data.</li></ul><h1 id="bdec" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Design Principles</h1><p id="b328" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">So far, we have seen how TimeSeries stores, configures and interacts with event datasets. Let’s see how we apply different techniques to improve the performance of our operations and provide better guarantees.</p><h1 id="737f" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Event Idempotency</h1><p id="83da" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">We prefer to bake in idempotency in all mutation endpoints, so that users can retry or hedge their requests safely. <a class="af nu" href="https://research.google/pubs/the-tail-at-scale/" rel="noopener ugc nofollow" target="_blank">Hedging</a> is when the client sends an identical competing request to the server, if the original request does not come back with a response in an expected amount of time. The client then responds with whichever request completes first. This is done to keep the tail latencies for an application relatively low. This can only be done safely if the mutations are idempotent. For TimeSeries, the combination of <strong class="my gv">event_time</strong>, <strong class="my gv">event_id</strong> and <strong class="my gv">event_item_key</strong> form the idempotency key for a given <strong class="my gv">time_series_id</strong> event.</p><h1 id="9ffa" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">SLO-based Hedging</h1><p id="751c" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">We assign Service Level Objectives (SLO) targets for different endpoints within TimeSeries, as an indication of what we think the performance of those endpoints should be <em class="oy">for a given namespace</em>. We can then hedge a request if the response does not come back in that configured amount of time.</p><pre class="pk pl pm pn po pv pw px bp py bb bk">"slos": {<br />  "read": {               // SLOs per endpoint<br />    "latency": {<br />      "target": "0.5s",   // hedge around this number<br />      "max": "1s"         // time-out around this number<br />    }<br />  },<br />  "write": {<br />    "latency": {<br />      "target": "0.01s",<br />      "max": "0.05s"<br />    }<br />  }<br />}</pre><h1 id="4152" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Partial Return</h1><p id="87eb" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">Sometimes, a client may be sensitive to latency and willing to accept a partial result set. A real-world example of this is real-time frequency capping. Precision is not critical in this case, but if the response is delayed, it becomes practically useless to the upstream client. Therefore, the client prefers to work with whatever data has been collected so far rather than timing out while waiting for all the data. The TimeSeries client supports partial returns around SLOs for this purpose. Importantly, we still maintain the latest order of events in this partial fetch.</p><h1 id="1d8b" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Adaptive Pagination</h1><p id="fabf" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">All reads start with a default fanout factor, scanning 8 partition buckets in parallel. However, if the service layer determines that the time_series dataset is dense — i.e., most reads are satisfied by reading the first few partition buckets — then it dynamically adjusts the fanout factor of future reads in order to reduce the read amplification on the underlying datastore. Conversely, if the dataset is sparse, we may want to increase this limit with a reasonable upper bound.</p><h1 id="7fcd" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Limited Write Window</h1><p id="d732" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">In most cases, the active range for writing data is smaller than the range for reading data — i.e., we want a range of time to become immutable as soon as possible so that we can apply optimizations on top of it. We control this by having a configurable “<strong class="my gv">acceptLimit</strong>” parameter that prevents users from writing events older than this time limit. For example, an accept limit of 4 hours means that users cannot write events older than <em class="oy">now() — 4 hours</em>. We sometimes raise this limit for backfilling historical data, but it is tuned back down for regular write operations. Once a range of data becomes immutable, we can safely do things like caching, compressing, and compacting it for reads.</p><h1 id="7223" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Buffering Writes</h1><p id="21e8" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">We frequently leverage this service for handling bursty workloads. Rather than overwhelming the underlying datastore with this load all at once, we aim to distribute it more evenly by allowing events to coalesce over short durations (typically seconds). These events accumulate in in-memory queues running on each instance. Dedicated consumers then steadily drain these queues, grouping the events by their partition key, and batching the writes to the underlying datastore.</p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi pj"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*pMVe_h3daBDLWdis%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*pMVe_h3daBDLWdis%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*pMVe_h3daBDLWdis%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*pMVe_h3daBDLWdis%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*pMVe_h3daBDLWdis%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*pMVe_h3daBDLWdis%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*pMVe_h3daBDLWdis%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*pMVe_h3daBDLWdis 640w, https://miro.medium.com/v2/resize:fit:720/0*pMVe_h3daBDLWdis 720w, https://miro.medium.com/v2/resize:fit:750/0*pMVe_h3daBDLWdis 750w, https://miro.medium.com/v2/resize:fit:786/0*pMVe_h3daBDLWdis 786w, https://miro.medium.com/v2/resize:fit:828/0*pMVe_h3daBDLWdis 828w, https://miro.medium.com/v2/resize:fit:1100/0*pMVe_h3daBDLWdis 1100w, https://miro.medium.com/v2/resize:fit:1400/0*pMVe_h3daBDLWdis 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="322" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="fa07" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">The queues are tailored to each datastore since their operational characteristics depend on the specific datastore being written to. For instance, the batch size for writing to Cassandra is significantly smaller than that for indexing into Elasticsearch, leading to different drain rates and batch sizes for the associated consumers.</p><p id="e4da" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">While using in-memory queues does increase JVM garbage collection, we have experienced substantial improvements by transitioning to JDK 21 with ZGC. To illustrate the impact, ZGC has reduced our tail latencies by an impressive 86%:</p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi pj"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*hj98LMk1UddaaDs-%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*hj98LMk1UddaaDs-%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*hj98LMk1UddaaDs-%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*hj98LMk1UddaaDs-%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*hj98LMk1UddaaDs-%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*hj98LMk1UddaaDs-%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*hj98LMk1UddaaDs-%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*hj98LMk1UddaaDs- 640w, https://miro.medium.com/v2/resize:fit:720/0*hj98LMk1UddaaDs- 720w, https://miro.medium.com/v2/resize:fit:750/0*hj98LMk1UddaaDs- 750w, https://miro.medium.com/v2/resize:fit:786/0*hj98LMk1UddaaDs- 786w, https://miro.medium.com/v2/resize:fit:828/0*hj98LMk1UddaaDs- 828w, https://miro.medium.com/v2/resize:fit:1100/0*hj98LMk1UddaaDs- 1100w, https://miro.medium.com/v2/resize:fit:1400/0*hj98LMk1UddaaDs- 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="345" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="e67f" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Because we use in-memory queues, we are prone to losing events in case of an instance crash. As such, these queues are only used for use cases that can tolerate some amount of data loss .e.g. tracing/logging. For use cases that need guaranteed durability and/or read-after-write consistency, these queues are effectively disabled and writes are flushed to the data store almost immediately.</p><h1 id="3b16" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Dynamic Compaction</h1><p id="b018" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">Once a time slice exits the active write window, we can leverage the immutability of the data to optimize it for read performance. This process may involve re-compacting immutable data using optimal compaction strategies, dynamically shrinking and/or splitting shards to optimize system resources, and other similar techniques to ensure fast and reliable performance.</p><p id="dbe4" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">The following section provides a glimpse into the real-world performance of some of our TimeSeries datasets.</p><h1 id="dd6b" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Real-world Performance</h1><p id="1625" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">The service can write data in the order of low single digit milliseconds</p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi pj"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*VWrQj2ya5PQWusBq%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*VWrQj2ya5PQWusBq%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*VWrQj2ya5PQWusBq%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*VWrQj2ya5PQWusBq%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*VWrQj2ya5PQWusBq%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*VWrQj2ya5PQWusBq%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*VWrQj2ya5PQWusBq%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*VWrQj2ya5PQWusBq 640w, https://miro.medium.com/v2/resize:fit:720/0*VWrQj2ya5PQWusBq 720w, https://miro.medium.com/v2/resize:fit:750/0*VWrQj2ya5PQWusBq 750w, https://miro.medium.com/v2/resize:fit:786/0*VWrQj2ya5PQWusBq 786w, https://miro.medium.com/v2/resize:fit:828/0*VWrQj2ya5PQWusBq 828w, https://miro.medium.com/v2/resize:fit:1100/0*VWrQj2ya5PQWusBq 1100w, https://miro.medium.com/v2/resize:fit:1400/0*VWrQj2ya5PQWusBq 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="344" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="92ad" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">while consistently maintaining stable point-read latencies:</p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi pj"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*23F_CzqsjMoI8GHB%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*23F_CzqsjMoI8GHB%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*23F_CzqsjMoI8GHB%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*23F_CzqsjMoI8GHB%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*23F_CzqsjMoI8GHB%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*23F_CzqsjMoI8GHB%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*23F_CzqsjMoI8GHB%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*23F_CzqsjMoI8GHB 640w, https://miro.medium.com/v2/resize:fit:720/0*23F_CzqsjMoI8GHB 720w, https://miro.medium.com/v2/resize:fit:750/0*23F_CzqsjMoI8GHB 750w, https://miro.medium.com/v2/resize:fit:786/0*23F_CzqsjMoI8GHB 786w, https://miro.medium.com/v2/resize:fit:828/0*23F_CzqsjMoI8GHB 828w, https://miro.medium.com/v2/resize:fit:1100/0*23F_CzqsjMoI8GHB 1100w, https://miro.medium.com/v2/resize:fit:1400/0*23F_CzqsjMoI8GHB 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="343" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="4890" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">At the time of writing this blog, the service was processing close to <em class="oy">15 million events/second</em> across all the different datasets at peak globally.</p><figure class="pk pl pm pn po pp ph pi paragraph-image"><div role="button" tabindex="0" class="pq pr fj ps bh pt"><div class="ph pi qv"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*dZFDUVX35Cj1MPOj%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*dZFDUVX35Cj1MPOj%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*dZFDUVX35Cj1MPOj%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*dZFDUVX35Cj1MPOj%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*dZFDUVX35Cj1MPOj%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*dZFDUVX35Cj1MPOj%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*dZFDUVX35Cj1MPOj%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*dZFDUVX35Cj1MPOj 640w, https://miro.medium.com/v2/resize:fit:720/0*dZFDUVX35Cj1MPOj 720w, https://miro.medium.com/v2/resize:fit:750/0*dZFDUVX35Cj1MPOj 750w, https://miro.medium.com/v2/resize:fit:786/0*dZFDUVX35Cj1MPOj 786w, https://miro.medium.com/v2/resize:fit:828/0*dZFDUVX35Cj1MPOj 828w, https://miro.medium.com/v2/resize:fit:1100/0*dZFDUVX35Cj1MPOj 1100w, https://miro.medium.com/v2/resize:fit:1400/0*dZFDUVX35Cj1MPOj 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pu c" width="700" height="366" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><h1 id="8793" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Time Series Usage @ Netflix</h1><p id="3d90" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">The TimeSeries Abstraction plays a vital role across key services at Netflix. Here are some impactful use cases:</p><ul class=""><li id="3163" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt oz pa pb bk"><strong class="my gv">Tracing and Insights: </strong>Logs traces across all apps and micro-services within Netflix, to understand service-to-service communication, aid in debugging of issues, and answer support requests.</li><li id="a327" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk"><strong class="my gv">User Interaction Tracking</strong>: Tracks millions of user interactions — such as video playbacks, searches, and content engagement — providing insights that enhance Netflix’s recommendation algorithms in real-time and improve the overall user experience.</li><li id="7ea2" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk"><strong class="my gv">Feature Rollout and Performance Analysis</strong>: Tracks the rollout and performance of new product features, enabling Netflix engineers to measure how users engage with features, which powers data-driven decisions about future improvements.</li><li id="660c" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk"><strong class="my gv">Asset Impression Tracking and Optimization</strong>: Tracks asset impressions ensuring content and assets are delivered efficiently while providing real-time feedback for optimizations.</li><li id="03cc" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk"><strong class="my gv">Billing and Subscription Management:</strong> Stores historical data related to billing and subscription management, ensuring accuracy in transaction records and supporting customer service inquiries.</li></ul><p id="a1ea" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">and more…</p><h1 id="3624" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Future Enhancements</h1><p id="451d" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">As the use cases evolve, and the need to make the abstraction even more cost effective grows, we aim to make many improvements to the service in the upcoming months. Some of them are:</p><ul class=""><li id="6046" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt oz pa pb bk"><strong class="my gv">Tiered Storage for Cost Efficiency: </strong>Support moving older, lesser-accessed data into cheaper object storage that has higher time to first byte, potentially saving Netflix millions of dollars.</li><li id="4073" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk"><strong class="my gv">Dynamic Event Bucketing: </strong>Support real-time partitioning of keys into optimally-sized partitions as events stream in, rather than having a <em class="oy">somewhat</em> static configuration at the time of provisioning a namespace. This strategy has a huge advantage of <em class="oy">not</em> partitioning time_series_ids that don’t need it, thus saving the overall cost of read amplification. Also, with Cassandra 4.x, we have noted major improvements in reading a subset of data in a wide partition that could lead us to be less aggressive with partitioning the entire dataset ahead of time.</li><li id="07ae" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk"><strong class="my gv">Caching: </strong>Take advantage of immutability of data and cache it intelligently for discrete time ranges.</li><li id="78ff" class="mw mx gu my b mz pc nb nc nd pd nf ng nh pe nj nk nl pf nn no np pg nr ns nt oz pa pb bk"><strong class="my gv">Count and other Aggregations: </strong>Some users are only interested in counting events in a given time interval rather than fetching all the event data for it.</li></ul><h1 id="d63b" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Conclusion</h1><p id="5ab6" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">The TimeSeries Abstraction is a vital component of Netflix’s online data infrastructure, playing a crucial role in supporting both real-time and long-term decision-making. Whether it’s monitoring system performance during high-traffic events or optimizing user engagement through behavior analytics, TimeSeries Abstraction ensures that Netflix operates seamlessly and efficiently on a global scale.</p><p id="5764" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">As Netflix continues to innovate and expand into new verticals, the TimeSeries Abstraction will remain a cornerstone of our platform, helping us push the boundaries of what’s possible in streaming and beyond.</p><p id="c076" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Stay tuned for Part 2, where we’ll introduce our <strong class="my gv">Distributed Counter Abstraction</strong>, a key element of <strong class="my gv">Netflix’s Composite Abstractions</strong>, built on top of the TimeSeries Abstraction.</p><h1 id="d6d2" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Acknowledgments</h1><p id="06af" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">Special thanks to our stunning colleagues who contributed to TimeSeries Abstraction’s success: <a class="af nu" href="https://www.linkedin.com/in/tomdevoe/" rel="noopener ugc nofollow" target="_blank">Tom DeVoe</a> <a class="af nu" href="https://www.linkedin.com/in/mengqingwang/" rel="noopener ugc nofollow" target="_blank">Mengqing Wang</a>, <a class="af nu" href="https://www.linkedin.com/in/kartik894/" rel="noopener ugc nofollow" target="_blank">Kartik Sathyanarayanan</a>, <a class="af nu" href="https://www.linkedin.com/in/jordan-west-8aa1731a3/" rel="noopener ugc nofollow" target="_blank">Jordan West</a>, <a class="af nu" href="https://www.linkedin.com/in/matt-lehman-39549719b/" rel="noopener ugc nofollow" target="_blank">Matt Lehman</a>, <a class="af nu" href="https://www.linkedin.com/in/cheng-wang-10323417/" rel="noopener ugc nofollow" target="_blank">Cheng Wang</a>, <a class="af nu" href="https://www.linkedin.com/in/clohfink/" rel="noopener ugc nofollow" target="_blank">Chris Lohfink</a> .</p></div>]]></description>
      <link>https://netflixtechblog.com/introducing-netflix-timeseries-data-abstraction-layer-31552f6326f8</link>
      <guid>https://netflixtechblog.com/introducing-netflix-timeseries-data-abstraction-layer-31552f6326f8</guid>
      <pubDate>Tue, 08 Oct 2024 19:05:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Introducing Netflix’s Key-Value Data Abstraction Layer]]></title>
      <description><![CDATA[<div><div></div><p id="3a78" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><a class="af nu" href="https://www.linkedin.com/in/vidhya-arvind-11908723" rel="noopener ugc nofollow" target="_blank">Vidhya Arvind</a>, <a class="af nu" href="https://www.linkedin.com/in/rummadis/" rel="noopener ugc nofollow" target="_blank">Rajasekhar Ummadisetty</a>, <a class="af nu" href="https://jolynch.github.io/" rel="noopener ugc nofollow" target="_blank">Joey Lynch</a>, <a class="af nu" href="https://www.linkedin.com/in/vinaychella" rel="noopener ugc nofollow" target="_blank">Vinay Chella</a></p><h1 id="8c5d" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Introduction</h1><p id="8fe5" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">At Netflix our ability to deliver seamless, high-quality, streaming experiences to millions of users hinges on robust, <em class="oy">global</em> backend infrastructure. Central to this infrastructure is our use of multiple online distributed databases such as <a class="af nu" href="https://cassandra.apache.org/" rel="noopener ugc nofollow" target="_blank">Apache Cassandra</a>, a NoSQL database known for its high availability and scalability. Cassandra serves as the backbone for a diverse array of use cases within Netflix, ranging from user sign-ups and storing viewing histories to supporting real-time analytics and live streaming.</p><p id="255e" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Over time as new key-value databases were introduced and service owners launched new use cases, we encountered numerous challenges with datastore misuse. Firstly, developers struggled to reason about consistency, durability and performance in this complex global deployment across multiple stores. Second, developers had to constantly re-learn new data modeling practices and common yet critical data access patterns. These include challenges with tail latency and idempotency, managing “wide” partitions with many rows, handling single large “fat” columns, and slow response pagination. Additionally, the tight coupling with multiple native database APIs — APIs that continually evolve and sometimes introduce backward-incompatible changes — resulted in org-wide engineering efforts to maintain and optimize our microservice’s data access.</p><p id="e872" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">To overcome these challenges, we developed a holistic approach that builds upon our <a class="af nu" href="https://netflixtechblog.medium.com/data-gateway-a-platform-for-growing-and-protecting-the-data-tier-f1ed8db8f5c6" rel="noopener">Data Gateway Platform</a>. This approach led to the creation of several foundational abstraction services, the most mature of which is our Key-Value (KV) Data Abstraction Layer (DAL). This abstraction simplifies data access, enhances the reliability of our infrastructure, and enables us to support the broad spectrum of use cases that Netflix demands with minimal developer effort.</p><p id="92d2" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">In this post, we dive deep into how Netflix’s KV abstraction works, the architectural principles guiding its design, the challenges we faced in scaling diverse use cases, and the technical innovations that have allowed us to achieve the performance and reliability required by Netflix’s global operations.</p><h1 id="9a98" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk"><strong class="al">The Key-Value Service</strong></h1><p id="4bde" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">The KV data abstraction service was introduced to solve the persistent challenges we faced with data access patterns in our distributed databases. Our goal was to build a versatile and efficient data storage solution that could handle a wide variety of use cases, ranging from the simplest hashmaps to more complex data structures, all while ensuring high availability, tunable consistency, and low latency.</p><h2 id="653a" class="oz nw gu bf nx pa pb dy ob pc pd ea of nh pe pf pg nl ph pi pj np pk pl pm pn bk">Data Model</h2><p id="adb9" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">At its core, the KV abstraction is built around a <strong class="my gv"><em class="oy">two-level map</em> </strong>architecture. The first level is a hashed string <strong class="my gv">ID</strong> (the primary key), and the second level is a <strong class="my gv"><em class="oy">sorted map of a key-value pair of bytes</em></strong>. This model supports both simple and complex data models, balancing flexibility and efficiency.</p><pre class="po pp pq pr ps pt pu pv bp pw bb bk">HashMap&lt;String, SortedMap&lt;Bytes, Bytes&gt;&gt;</pre><p id="9c5f" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">For complex data models such as structured <code class="cx qc qd qe pu b">Records</code> or time-ordered <code class="cx qc qd qe pu b">Events</code>, this two-level approach handles hierarchical structures effectively, allowing related data to be retrieved together. For simpler use cases, it also represents flat key-value <code class="cx qc qd qe pu b">Maps</code> (e.g. <code class="cx qc qd qe pu b">id → {"" → value}</code>) or named <code class="cx qc qd qe pu b">Sets</code> (e.g.<code class="cx qc qd qe pu b">id → {key → ""}</code>). This adaptability allows the KV abstraction to be used in hundreds of diverse use cases, making it a versatile solution for managing both simple and complex data models in large-scale infrastructures like Netflix.</p><p id="e359" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">The KV data can be visualized at a high level, as shown in the diagram below, where three records are shown.</p><figure class="po pp pq pr ps qi qf qg paragraph-image"><div role="button" tabindex="0" class="qj qk fj ql bh qm"><div class="qf qg qh"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*9Ny8Uc-diSDnVGnk%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*9Ny8Uc-diSDnVGnk%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*9Ny8Uc-diSDnVGnk%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*9Ny8Uc-diSDnVGnk%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*9Ny8Uc-diSDnVGnk%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*9Ny8Uc-diSDnVGnk%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*9Ny8Uc-diSDnVGnk%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*9Ny8Uc-diSDnVGnk 640w, https://miro.medium.com/v2/resize:fit:720/0*9Ny8Uc-diSDnVGnk 720w, https://miro.medium.com/v2/resize:fit:750/0*9Ny8Uc-diSDnVGnk 750w, https://miro.medium.com/v2/resize:fit:786/0*9Ny8Uc-diSDnVGnk 786w, https://miro.medium.com/v2/resize:fit:828/0*9Ny8Uc-diSDnVGnk 828w, https://miro.medium.com/v2/resize:fit:1100/0*9Ny8Uc-diSDnVGnk 1100w, https://miro.medium.com/v2/resize:fit:1400/0*9Ny8Uc-diSDnVGnk 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md qn c" width="700" height="442" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><pre class="po pp pq pr ps pt pu pv bp pw bb bk">message Item (   <br />  Bytes    key,<br />  Bytes    value,<br />  Metadata metadata,<br />  Integer  chunk<br />)</pre><h2 id="9585" class="oz nw gu bf nx pa pb dy ob pc pd ea of nh pe pf pg nl ph pi pj np pk pl pm pn bk">Database Agnostic Abstraction</h2><p id="7654" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">The KV abstraction is designed to hide the implementation details of the underlying database, offering a consistent interface to application developers regardless of the optimal storage system for that use case. While Cassandra is one example, the abstraction works with multiple data stores like <a class="af nu" href="https://github.com/Netflix/EVCache" rel="noopener ugc nofollow" target="_blank">EVCache</a>, <a class="af nu" href="https://aws.amazon.com/dynamodb/" rel="noopener ugc nofollow" target="_blank">DynamoDB</a>, <a class="af nu" href="https://rocksdb.org/" rel="noopener ugc nofollow" target="_blank">RocksDB</a>, etc…</p><p id="6b60" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">For example, when implemented with Cassandra, the abstraction leverages Cassandra’s partitioning and clustering capabilities. The record <strong class="my gv"><em class="oy">ID</em></strong> acts as the partition key, and the item <strong class="my gv"><em class="oy">key</em></strong> as the clustering column:</p><figure class="po pp pq pr ps qi qf qg paragraph-image"><div role="button" tabindex="0" class="qj qk fj ql bh qm"><div class="qf qg qo"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*tMhXVTWqtHt24l1oflpAJQ.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*tMhXVTWqtHt24l1oflpAJQ.png" /><img alt="" class="bh md qn c" width="700" height="194" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="17d6" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">The corresponding Data Definition Language (DDL) for this structure in Cassandra is:</p><pre class="po pp pq pr ps pt pu pv bp pw bb bk">CREATE TABLE IF NOT EXISTS &lt;ns&gt;.&lt;table&gt; (<br />  id             text,<br />  key            blob,<br />  value          blob,<br />  value_metadata blob,PRIMARY KEY (id, key))<br />WITH CLUSTERING ORDER BY (key &lt;ASC|DESC&gt;)</pre><h2 id="dd6f" class="oz nw gu bf nx pa pb dy ob pc pd ea of nh pe pf pg nl ph pi pj np pk pl pm pn bk">Namespace: Logical and Physical Configuration</h2><p id="6e76" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">A <strong class="my gv">namespace</strong> defines where and how data is stored, providing logical and physical separation while abstracting the underlying storage systems. It also serves as central configuration of access patterns such as consistency or latency targets. Each namespace may use different backends: Cassandra, EVCache, or combinations of multiple. This flexibility allows our Data Platform to route different use cases to the most suitable storage system based on performance, durability, and consistency needs. Developers just provide their data problem rather than a database solution!</p><p id="898b" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">In this example configuration, the <code class="cx qc qd qe pu b">ngsegment</code> namespace is backed by both a Cassandra cluster and an EVCache caching layer, allowing for highly durable persistent storage <em class="oy">and</em> lower-latency point reads.</p><pre class="po pp pq pr ps pt pu pv bp pw bb bk">"persistence_configuration":[                                                   <br />  {                                                                           <br />    "id":"PRIMARY_STORAGE",                                                 <br />    "physical_storage": {                                                    <br />      "type":"CASSANDRA",                                                 <br />      "cluster":"cassandra_kv_ngsegment",                                <br />      "dataset":"ngsegment",                                             <br />      "table":"ngsegment",                                               <br />      "regions": ["us-east-1"],<br />      "config": {<br />        "consistency_scope": "LOCAL",<br />        "consistency_target": "READ_YOUR_WRITES"<br />      }                                            <br />    }                                                                       <br />  },                                                                          <br />  {                                                                           <br />    "id":"CACHE",                                                           <br />    "physical_storage": {                                                    <br />      "type":"CACHE",                                                     <br />      "cluster":"evcache_kv_ngsegment"                                   <br />     },                                                                      <br />     "config": {                                                              <br />       "default_cache_ttl": 180s                                             <br />     }                                                                       <br />  }                                                                           <br />] <br /></pre><h1 id="d313" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk"><strong class="al">Key APIs of the KV Abstraction</strong></h1><p id="0ffe" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">To support diverse use-cases, the KV abstraction provides four basic CRUD APIs:</p><h2 id="b9f2" class="oz nw gu bf nx pa pb dy ob pc pd ea of nh pe pf pg nl ph pi pj np pk pl pm pn bk">PutItems <strong class="al">— Write one or more Items to a Record</strong></h2><p id="4d14" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">The <code class="cx qc qd qe pu b">PutItems</code> API is an upsert operation, it can insert new data or update existing data in the two-level map structure.</p><pre class="po pp pq pr ps pt pu pv bp pw bb bk">message PutItemRequest (<br />  IdempotencyToken idempotency_token,<br />  string           namespace, <br />  string           id, <br />  List&lt;Item&gt;       items<br />)</pre><p id="d3db" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">As you can see, the request includes the namespace, Record ID, one or more items, and an <strong class="my gv">idempotency token</strong> to ensure retries of the same write are safe. Chunked data can be written by staging chunks and then committing them with appropriate metadata (e.g. number of chunks).</p><h2 id="d04a" class="oz nw gu bf nx pa pb dy ob pc pd ea of nh pe pf pg nl ph pi pj np pk pl pm pn bk">GetItems <strong class="al">— Read one or more Items from a Record</strong></h2><p id="487c" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">The <code class="cx qc qd qe pu b">GetItems</code>API provides a structured and adaptive way to fetch data using ID, predicates, and selection mechanisms. This approach balances the need to retrieve large volumes of data while meeting stringent Service Level Objectives (SLOs) for performance and reliability.</p><pre class="po pp pq pr ps pt pu pv bp pw bb bk">message GetItemsRequest (<br />  String              namespace,<br />  String              id,<br />  Predicate           predicate,<br />  Selection           selection,<br />  Map&lt;String, Struct&gt; signals<br />)</pre><p id="ed18" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">The <code class="cx qc qd qe pu b">GetItemsRequest</code> includes several key parameters:</p><ul class=""><li id="f54d" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt qp qq qr bk"><strong class="my gv">Namespace</strong>: Specifies the logical dataset or table</li><li id="a823" class="mw mx gu my b mz qs nb nc nd qt nf ng nh qu nj nk nl qv nn no np qw nr ns nt qp qq qr bk"><strong class="my gv">Id</strong>: Identifies the entry in the top-level HashMap</li><li id="b7f8" class="mw mx gu my b mz qs nb nc nd qt nf ng nh qu nj nk nl qv nn no np qw nr ns nt qp qq qr bk"><strong class="my gv">Predicate</strong>: Filters the matching items and can retrieve all items (<code class="cx qc qd qe pu b">match_all</code>), specific items (<code class="cx qc qd qe pu b">match_keys</code>), or a range (<code class="cx qc qd qe pu b">match_range</code>)</li><li id="a2e5" class="mw mx gu my b mz qs nb nc nd qt nf ng nh qu nj nk nl qv nn no np qw nr ns nt qp qq qr bk"><strong class="my gv">Selection</strong>: Narrows returned responses for example <code class="cx qc qd qe pu b">page_size_bytes</code> for pagination, <code class="cx qc qd qe pu b">item_limit</code> for limiting the total number of items across pages and <code class="cx qc qd qe pu b">include</code>/<code class="cx qc qd qe pu b">exclude</code> to include or exclude large values from responses</li><li id="9b65" class="mw mx gu my b mz qs nb nc nd qt nf ng nh qu nj nk nl qv nn no np qw nr ns nt qp qq qr bk"><strong class="my gv">Signals:</strong> Provides in-band signaling to indicate client capabilities, such as supporting client compression or chunking.</li></ul><p id="697f" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">The <code class="cx qc qd qe pu b">GetItemResponse</code> message contains the matching data:</p><pre class="po pp pq pr ps pt pu pv bp pw bb bk">message GetItemResponse (<br />  List&lt;Item&gt;       items,<br />  Optional&lt;String&gt; next_page_token<br />)</pre><ul class=""><li id="9d6e" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt qp qq qr bk"><strong class="my gv">Items</strong>: A list of retrieved items based on the <code class="cx qc qd qe pu b">Predicate</code> and <code class="cx qc qd qe pu b">Selection</code> defined in the request.</li><li id="5617" class="mw mx gu my b mz qs nb nc nd qt nf ng nh qu nj nk nl qv nn no np qw nr ns nt qp qq qr bk"><strong class="my gv">Next Page Token</strong>: An optional token indicating the position for subsequent reads if needed, essential for handling large data sets across multiple requests. Pagination is a critical component for efficiently managing data retrieval, especially when dealing with large datasets that could exceed typical response size limits.</li></ul><h2 id="a3d0" class="oz nw gu bf nx pa pb dy ob pc pd ea of nh pe pf pg nl ph pi pj np pk pl pm pn bk"><strong class="al">DeleteItems — Delete one or more Items from a Record</strong></h2><p id="b258" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">The <code class="cx qc qd qe pu b">DeleteItems</code> API provides flexible options for removing data, including record-level, item-level, and range deletes — all while supporting idempotency.</p><pre class="po pp pq pr ps pt pu pv bp pw bb bk">message DeleteItemsRequest (<br />  IdempotencyToken idempotency_token,<br />  String           namespace,<br />  String           id,<br />  Predicate        predicate<br />)<br /></pre><p id="b5bd" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Just like in the <code class="cx qc qd qe pu b">GetItems</code> API, the <code class="cx qc qd qe pu b">Predicate</code> allows one or more Items to be addressed at once:</p><ul class=""><li id="4e42" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt qp qq qr bk"><strong class="my gv">Record-Level Deletes (match_all)</strong>: Removes the entire record in constant latency regardless of the number of items in the record.</li><li id="0b32" class="mw mx gu my b mz qs nb nc nd qt nf ng nh qu nj nk nl qv nn no np qw nr ns nt qp qq qr bk"><strong class="my gv">Item-Range Deletes (match_range)</strong>: This deletes a range of items within a Record. Useful for keeping “n-newest” or prefix path deletion.</li><li id="7525" class="mw mx gu my b mz qs nb nc nd qt nf ng nh qu nj nk nl qv nn no np qw nr ns nt qp qq qr bk"><strong class="my gv">Item-Level Deletes (match_keys)</strong>: Deletes one or more individual items.</li></ul><p id="4f76" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Some storage engines (any store which defers true deletion) such as Cassandra struggle with high volumes of deletes due to tombstone and compaction overhead. Key-Value optimizes both record and range deletes to generate a single tombstone for the operation — you can learn more about tombstones in <a class="af nu" href="https://thelastpickle.com/blog/2016/07/27/about-deletes-and-tombstones.html" rel="noopener ugc nofollow" target="_blank">About Deletes and Tombstones</a>.</p><p id="3569" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Item-level deletes create many tombstones but KV hides that storage engine complexity via <strong class="my gv">TTL-based deletes with jitter</strong>. Instead of immediate deletion, item metadata is updated as expired with randomly jittered TTL applied to stagger deletions. This technique maintains read pagination protections. While this doesn’t completely solve the problem it reduces load spikes and helps maintain consistent performance while compaction catches up. These strategies help maintain system performance, reduce read overhead, and meet SLOs by minimizing the impact of deletes.</p><h2 id="9f4f" class="oz nw gu bf nx pa pb dy ob pc pd ea of nh pe pf pg nl ph pi pj np pk pl pm pn bk">Complex Mutate and Scan APIs</h2><p id="d6e0" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">Beyond simple CRUD on single Records, KV also supports complex multi-item and multi-record mutations and scans via <code class="cx qc qd qe pu b">MutateItems</code> and <code class="cx qc qd qe pu b">ScanItems</code> APIs. <code class="cx qc qd qe pu b">PutItems</code> also supports atomic writes of large blob data within a single <code class="cx qc qd qe pu b">Item</code> via a chunked protocol. These complex APIs require careful consideration to ensure predictable linear low-latency and we will share details on their implementation in a future post.</p><h1 id="4e6f" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Design Philosophies for reliable and predictable performance</h1><h2 id="c223" class="oz nw gu bf nx pa pb dy ob pc pd ea of nh pe pf pg nl ph pi pj np pk pl pm pn bk">Idempotency to fight tail latencies</h2><p id="d7cd" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">To ensure data integrity the <code class="cx qc qd qe pu b">PutItems</code> and <code class="cx qc qd qe pu b">DeleteItems</code> APIs use <strong class="my gv">idempotency tokens</strong>, which uniquely identify each mutative operation and guarantee that operations are logically executed in order, even when hedged or retried for latency reasons. This is especially crucial in last-write-wins databases like Cassandra, where ensuring the correct order and de-duplication of requests is vital.</p><p id="b8d5" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">In the Key-Value abstraction, idempotency tokens contain a generation timestamp and random nonce token. Either or both may be required by backing storage engines to de-duplicate mutations.</p><pre class="po pp pq pr ps pt pu pv bp pw bb bk">message IdempotencyToken (<br />  Timestamp generation_time,<br />  String    token<br />)</pre><p id="3089" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">At Netflix, <strong class="my gv">client-generated monotonic tokens</strong> are preferred due to their reliability, especially in environments where network delays could impact server-side token generation. This combines a client provided monotonic <code class="cx qc qd qe pu b">generation_time</code> timestamp with a 128 bit random UUID <code class="cx qc qd qe pu b">token</code>. Although clock-based token generation can suffer from clock skew, our tests on EC2 Nitro instances show drift is minimal (under 1 millisecond). In some cases that require stronger ordering, regionally unique tokens can be generated using tools like Zookeeper, or globally unique tokens such as a transaction IDs can be used.</p><p id="abaa" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">The following graphs illustrate the observed <a class="af nu" href="https://docs.google.com/document/d/1XLBjQ9scZCy-xIo51Rs--CSdFV781fnp5hXdXTBAk1k/edit" rel="noopener ugc nofollow" target="_blank">clock skew</a> on our Cassandra fleet, suggesting the safety of this technique on modern cloud VMs with direct access to high-quality clocks. To further maintain safety, KV servers reject writes bearing tokens with large drift both preventing silent write discard (write has timestamp far in past) and immutable doomstones (write has a timestamp far in future) in storage engines vulnerable to those.</p><figure class="po pp pq pr ps qi qf qg paragraph-image"><div role="button" tabindex="0" class="qj qk fj ql bh qm"><div class="qf qg qx"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*gTmQpPIyZcKDb4Fb%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*gTmQpPIyZcKDb4Fb%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*gTmQpPIyZcKDb4Fb%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*gTmQpPIyZcKDb4Fb%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*gTmQpPIyZcKDb4Fb%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*gTmQpPIyZcKDb4Fb%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*gTmQpPIyZcKDb4Fb%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*gTmQpPIyZcKDb4Fb 640w, https://miro.medium.com/v2/resize:fit:720/0*gTmQpPIyZcKDb4Fb 720w, https://miro.medium.com/v2/resize:fit:750/0*gTmQpPIyZcKDb4Fb 750w, https://miro.medium.com/v2/resize:fit:786/0*gTmQpPIyZcKDb4Fb 786w, https://miro.medium.com/v2/resize:fit:828/0*gTmQpPIyZcKDb4Fb 828w, https://miro.medium.com/v2/resize:fit:1100/0*gTmQpPIyZcKDb4Fb 1100w, https://miro.medium.com/v2/resize:fit:1400/0*gTmQpPIyZcKDb4Fb 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md qn c" width="700" height="685" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><h2 id="7b42" class="oz nw gu bf nx pa pb dy ob pc pd ea of nh pe pf pg nl ph pi pj np pk pl pm pn bk">Handling Large Data through Chunking</h2><p id="4ff0" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">Key-Value is also designed to efficiently handle large blobs, a common challenge for traditional key-value stores. Databases often face limitations on the amount of data that can be stored per key or partition. To address these constraints, KV uses transparent <strong class="my gv">chunking</strong> to manage large data efficiently.</p><p id="2392" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">For items smaller than 1 MiB, data is stored directly in the main backing storage (e.g. Cassandra), ensuring fast and efficient access. However, for larger items, only the <strong class="my gv">id</strong>, <strong class="my gv">key</strong>, and <strong class="my gv">metadata</strong> are stored in the primary storage, while the actual data is split into smaller chunks and stored separately in chunk storage. This chunk storage can also be Cassandra but with a different partitioning scheme optimized for handling large values. The idempotency token ties all these writes together into one atomic operation.</p><p id="605e" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">By splitting large items into chunks, we ensure that latency scales linearly with the size of the data, making the system both predictable and efficient. A future blog post will describe the <strong class="my gv">chunking architecture</strong> in more detail, including its intricacies and optimization strategies.</p><h2 id="2b4f" class="oz nw gu bf nx pa pb dy ob pc pd ea of nh pe pf pg nl ph pi pj np pk pl pm pn bk">Client-Side Compression</h2><p id="5ad2" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">The KV abstraction leverages client-side payload compression to optimize performance, especially for large data transfers. While many databases offer server-side compression, handling compression on the client side reduces expensive server CPU usage, network bandwidth, and disk I/O. In one of our deployments, which helps power Netflix’s search, enabling client-side compression reduced payload sizes by 75%, significantly improving cost efficiency.</p><h2 id="5805" class="oz nw gu bf nx pa pb dy ob pc pd ea of nh pe pf pg nl ph pi pj np pk pl pm pn bk">Smarter Pagination</h2><p id="739c" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">We chose payload size in bytes as the limit per response page rather than the number of items because it allows us to provide predictable operation SLOs. For instance, we can provide a single-digit millisecond SLO on a 2 MiB page read. Conversely, using the number of items per page as the limit would result in unpredictable latencies due to significant variations in item size. A request for 10 items per page could result in vastly different latencies if each item was 1 KiB versus 1 MiB.</p><p id="21ce" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Using bytes as a limit poses challenges as few backing stores support byte-based pagination; most data stores use the number of results e.g. DynamoDB and Cassandra limit by number of items or rows. To address this, we use a static limit for the initial queries to the backing store, query with this limit, and process the results. If more data is needed to meet the byte limit, additional queries are executed until the limit is met, the excess result is discarded and a page token is generated.</p><p id="51d1" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">This static limit can lead to inefficiencies, one large item in the result may cause us to discard many results, while small items may require multiple iterations to fill a page, resulting in read amplification. To mitigate these issues, we implemented <em class="oy">adaptive</em> pagination which dynamically tunes the limits based on observed data.</p><h2 id="6f49" class="oz nw gu bf nx pa pb dy ob pc pd ea of nh pe pf pg nl ph pi pj np pk pl pm pn bk">Adaptive Pagination</h2><p id="d9f6" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">When an initial request is made, a query is executed in the storage engine, and the results are retrieved. As the consumer processes these results, the system tracks the number of items consumed and the total size used. This data helps calculate an approximate item size, which is stored in the page token. For subsequent page requests, this stored information allows the server to apply the appropriate limits to the underlying storage, reducing unnecessary work and minimizing read amplification.</p><p id="c57f" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">While this method is effective for follow-up page requests, what happens with the initial request? In addition to storing item size information in the page token, the server also estimates the average item size for a given namespace and caches it locally. This cached estimate helps the server set a more optimal limit on the backing store for the initial request, improving efficiency. The server continuously adjusts this limit based on recent query patterns or other factors to keep it accurate. For subsequent pages, the server uses both the cached data and the information in the page token to fine-tune the limits.</p><figure class="po pp pq pr ps qi qf qg paragraph-image"><div role="button" tabindex="0" class="qj qk fj ql bh qm"><div class="qf qg qy"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*yg8xyQEoEmvKYoOV%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*yg8xyQEoEmvKYoOV%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*yg8xyQEoEmvKYoOV%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*yg8xyQEoEmvKYoOV%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*yg8xyQEoEmvKYoOV%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*yg8xyQEoEmvKYoOV%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*yg8xyQEoEmvKYoOV%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*yg8xyQEoEmvKYoOV 640w, https://miro.medium.com/v2/resize:fit:720/0*yg8xyQEoEmvKYoOV 720w, https://miro.medium.com/v2/resize:fit:750/0*yg8xyQEoEmvKYoOV 750w, https://miro.medium.com/v2/resize:fit:786/0*yg8xyQEoEmvKYoOV 786w, https://miro.medium.com/v2/resize:fit:828/0*yg8xyQEoEmvKYoOV 828w, https://miro.medium.com/v2/resize:fit:1100/0*yg8xyQEoEmvKYoOV 1100w, https://miro.medium.com/v2/resize:fit:1400/0*yg8xyQEoEmvKYoOV 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md qn c" width="700" height="559" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="e11f" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">In addition to adaptive pagination, a mechanism is in place to send a response early if the server detects that processing the request is at risk of exceeding the request’s latency SLO.</p><p id="8c43" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">For example, let us assume a client submits a <code class="cx qc qd qe pu b">GetItems</code> request with a per-page limit of 2 MiB and a maximum end-to-end latency limit of 500ms. While processing this request, the server retrieves data from the backing store. This particular record has thousands of small items so it would normally take longer than the 500ms SLO to gather the full page of data. If this happens, the client would receive an SLO violation error, causing the request to fail even though there is nothing exceptional. To prevent this, the server tracks the elapsed time while fetching data. If it determines that continuing to retrieve more data might breach the SLO, the server will stop processing further results and return a response with a pagination token.</p><figure class="po pp pq pr ps qi qf qg paragraph-image"><div role="button" tabindex="0" class="qj qk fj ql bh qm"><div class="qf qg qz"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*hEkIfkUJ4KDnbbGx%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*hEkIfkUJ4KDnbbGx%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*hEkIfkUJ4KDnbbGx%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*hEkIfkUJ4KDnbbGx%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*hEkIfkUJ4KDnbbGx%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*hEkIfkUJ4KDnbbGx%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*hEkIfkUJ4KDnbbGx%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*hEkIfkUJ4KDnbbGx 640w, https://miro.medium.com/v2/resize:fit:720/0*hEkIfkUJ4KDnbbGx 720w, https://miro.medium.com/v2/resize:fit:750/0*hEkIfkUJ4KDnbbGx 750w, https://miro.medium.com/v2/resize:fit:786/0*hEkIfkUJ4KDnbbGx 786w, https://miro.medium.com/v2/resize:fit:828/0*hEkIfkUJ4KDnbbGx 828w, https://miro.medium.com/v2/resize:fit:1100/0*hEkIfkUJ4KDnbbGx 1100w, https://miro.medium.com/v2/resize:fit:1400/0*hEkIfkUJ4KDnbbGx 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md qn c" width="700" height="296" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="858e" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">This approach ensures that requests are processed within the SLO, even if the full page size isn’t met, giving clients predictable progress. Furthermore, if the client is a gRPC server with proper deadlines, the client is smart enough not to issue further requests, reducing useless work.</p><p id="c909" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">If you want to know more, the <a class="af nu" href="https://www.infoq.com/articles/netflix-highly-reliable-stateful-systems/" rel="noopener ugc nofollow" target="_blank">How Netflix Ensures Highly-Reliable Online Stateful Systems</a> article talks in further detail about these and many other techniques.</p><h2 id="c0c5" class="oz nw gu bf nx pa pb dy ob pc pd ea of nh pe pf pg nl ph pi pj np pk pl pm pn bk">Signaling</h2><p id="2e5c" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">KV uses in-band messaging we call <em class="oy">signaling</em> that allows the dynamic configuration of the client and enables it to communicate its capabilities to the server. This ensures that configuration settings and tuning parameters can be exchanged seamlessly between the client and server. Without signaling, the client would need static configuration — requiring a redeployment for each change — or, with dynamic configuration, would require coordination with the client team.</p><p id="f164" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">For server-side signals, when the client is initialized, it sends a handshake to the server. The server responds back with signals, such as target or max latency SLOs, allowing the client to dynamically adjust timeouts and hedging policies. Handshakes are then made periodically in the background to keep the configuration current. For client-communicated signals, the client, along with each request, communicates its capabilities, such as whether it can handle compression, chunking, and other features.</p><figure class="po pp pq pr ps qi qf qg paragraph-image"><div role="button" tabindex="0" class="qj qk fj ql bh qm"><div class="qf qg ra"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*sVOLoSeIKpzDMQ5N%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*sVOLoSeIKpzDMQ5N%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*sVOLoSeIKpzDMQ5N%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*sVOLoSeIKpzDMQ5N%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*sVOLoSeIKpzDMQ5N%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*sVOLoSeIKpzDMQ5N%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*sVOLoSeIKpzDMQ5N%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*sVOLoSeIKpzDMQ5N 640w, https://miro.medium.com/v2/resize:fit:720/0*sVOLoSeIKpzDMQ5N 720w, https://miro.medium.com/v2/resize:fit:750/0*sVOLoSeIKpzDMQ5N 750w, https://miro.medium.com/v2/resize:fit:786/0*sVOLoSeIKpzDMQ5N 786w, https://miro.medium.com/v2/resize:fit:828/0*sVOLoSeIKpzDMQ5N 828w, https://miro.medium.com/v2/resize:fit:1100/0*sVOLoSeIKpzDMQ5N 1100w, https://miro.medium.com/v2/resize:fit:1400/0*sVOLoSeIKpzDMQ5N 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md qn c" width="700" height="264" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><h1 id="ea21" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">KV Usage @ Netflix</h1><p id="18e7" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">The KV abstraction powers several key Netflix use cases, including:</p><ul class=""><li id="f38f" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt qp qq qr bk"><strong class="my gv">Streaming Metadata</strong>: High-throughput, low-latency access to streaming metadata, ensuring personalized content delivery in real-time.</li><li id="555b" class="mw mx gu my b mz qs nb nc nd qt nf ng nh qu nj nk nl qv nn no np qw nr ns nt qp qq qr bk"><strong class="my gv">User Profiles</strong>: Efficient storage and retrieval of user preferences and history, enabling seamless, personalized experiences across devices.</li><li id="61d8" class="mw mx gu my b mz qs nb nc nd qt nf ng nh qu nj nk nl qv nn no np qw nr ns nt qp qq qr bk"><strong class="my gv">Messaging</strong>: Storage and retrieval of <a class="af nu" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/pushy-to-the-limit-evolving-netflixs-websocket-proxy-for-the-future-b468bc0ff658">push registry</a> for messaging needs, enabling the millions of requests to flow through.</li><li id="bcc9" class="mw mx gu my b mz qs nb nc nd qt nf ng nh qu nj nk nl qv nn no np qw nr ns nt qp qq qr bk"><strong class="my gv">Real-Time Analytics</strong>: This persists large-scale impression and provides insights into user behavior and system performance, <a class="af nu" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/bulldozer-batch-data-moving-from-data-warehouse-to-online-key-value-stores-41bac13863f8">moving data from offline to online</a> and vice versa.</li></ul><h1 id="332b" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Future Enhancements</h1><p id="6a18" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">Looking forward, we plan to enhance the KV abstraction with:</p><ul class=""><li id="27c2" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt qp qq qr bk"><strong class="my gv">Lifecycle Management</strong>: Fine-grained control over data retention and deletion.</li><li id="3282" class="mw mx gu my b mz qs nb nc nd qt nf ng nh qu nj nk nl qv nn no np qw nr ns nt qp qq qr bk"><strong class="my gv">Summarization</strong>: Techniques to improve retrieval efficiency by summarizing records with many items into fewer backing rows.</li><li id="249b" class="mw mx gu my b mz qs nb nc nd qt nf ng nh qu nj nk nl qv nn no np qw nr ns nt qp qq qr bk"><strong class="my gv">New Storage Engines</strong>: Integration with more storage systems to support new use cases.</li><li id="1ebe" class="mw mx gu my b mz qs nb nc nd qt nf ng nh qu nj nk nl qv nn no np qw nr ns nt qp qq qr bk"><strong class="my gv">Dictionary Compression</strong>: Further reducing data size while maintaining performance.</li></ul><h1 id="496e" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Conclusion</h1><p id="b19f" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">The Key-Value service at Netflix is a flexible, cost-effective solution that supports a wide range of data patterns and use cases, from low to high traffic scenarios, including critical Netflix streaming use-cases. The simple yet robust design allows it to handle diverse data models like HashMaps, Sets, Event storage, Lists, and Graphs. It abstracts the complexity of the underlying databases from our developers, which enables our application engineers to focus on solving business problems instead of becoming experts in every storage engine and their distributed <a class="af nu" href="https://jepsen.io/consistency" rel="noopener ugc nofollow" target="_blank">consistency models</a>. As Netflix continues to innovate in online datastores, the KV abstraction remains a central component in managing data efficiently and reliably at scale, ensuring a solid foundation for future growth.</p><p id="683f" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><strong class="my gv"><em class="oy">Acknowledgments:</em></strong><em class="oy"> Special thanks to our stunning colleagues who contributed to Key Value’s success: </em><a class="af nu" href="https://www.linkedin.com/in/william-schor/" rel="noopener ugc nofollow" target="_blank"><em class="oy">William Schor</em></a><em class="oy">, </em><a class="af nu" href="https://www.linkedin.com/in/mengqingwang/" rel="noopener ugc nofollow" target="_blank"><em class="oy">Mengqing Wang</em></a><em class="oy">, </em><a class="af nu" href="https://www.linkedin.com/in/cthumuluru/" rel="noopener ugc nofollow" target="_blank"><em class="oy">Chandrasekhar Thumuluru</em></a><em class="oy">, </em><a class="af nu" href="https://www.linkedin.com/in/john-l-693b7915a/" rel="noopener ugc nofollow" target="_blank"><em class="oy">John Lu</em></a><em class="oy">, </em><a class="af nu" href="https://www.linkedin.com/in/georgecampbell/" rel="noopener ugc nofollow" target="_blank"><em class="oy">George Cambell</em></a><em class="oy">, </em><a class="af nu" href="https://www.linkedin.com/in/akhaku/" rel="noopener ugc nofollow" target="_blank"><em class="oy">Ammar Khaku</em></a><em class="oy">, </em><a class="af nu" href="https://www.linkedin.com/in/jordan-west-8aa1731a3/" rel="noopener ugc nofollow" target="_blank"><em class="oy">Jordan West</em></a><em class="oy">, </em><a class="af nu" href="https://www.linkedin.com/in/clohfink/" rel="noopener ugc nofollow" target="_blank"><em class="oy">Chris Lohfink</em></a><em class="oy">, </em><a class="af nu" href="https://www.linkedin.com/in/matt-lehman-39549719b/" rel="noopener ugc nofollow" target="_blank"><em class="oy">Matt Lehman</em></a><em class="oy">, and the whole online datastores team (ODS, f.k.a CDE).</em></p></div>]]></description>
      <link>https://netflixtechblog.com/introducing-netflixs-key-value-data-abstraction-layer-1ea8a0a11b30</link>
      <guid>https://netflixtechblog.com/introducing-netflixs-key-value-data-abstraction-layer-1ea8a0a11b30</guid>
      <pubDate>Thu, 19 Sep 2024 00:49:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Pushy to the Limit: Evolving Netflix’s WebSocket proxy for the future]]></title>
      <description><![CDATA[<div><div></div><p id="7b9e" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><em class="nu">By </em><a class="af nv" href="https://www.linkedin.com/in/kyagna/" rel="noopener ugc nofollow" target="_blank"><em class="nu">Karthik Yagna</em></a><em class="nu">, </em><a class="af nv" href="https://www.linkedin.com/in/baskar-o-n-46477b3/" rel="noopener ugc nofollow" target="_blank"><em class="nu">Baskar Odayarkoil</em></a><em class="nu">, and </em><a class="af nv" href="https://www.linkedin.com/in/alexander-ellis/" rel="noopener ugc nofollow" target="_blank"><em class="nu">Alex Ellis</em></a></p><p id="e6ee" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Pushy is Netflix’s WebSocket server that maintains persistent WebSocket connections with devices running the Netflix application. This allows data to be sent to the device from backend services on demand, without the need for continually polling requests from the device. Over the last few years, Pushy has seen tremendous growth, evolving from its role as a best-effort message delivery service to be an integral part of the Netflix ecosystem. This post describes how we’ve grown and scaled Pushy to meet its new and future needs, as it handles hundreds of millions of concurrent WebSocket connections, delivers hundreds of thousands of messages per second, and maintains a steady 99.999% message delivery reliability rate.</p><h1 id="c697" class="nw nx gu bf ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os ot bk">History &amp; motivation</h1><p id="2131" class="pw-post-body-paragraph mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt gn bk">There were two main motivating use cases that drove Pushy’s initial development and usage. The first was voice control, where you can play a title or search using your virtual assistant with a voice command like “Show me Stranger Things on Netflix.” (See <a class="af nv" href="https://help.netflix.com/en/node/111997" rel="noopener ugc nofollow" target="_blank"><em class="nu">How to use voice controls with Netflix</em></a> if you want to do this yourself!).</p><p id="3c8c" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">If we consider the Alexa use case, we can see how this partnership with Amazon enabled this to work. Once they receive the voice command, we allow them to make an authenticated call through <a class="af nv" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/open-sourcing-zuul-2-82ea476cb2b3">apiproxy</a>, our streaming edge proxy, to our internal voice service. This call includes metadata, such as the user’s information and details about the command, such as the specific show to play. The voice service then constructs a message for the device and places it on the message queue, which is then processed and sent to Pushy to deliver to the device. Finally, the device receives the message, and the action, such as “Show me Stranger Things on Netflix”, is performed. This initial functionality was built out for FireTVs and was expanded from there.</p><figure class="pc pd pe pf pg ph oz pa paragraph-image"><div role="button" tabindex="0" class="pi pj fj pk bh pl"><div class="oz pa pb"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*WQ1W30ChfWrEmmR5%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*WQ1W30ChfWrEmmR5%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*WQ1W30ChfWrEmmR5%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*WQ1W30ChfWrEmmR5%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*WQ1W30ChfWrEmmR5%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*WQ1W30ChfWrEmmR5%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*WQ1W30ChfWrEmmR5%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*WQ1W30ChfWrEmmR5 640w, https://miro.medium.com/v2/resize:fit:720/0*WQ1W30ChfWrEmmR5 720w, https://miro.medium.com/v2/resize:fit:750/0*WQ1W30ChfWrEmmR5 750w, https://miro.medium.com/v2/resize:fit:786/0*WQ1W30ChfWrEmmR5 786w, https://miro.medium.com/v2/resize:fit:828/0*WQ1W30ChfWrEmmR5 828w, https://miro.medium.com/v2/resize:fit:1100/0*WQ1W30ChfWrEmmR5 1100w, https://miro.medium.com/v2/resize:fit:1400/0*WQ1W30ChfWrEmmR5 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="Sample system diagram for an Alexa voice command, with the voice command entering Netflix’s cloud infrastructure via apiproxy and existing via a server-side message through Pushy to the device." class="bh md pm c" width="700" height="347" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="pn ff po oz pa pp pq bf b bg z du"><em class="pr">Sample system diagram for an Alexa voice command. Where aws ends and the internet begins is an exercise left to the reader.</em></figcaption></figure><p id="022f" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">The other main use case was RENO, the Rapid Event Notification System mentioned above. Before the integration with Pushy, the TV UI would continuously poll a backend service to see if there were any row updates to get the latest information. These requests would happen every few seconds, which ended up creating extraneous requests to the backend and were costly for devices, which are frequently resource constrained. The integration with WebSockets and Pushy alleviated both of these points, allowing the origin service to send row updates as they were ready, resulting in lower request rates and cost savings.</p><p id="63a0" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">For more background on Pushy, you can see <a class="af nv" href="https://www.youtube.com/watch?v=6w6E_B55p0E" rel="noopener ugc nofollow" target="_blank">this InfoQ talk by Susheel Aroskar</a>. Since that presentation, Pushy has grown in both size and scope, and this article will be discussing the investments we’ve made to evolve Pushy for the next generation of features.</p><h1 id="2f72" class="nw nx gu bf ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os ot bk">Client Reach</h1><p id="5d47" class="pw-post-body-paragraph mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt gn bk">This integration was initially rolled out for Fire TVs, PS4s, Samsung TVs, and LG TVs, leading to a reach of about 30 million candidate devices. With these clear benefits, we continued to build out this functionality for more devices, enabling the same efficiency wins. As of today, we’ve expanded our list of candidate devices even further to nearly a billion devices, including mobile devices running the Netflix app and the website experience. We’ve even extended support to older devices that lack modern capabilities, like support for TLS and HTTPS requests. For those, we’ve enabled secure communication from client to Pushy via an encryption/decryption layer on each, allowing for confidential messages to flow between the device and server.</p><h1 id="d786" class="nw nx gu bf ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os ot bk">Scaling to handle that growth (and more)</h1><h2 id="99cc" class="ps nx gu bf ny pt pu dy oc pv pw ea og nh px py pz nl qa qb qc np qd qe qf qg bk">Growth</h2><p id="7458" class="pw-post-body-paragraph mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt gn bk">With that extended reach, Pushy has gotten busier. Over the last five years, Pushy has gone from tens of millions of concurrent connections to hundreds of millions of concurrent connections, and it regularly reaches 300,000 messages sent per second. To support this growth, we’ve revisited Pushy’s past assumptions and design decisions with an eye towards both Pushy’s future role and future stability. Pushy had been relatively hands-free operationally over the last few years, and as we updated Pushy to fit its evolving role, our goal was also to get it into a stable state for the next few years. This is particularly important as we build out new functionality that relies on Pushy; a strong, stable infrastructure foundation allows our partners to continue to build on top of Pushy with confidence.</p><p id="f063" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Throughout this evolution, we’ve been able to maintain high availability and a consistent message delivery rate, with Pushy successfully maintaining 99.999% reliability for message delivery over the last few months. When our partners want to deliver a message to a device, it’s our job to make sure they can do so.</p><p id="4520" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Here are a few of the ways we’ve evolved Pushy to handle its growing scale.</p><figure class="pc pd pe pf pg ph oz pa paragraph-image"><div role="button" tabindex="0" class="pi pj fj pk bh pl"><div class="oz pa qh"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*6yETYqbh6V9LhZcI%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*6yETYqbh6V9LhZcI%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*6yETYqbh6V9LhZcI%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*6yETYqbh6V9LhZcI%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*6yETYqbh6V9LhZcI%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*6yETYqbh6V9LhZcI%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*6yETYqbh6V9LhZcI%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*6yETYqbh6V9LhZcI 640w, https://miro.medium.com/v2/resize:fit:720/0*6yETYqbh6V9LhZcI 720w, https://miro.medium.com/v2/resize:fit:750/0*6yETYqbh6V9LhZcI 750w, https://miro.medium.com/v2/resize:fit:786/0*6yETYqbh6V9LhZcI 786w, https://miro.medium.com/v2/resize:fit:828/0*6yETYqbh6V9LhZcI 828w, https://miro.medium.com/v2/resize:fit:1100/0*6yETYqbh6V9LhZcI 1100w, https://miro.medium.com/v2/resize:fit:1400/0*6yETYqbh6V9LhZcI 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="A few of the related services in Pushy’s immediate ecosystem and the changes we’ve made for them." class="bh md pm c" width="700" height="364" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="pn ff po oz pa pp pq bf b bg z du">A few of the related services in Pushy’s immediate ecosystem and the changes we’ve made for them.</figcaption></figure><h2 id="a76c" class="ps nx gu bf ny pt pu dy oc pv pw ea og nh px py pz nl qa qb qc np qd qe qf qg bk">Message processor</h2><p id="7fed" class="pw-post-body-paragraph mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt gn bk">One aspect that we invested in was the evolution of the asynchronous message processor. The previous version of the message processor was a Mantis stream-processing job that processed messages from the message queue. It was very efficient, but it had a set job size, requiring manual intervention if we wanted to horizontally scale it, and it required manual intervention when rolling out a new version.</p><p id="c757" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">It served Pushy’s needs well for many years. As the scale of the messages being processed increased and we were making more code changes in the message processor, we found ourselves looking for something more flexible. In particular, we were looking for some of the features we enjoy with our other services: automatic horizontal scaling, canaries, automated red/black rollouts, and more observability. With this in mind, we rewrote the message processor as a standalone Spring Boot service using Netflix paved-path components. Its job is the same, but it does so with easy rollouts, canary configuration that lets us roll changes safely, and autoscaling policies we’ve defined to let it handle varying volumes.</p><p id="8436" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Rewriting always comes with a risk, and it’s never the first solution we reach for, particularly when working with a system that’s in place and working well. In this case, we found that the burden from maintaining and improving the custom stream processing job was increasing, and we made the judgment call to do the rewrite. Part of the reason we did so was the clear role that the message processor played — we weren’t rewriting a huge monolithic service, but instead a well-scoped component that had explicit goals, well-defined success criteria, and a clear path towards improvement. Since the rewrite was completed in mid-2023, the message processor component has been completely zero touch, happily automated and running reliably on its own.</p><h2 id="93ca" class="ps nx gu bf ny pt pu dy oc pv pw ea og nh px py pz nl qa qb qc np qd qe qf qg bk">Push Registry</h2><p id="4448" class="pw-post-body-paragraph mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt gn bk">For most of its life, Pushy has used <a class="af nv" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/introducing-dynomite-making-non-distributed-databases-distributed-c7bce3d89404">Dynomite</a> for keeping track of device connection metadata in its Push Registry. Dynomite is a Netflix open source wrapper around Redis that provides a few additional features like auto-sharding and cross-region replication, and it provided Pushy with low latency and easy record expiry, both of which are critical for Pushy’s workload.</p><p id="ddbc" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">As Pushy’s portfolio grew, we experienced some pain points with Dynomite. Dynomite had great performance, but it required manual scaling as the system grew. The folks on the Cloud Data Engineering (CDE) team, the ones building the paved path for internal data at Netflix, graciously helped us scale it up and make adjustments, but it ended up being an involved process as we kept growing.</p><p id="04b6" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">These pain points coincided with the introduction of KeyValue, which was a new offering from the CDE team that is roughly “HashMap as a service” for Netflix developers. KeyValue is an abstraction over the storage engine itself, which allows us to choose the best storage engine that meets our SLO needs. In our case, we value low latency — the faster we can read from KeyValue, the faster these messages can get delivered. With CDE’s help, we migrated our Push Registry to use KV instead, and we have been extremely satisfied with the result. After tuning our store for Pushy’s needs, it has been on autopilot since, appropriately scaling and serving our requests with very low latency.</p><h2 id="76b8" class="ps nx gu bf ny pt pu dy oc pv pw ea og nh px py pz nl qa qb qc np qd qe qf qg bk">Scaling Pushy horizontally and vertically</h2><p id="0445" class="pw-post-body-paragraph mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt gn bk">Most of the other services our team runs, like apiproxy, the streaming edge proxy, are CPU bound, and we have autoscaling policies that scale them horizontally when we see an increase in CPU usage. This maps well to their workload — more HTTP requests means more CPU used, and we can scale up and down accordingly.</p><p id="e4a2" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Pushy has slightly different performance characteristics, with each node maintaining many connections and delivering messages on demand. In Pushy’s case, CPU usage is consistently low, since most of the connections are parked and waiting for an occasional message. Instead of relying on CPU, we scale Pushy on the number of connections, with exponential scaling to scale faster after higher thresholds are reached. We load balance the initial HTTP requests to establish the connections and rely on a reconnect protocol where devices will reconnect every 30 minutes or so, with some staggering, that gives us a steady stream of reconnecting devices to balance connections across all available instances.</p><p id="cd71" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">For a few years, our scaling policy had been that we would add new instances when the average number of connections reached 60,000 connections per instance. For a couple hundred million devices, this meant that we were regularly running thousands of Pushy instances. We can horizontally scale Pushy to our heart’s content, but we would be less content with our bill and would have to shard Pushy further to get around NLB connection limits. This evolution effort aligned well with an internal focus on cost efficiency, and we used this as an opportunity to revisit these earlier assumptions with an eye towards efficiency.</p><p id="7d83" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Both of these would be helped by increasing the number of connections that each Pushy node could handle, reducing the total number of Pushy instances and running more efficiently with the right balance between instance type, instance cost, and maximum concurrent connections. It would also allow us to have more breathing room with the NLB limits, reducing the toil of additional sharding as we continue to grow. That being said, increasing the number of connections per node is not without its own drawbacks. When a Pushy instance goes down, the devices that were connected to it will immediately try to reconnect. By increasing the number of connections per instance, it means that we would be increasing the number of devices that would be immediately trying to reconnect. We could have a million connections per instance, but a down node would lead to a thundering herd of a million devices reconnecting at the same time.</p><p id="806f" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">This delicate balance led to us doing a deep evaluation of many instance types and performance tuning options. Striking that balance, we ended up with instances that handle an average of 200,000 connections per node, with breathing room to go up to 400,000 connections if we had to. This makes for a nice balance between CPU usage, memory usage, and the thundering herd when a device connects. We’ve also enhanced our autoscaling policies to scale exponentially; the farther we are past our target average connection count, the more instances we’ll add. These improvements have enabled Pushy to be almost entirely hands off operationally, giving us plenty of flexibility as more devices come online in different patterns.</p><h2 id="b5d7" class="ps nx gu bf ny pt pu dy oc pv pw ea og nh px py pz nl qa qb qc np qd qe qf qg bk">Reliability &amp; building a stable foundation</h2><p id="fcec" class="pw-post-body-paragraph mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt gn bk">Alongside these efforts to scale Pushy for the future, we also took a close look at our reliability after finding some connectivity edge cases during recent feature development. We found a few areas for improvement around the connection between Pushy and the device, with failures due to Pushy attempting to send messages on a connection that had failed without notifying Pushy. Ideally something like a silent failure wouldn’t happen, but we frequently see odd client behavior, particularly on older devices.</p><p id="215b" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">In collaboration with the client teams, we were able to make some improvements. On the client side, better connection handling and improvements around the reconnect flow meant that they were more likely to reconnect appropriately. In Pushy, we added additional heartbeats, idle connection cleanup, and better connection tracking, which meant that we were keeping around fewer and fewer stale connections.</p><p id="ab19" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">While these improvements were mostly around those edge cases for the feature development, they had the side benefit of bumping our message delivery rates up even further. We already had a good message delivery rate, but this additional bump has enabled Pushy to regularly average 5 9s of message delivery reliability.</p><figure class="pc pd pe pf pg ph oz pa paragraph-image"><div role="button" tabindex="0" class="pi pj fj pk bh pl"><div class="oz pa qi"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*SFyzjaMH524tYkkQ%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*SFyzjaMH524tYkkQ%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*SFyzjaMH524tYkkQ%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*SFyzjaMH524tYkkQ%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*SFyzjaMH524tYkkQ%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*SFyzjaMH524tYkkQ%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*SFyzjaMH524tYkkQ%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*SFyzjaMH524tYkkQ 640w, https://miro.medium.com/v2/resize:fit:720/0*SFyzjaMH524tYkkQ 720w, https://miro.medium.com/v2/resize:fit:750/0*SFyzjaMH524tYkkQ 750w, https://miro.medium.com/v2/resize:fit:786/0*SFyzjaMH524tYkkQ 786w, https://miro.medium.com/v2/resize:fit:828/0*SFyzjaMH524tYkkQ 828w, https://miro.medium.com/v2/resize:fit:1100/0*SFyzjaMH524tYkkQ 1100w, https://miro.medium.com/v2/resize:fit:1400/0*SFyzjaMH524tYkkQ 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="Push message delivery success rate over a recent 2-week period, staying consistently over 5 9s of reliability." class="bh md pm c" width="700" height="220" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="pn ff po oz pa pp pq bf b bg z du"><em class="pr">Push message delivery success rate over a recent 2-week period.</em></figcaption></figure><h1 id="4fa9" class="nw nx gu bf ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os ot bk">Recent developments</h1><p id="22ab" class="pw-post-body-paragraph mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt gn bk">With this stable foundation and all of these connections, what can we now do with them? This question has been the driving force behind nearly all of the recent features built on top of Pushy, and it’s an exciting question to ask, particularly as an infrastructure team.</p><h2 id="2591" class="ps nx gu bf ny pt pu dy oc pv pw ea og nh px py pz nl qa qb qc np qd qe qf qg bk">Shift towards direct push</h2><p id="832c" class="pw-post-body-paragraph mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt gn bk">The first change from Pushy’s traditional role is what we call direct push; instead of a backend service dropping the message on the asynchronous message queue, it can instead leverage the Push library to skip the asynchronous queue entirely. When called to deliver a message in the direct path, the Push library will look up the Pushy connected to the target device in the Push Registry, then send the message directly to that Pushy. Pushy will respond with a status code reflecting whether it was able to successfully deliver the message or it encountered an error, and the Push library will bubble that up to the calling code in the service.</p><figure class="pc pd pe pf pg ph oz pa paragraph-image"><div role="button" tabindex="0" class="pi pj fj pk bh pl"><div class="oz pa qj"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*PJHwCgmRYIYMVPcl%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*PJHwCgmRYIYMVPcl%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*PJHwCgmRYIYMVPcl%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*PJHwCgmRYIYMVPcl%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*PJHwCgmRYIYMVPcl%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*PJHwCgmRYIYMVPcl%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*PJHwCgmRYIYMVPcl%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*PJHwCgmRYIYMVPcl 640w, https://miro.medium.com/v2/resize:fit:720/0*PJHwCgmRYIYMVPcl 720w, https://miro.medium.com/v2/resize:fit:750/0*PJHwCgmRYIYMVPcl 750w, https://miro.medium.com/v2/resize:fit:786/0*PJHwCgmRYIYMVPcl 786w, https://miro.medium.com/v2/resize:fit:828/0*PJHwCgmRYIYMVPcl 828w, https://miro.medium.com/v2/resize:fit:1100/0*PJHwCgmRYIYMVPcl 1100w, https://miro.medium.com/v2/resize:fit:1400/0*PJHwCgmRYIYMVPcl 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="The system diagram for the direct and indirect push paths. The direct push path goes directly from a backend service to Pushy, while the indirect path goes to a decoupled message queue, which is then handled by a message processor and sent on to Pushy." class="bh md pm c" width="700" height="273" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="pn ff po oz pa pp pq bf b bg z du">The system diagram for the direct and indirect push paths.</figcaption></figure><p id="8c27" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Susheel, the original author of Pushy, added this functionality as an optional path, but for years, nearly all backend services relied on the indirect path with its “best-effort” being good enough for their use cases. In recent years, we’ve seen usage of this direct path really take off as the needs of backend services have grown. In particular, rather than being just best effort, these direct messages allow the calling service to have immediate feedback about the delivery, letting them retry if a device they’re targeting has gone offline.</p><p id="80f7" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">These days, messages sent via direct push make up the majority of messages sent through Pushy. For example, for a recent 24 hour period, direct messages averaged around 160,000 messages per second and indirect averaged at around 50,000 messages per second..</p><figure class="pc pd pe pf pg ph oz pa paragraph-image"><div class="oz pa qk"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*oCI-seLx9OMSYZQk%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*oCI-seLx9OMSYZQk%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*oCI-seLx9OMSYZQk%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*oCI-seLx9OMSYZQk%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*oCI-seLx9OMSYZQk%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*oCI-seLx9OMSYZQk%201100w,%20https://miro.medium.com/v2/resize:fit:1306/format:webp/0*oCI-seLx9OMSYZQk%201306w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 653px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*oCI-seLx9OMSYZQk 640w, https://miro.medium.com/v2/resize:fit:720/0*oCI-seLx9OMSYZQk 720w, https://miro.medium.com/v2/resize:fit:750/0*oCI-seLx9OMSYZQk 750w, https://miro.medium.com/v2/resize:fit:786/0*oCI-seLx9OMSYZQk 786w, https://miro.medium.com/v2/resize:fit:828/0*oCI-seLx9OMSYZQk 828w, https://miro.medium.com/v2/resize:fit:1100/0*oCI-seLx9OMSYZQk 1100w, https://miro.medium.com/v2/resize:fit:1306/0*oCI-seLx9OMSYZQk 1306w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 653px" /><img alt="Graph of direct vs indirect messages per second, showing around 150,000 direct messages per second and around 50,000 indirect messages per second." class="bh md pm c" width="653" height="506" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div><figcaption class="pn ff po oz pa pp pq bf b bg z du">Graph of direct vs indirect messages per second.</figcaption></figure><h2 id="7caa" class="ps nx gu bf ny pt pu dy oc pv pw ea og nh px py pz nl qa qb qc np qd qe qf qg bk">Device to device messaging</h2><p id="f571" class="pw-post-body-paragraph mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt gn bk">As we’ve thought through this evolving use case, our concept of a message sender has also evolved. What if we wanted to move past Pushy’s pattern of delivering server-side messages? What if we wanted to have a device send a message to a backend service, or maybe even to another device? Our messages had traditionally been unidirectional as we send messages from the server to the device, but we now leverage these bidirectional connections and direct device messaging to enable what we call device to device messaging. This device to device messaging supported early phone-to-TV communication in support of games like Triviaverse, and it’s the messaging foundation for our <a class="af nv" href="https://help.netflix.com/en/node/132821" rel="noopener ugc nofollow" target="_blank">Companion Mode</a> as TVs and phones communicate back and forth.</p><figure class="pc pd pe pf pg ph oz pa paragraph-image"><div class="oz pa ql"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*rA3HZj7YEo5Sp4Xjp5c6EA.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*rA3HZj7YEo5Sp4Xjp5c6EA.png" /><img alt="A screenshot of one of the authors playing Triviaquest with a mobile device as the controller." class="bh md pm c" width="596" height="1049" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div><figcaption class="pn ff po oz pa pp pq bf b bg z du">A screenshot of one of the authors playing Triviaquest with a mobile device as the controller.</figcaption></figure><p id="15e9" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">This requires higher level knowledge of the system, where we need to know not just information about a single device, but more broader information, like what devices are connected for an account that the phone can pair with. This also enables things like subscribing to device events to know when another device comes online and when they’re available to pair or send a message to. This has been built out with an additional service that receives device connection information from Pushy. These events, sent over a Kafka topic, let the service keep track of the device list for a given account. Devices can subscribe to these events, allowing them to receive a message from the service when another device for the same account comes online.</p><figure class="pc pd pe pf pg ph oz pa paragraph-image"><div role="button" tabindex="0" class="pi pj fj pk bh pl"><div class="oz pa qm"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*PhEf0jXvhXbx6kwN%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*PhEf0jXvhXbx6kwN%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*PhEf0jXvhXbx6kwN%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*PhEf0jXvhXbx6kwN%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*PhEf0jXvhXbx6kwN%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*PhEf0jXvhXbx6kwN%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*PhEf0jXvhXbx6kwN%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*PhEf0jXvhXbx6kwN 640w, https://miro.medium.com/v2/resize:fit:720/0*PhEf0jXvhXbx6kwN 720w, https://miro.medium.com/v2/resize:fit:750/0*PhEf0jXvhXbx6kwN 750w, https://miro.medium.com/v2/resize:fit:786/0*PhEf0jXvhXbx6kwN 786w, https://miro.medium.com/v2/resize:fit:828/0*PhEf0jXvhXbx6kwN 828w, https://miro.medium.com/v2/resize:fit:1100/0*PhEf0jXvhXbx6kwN 1100w, https://miro.medium.com/v2/resize:fit:1400/0*PhEf0jXvhXbx6kwN 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="Pushy and its relationship with the Device List Service for discovering other devices. Pushy reaches out to the Device List Service, and when it receives the device list in response, propagates that back to the requesting device." class="bh md pm c" width="700" height="314" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="pn ff po oz pa pp pq bf b bg z du">Pushy and its relationship with the Device List Service for discovering other devices.</figcaption></figure><p id="7742" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">This device list enables the discoverability aspect of these device to device messages. Once the devices have this knowledge of the other devices connected for the same account, they’re able to choose a target device from this list that they can then send messages to.</p><p id="e182" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Once a device has that list, it can send a message to Pushy over its WebSocket connection with that device as the target in what we call a <em class="nu">device to device message</em> (1 in the diagram below). Pushy looks up the target device’s metadata in the Push registry (2) and sends the message to the second Pushy that the target device is connected to (3), as if it was the backend service in the direct push pattern above. That Pushy delivers the message to the target device (4), and the original Pushy will receive a status code in response, which it can pass back to the source device (5).</p><figure class="pc pd pe pf pg ph oz pa paragraph-image"><div role="button" tabindex="0" class="pi pj fj pk bh pl"><div class="oz pa qn"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*dEQ1TpVfTQNs3eg4%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*dEQ1TpVfTQNs3eg4%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*dEQ1TpVfTQNs3eg4%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*dEQ1TpVfTQNs3eg4%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*dEQ1TpVfTQNs3eg4%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*dEQ1TpVfTQNs3eg4%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*dEQ1TpVfTQNs3eg4%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*dEQ1TpVfTQNs3eg4 640w, https://miro.medium.com/v2/resize:fit:720/0*dEQ1TpVfTQNs3eg4 720w, https://miro.medium.com/v2/resize:fit:750/0*dEQ1TpVfTQNs3eg4 750w, https://miro.medium.com/v2/resize:fit:786/0*dEQ1TpVfTQNs3eg4 786w, https://miro.medium.com/v2/resize:fit:828/0*dEQ1TpVfTQNs3eg4 828w, https://miro.medium.com/v2/resize:fit:1100/0*dEQ1TpVfTQNs3eg4 1100w, https://miro.medium.com/v2/resize:fit:1400/0*dEQ1TpVfTQNs3eg4 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="A basic order of events for a device to device message." class="bh md pm c" width="700" height="310" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="pn ff po oz pa pp pq bf b bg z du">A basic order of events for a device to device message.</figcaption></figure><h2 id="8a43" class="ps nx gu bf ny pt pu dy oc pv pw ea og nh px py pz nl qa qb qc np qd qe qf qg bk">The messaging protocol</h2><p id="1ace" class="pw-post-body-paragraph mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt gn bk">We’ve defined a basic JSON-based message protocol for device to device messaging that lets these messages be passed from the source device to the target device. As a networking team, we naturally lean towards abstracting the communication layer with encapsulation wherever possible. This generalized message means that device teams are able to define their own protocols on top of these messages — Pushy would just be the transport layer, happily forwarding messages back and forth.</p><figure class="pc pd pe pf pg ph oz pa paragraph-image"><div class="oz pa qo"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*4-ijw8c0BTX9r20jVIgKNA.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*4-ijw8c0BTX9r20jVIgKNA.png" /><img alt="A simple block diagram showing the client app protocol on top of the device to device protocol, which itself is on top of the WebSocket &amp; Pushy protocol." class="bh md pm c" width="354" height="194" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div><figcaption class="pn ff po oz pa pp pq bf b bg z du">The client app protocol, built on top of the device to device protocol, built on top of Pushy.</figcaption></figure><p id="c46f" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">This generalization paid off in terms of investment and operational support. We built the majority of this functionality in October 2022, and we’ve only needed small tweaks since then. We needed nearly no modifications as client teams built out the functionality on top of this layer, defining the higher level application-specific protocols that powered the features they were building. We really do enjoy working with our partner teams, but if we’re able to give them the freedom to build on top of our infrastructure layer without us getting involved, then we’re able to increase their velocity, make their lives easier, and play our infrastructure roles as message platform providers.</p><p id="2933" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">With early features in experimentation, Pushy sees an average of 1000 device to device messages per second, a number that will only continue to grow.</p><figure class="pc pd pe pf pg ph oz pa paragraph-image"><div role="button" tabindex="0" class="pi pj fj pk bh pl"><div class="oz pa qp"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*6gn9UvREat4OqRoU%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*6gn9UvREat4OqRoU%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*6gn9UvREat4OqRoU%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*6gn9UvREat4OqRoU%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*6gn9UvREat4OqRoU%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*6gn9UvREat4OqRoU%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*6gn9UvREat4OqRoU%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*6gn9UvREat4OqRoU 640w, https://miro.medium.com/v2/resize:fit:720/0*6gn9UvREat4OqRoU 720w, https://miro.medium.com/v2/resize:fit:750/0*6gn9UvREat4OqRoU 750w, https://miro.medium.com/v2/resize:fit:786/0*6gn9UvREat4OqRoU 786w, https://miro.medium.com/v2/resize:fit:828/0*6gn9UvREat4OqRoU 828w, https://miro.medium.com/v2/resize:fit:1100/0*6gn9UvREat4OqRoU 1100w, https://miro.medium.com/v2/resize:fit:1400/0*6gn9UvREat4OqRoU 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="Graph of device to device messages per second, showing an average of 1000 messages per second." class="bh md pm c" width="700" height="425" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="pn ff po oz pa pp pq bf b bg z du">Graph of device to device messages per second.</figcaption></figure><h2 id="03fc" class="ps nx gu bf ny pt pu dy oc pv pw ea og nh px py pz nl qa qb qc np qd qe qf qg bk">The Netty-gritty details</h2><p id="aaec" class="pw-post-body-paragraph mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt gn bk">In Pushy, we handle incoming WebSocket messages in our PushClientProtocolHandler (<a class="af nv" href="https://github.com/Netflix/zuul/blob/99ef8841c8b7b82536d5fb193fd751c675c9ad0d/zuul-core/src/main/java/com/netflix/zuul/netty/server/push/PushClientProtocolHandler.java" rel="noopener ugc nofollow" target="_blank">code pointer to class in Zuul that we extend</a>), which extends Netty’s ChannelInboundHandlerAdapter and is added to the Netty pipeline for each client connection. We listen for incoming WebSocket messages from the connected device in its channelRead method and parse the incoming message. If it’s a device to device message, we pass the message, the ChannelHandlerContext, and the PushUserAuth information about the connection’s identity to our DeviceToDeviceManager.</p><figure class="pc pd pe pf pg ph oz pa paragraph-image"><div role="button" tabindex="0" class="pi pj fj pk bh pl"><div class="oz pa qq"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*cp-lfclw0ayykX2H%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*cp-lfclw0ayykX2H%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*cp-lfclw0ayykX2H%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*cp-lfclw0ayykX2H%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*cp-lfclw0ayykX2H%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*cp-lfclw0ayykX2H%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*cp-lfclw0ayykX2H%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*cp-lfclw0ayykX2H 640w, https://miro.medium.com/v2/resize:fit:720/0*cp-lfclw0ayykX2H 720w, https://miro.medium.com/v2/resize:fit:750/0*cp-lfclw0ayykX2H 750w, https://miro.medium.com/v2/resize:fit:786/0*cp-lfclw0ayykX2H 786w, https://miro.medium.com/v2/resize:fit:828/0*cp-lfclw0ayykX2H 828w, https://miro.medium.com/v2/resize:fit:1100/0*cp-lfclw0ayykX2H 1100w, https://miro.medium.com/v2/resize:fit:1400/0*cp-lfclw0ayykX2H 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="A rough overview of the internal organization for these components, with the code classes described above. Inside Pushy, a Push Client Protocol handler inside a Netty Channel calls out to the Device to Device manager, which itself calls out to the Push Message Sender class that forwards the message on to the other Pushy." class="bh md pm c" width="700" height="500" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="pn ff po oz pa pp pq bf b bg z du">A rough overview of the internal organization for these components.</figcaption></figure><p id="973c" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">The DeviceToDeviceManager is responsible for validating the message, doing some bookkeeping, and kicking off an async call that validates that the device is an authorized target, looks up the Pushy for the target device in the local cache (or makes a call to the data store if it’s not found), and forwards on the message. We run this asynchronously to avoid any event loop blocking due to these calls. The DeviceToDeviceManager is also responsible for observability, with metrics around cache hits, calls to the data store, message delivery rates, and latency percentile measurements. We’ve relied heavily on these metrics for alerts and optimizations — Pushy really is a metrics service that occasionally will deliver a message or two!</p><h2 id="3a46" class="ps nx gu bf ny pt pu dy oc pv pw ea og nh px py pz nl qa qb qc np qd qe qf qg bk">Security</h2><p id="d15f" class="pw-post-body-paragraph mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt gn bk">As the edge of the Netflix cloud, security considerations are always top of mind. With every connection over HTTPS, we’ve limited these messages to just authenticated WebSocket connections, added rate limiting, and added authorization checks to ensure that a device is able to target another device — you may have the best intentions in mind, but I’d strongly prefer it if you weren’t able to send arbitrary data to my personal TV from yours (and vice versa, I’m sure!).</p><h2 id="f500" class="ps nx gu bf ny pt pu dy oc pv pw ea og nh px py pz nl qa qb qc np qd qe qf qg bk">Latency and other considerations</h2><p id="c746" class="pw-post-body-paragraph mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt gn bk">One main consideration with the products built on top of this is latency, particularly when this feature is used for anything interactive within the Netflix app.</p><p id="9d25" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">We’ve added caching to Pushy to reduce the number of lookups in the hotpath for things that are unlikely to change frequently, like a device’s allowed list of targets and the Pushy instance the target device is connected to. We have to do some lookups on the initial messages to know where to send them, but it enables us to send subsequent messages faster without any KeyValue lookups. For these requests where caching removed KeyValue from the hot path, we were able to greatly speed things up. From the incoming message arriving at Pushy to the response being sent back to the device, we reduced median latency to less than a millisecond, with the 99th percentile of latency at less than 4ms.</p><p id="b708" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Our KeyValue latency is usually very low, but we have seen brief periods of elevated read latencies due to underlying issues in our KeyValue datastore. Overall latencies increased for other parts of Pushy, like client registration, but we saw very little increase in device to device latency with this caching in place.</p><h1 id="da41" class="nw nx gu bf ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os ot bk">Cultural aspects that enable this work</h1><p id="17eb" class="pw-post-body-paragraph mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt gn bk">Pushy’s scale and system design considerations make the work technically interesting, but we also deliberately focus on non-technical aspects that have helped to drive Pushy’s growth. We focus on iterative development that solves the hardest problem first, with projects frequently starting with quick hacks or prototypes to prove out a feature. As we do this initial version, we do our best to keep an eye towards the future, allowing us to move quickly from supporting a single, focused use case to a broad, generalized solution. For example, for our cross-device messaging, we were able to solve hard problems in the early work for <em class="nu">Triviaverse</em> that we later leveraged for the generic device to device solution.</p><p id="160a" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">As one can immediately see in the system diagrams above, Pushy does not exist in a vacuum, with projects frequently involving at least half a dozen teams. Trust, experience, communication, and strong relationships all enable this to work. Our team wouldn’t exist without our platform users, and we certainly wouldn’t be here writing this post without all of the work our product and client teams do. This has also emphasized the importance of building and sharing — if we’re able to get a prototype together with a device team, we’re able to then show it off to seed ideas from other teams. It’s one thing to mention that you can send these messages, but it’s another to show off the TV responding to the first click of the phone controller button!</p><h1 id="35a3" class="nw nx gu bf ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os ot bk">The future of Pushy</h1><p id="1f62" class="pw-post-body-paragraph mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt gn bk">If there’s anything certain in this world, it’s that Pushy will continue to grow and evolve. We have many new features in the works, like WebSocket message proxying, WebSocket message tracing, a global broadcast mechanism, and subscription functionality in support of Games and Live. With all of this investment, Pushy is a stable, reinforced foundation, ready for this next generation of features.</p><p id="4db1" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">We’ll be writing about those new features as well — stay tuned for future posts.</p><p id="df27" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><em class="nu">Special thanks to our stunning colleagues </em><a class="af nv" href="https://www.linkedin.com/in/jeremy-kelly-526a30180/" rel="noopener ugc nofollow" target="_blank"><em class="nu">Jeremy Kelly</em></a><em class="nu"> and </em><a class="af nv" href="https://www.linkedin.com/in/justin-guerra-3282262b/" rel="noopener ugc nofollow" target="_blank"><em class="nu">Justin Guerra</em></a><em class="nu"> who have both been invaluable to Pushy’s growth and the WebSocket ecosystem at large. We would also like to thank our larger teams and our numerous partners for their great work; it truly takes a village!</em></p></div>]]></description>
      <link>https://netflixtechblog.com/pushy-to-the-limit-evolving-netflixs-websocket-proxy-for-the-future-b468bc0ff658</link>
      <guid>https://netflixtechblog.com/pushy-to-the-limit-evolving-netflixs-websocket-proxy-for-the-future-b468bc0ff658</guid>
      <pubDate>Tue, 10 Sep 2024 21:15:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Noisy Neighbor Detection with eBPF]]></title>
      <description><![CDATA[<div class="ab cb"><div class="ci bh fz ga gb gc"><div><div></div><p id="3108" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><em class="nu">By </em><a class="af nv" href="https://www.linkedin.com/in/josefernandezmn/" rel="noopener ugc nofollow" target="_blank"><em class="nu">Jose Fernandez</em></a><em class="nu">, </em><a class="af nv" href="https://www.linkedin.com/in/sebastien-dabdoub-2a5a0958/" rel="noopener ugc nofollow" target="_blank"><em class="nu">Sebastien Dabdoub</em></a><em class="nu">, </em><a class="af nv" href="https://www.linkedin.com/in/jason-koch-5692172/" rel="noopener ugc nofollow" target="_blank"><em class="nu">Jason Kock</em></a><em class="nu">, </em><a class="af nv" href="https://www.linkedin.com/in/artemtkachuk/" rel="noopener ugc nofollow" target="_blank"><em class="nu">Artem Tkachuk</em></a></p><p id="8709" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">The Compute and Performance Engineering teams at Netflix regularly investigate performance issues in our multi-tenant environment. The first step is determining whether the problem originates from the application or the underlying infrastructure. One issue that often complicates this process is the "noisy neighbor" problem. On <a class="af nv" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/titus-the-netflix-container-management-platform-is-now-open-source-f868c9fb5436">Titus</a>, our multi-tenant compute platform, a "noisy neighbor" refers to a container or system service that heavily utilizes the server's resources, causing performance degradation in adjacent containers. We usually focus on CPU utilization because it is our workload's most frequent source of noisy neighbor issues.</p><p id="f191" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Detecting the effects of noisy neighbors is complex. Traditional performance analysis tools such as<a class="af nv" href="https://www.brendangregg.com/perf.html" rel="noopener ugc nofollow" target="_blank"> perf</a> can introduce significant overhead, risking further performance degradation. Additionally, these tools are typically deployed after the fact, which is too late for effective investigation.Another challenge is that debugging noisy neighbor issues requires significant low-level expertise and specialized tooling<em class="nu">. </em>In this blog post, we'll reveal how we leveraged eBPF to achieve continuous, low-overhead instrumentation of the Linux scheduler, enabling effective self-serve monitoring of noisy neighbor issues. Learn how Linux kernel instrumentation can improve your infrastructure observability with deeper insights and enhanced monitoring.</p><h1 id="1e46" class="nw nx gu bf ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os ot bk">Continuous Instrumentation of the Linux Scheduler</h1><p id="792f" class="pw-post-body-paragraph mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt gn bk">To ensure the reliability of our workloads that depend on low latency responses, we instrumented the <a class="af nv" href="https://en.wikipedia.org/wiki/Run_queue" rel="noopener ugc nofollow" target="_blank">run queue</a> latency for each container, which measures the time processes spend in the scheduling queue before being dispatched to the CPU. Extended waiting in this queue can be a telltale of performance issues, especially when containers are not utilizing their total CPU allocation. Continuous instrumentation is critical to catching such matters as they emerge, and eBPF, with its hooks into the Linux scheduler with minimal overhead, enabled us to monitor run queue latency efficiently.</p><p id="048b" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">To emit a run queue latency metric, we leveraged three eBPF hooks: <strong class="my gv">sched_wakeup, sched_wakeup_new,</strong> and <strong class="my gv">sched_switch</strong>.</p></div></div><div class="oz"><div class="ab cb"><div class="ly pa lz pb ma pc cf pd cg pe ci bh"><figure class="pi pj pk pl pm oz pn po paragraph-image"><div role="button" tabindex="0" class="pp pq fj pr bh ps"><div class="pf pg ph"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*6bapyclfXZPsUIaXFM-xaQ.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*6bapyclfXZPsUIaXFM-xaQ.png" /><img alt="" class="bh md pt c" width="1000" height="563" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure></div></div></div><div class="ab cb"><div class="ci bh fz ga gb gc"><p id="184b" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">The <strong class="my gv">sched_wakeup </strong>and <strong class="my gv">sched_wakeup_new</strong> hooks are invoked when a process changes state from 'sleeping' to 'runnable.' They let us identify when a process is ready to run and is waiting for CPU time. During this event, we generate a timestamp and store it in an eBPF hash map using the process ID as the key.</p><pre class="pi pj pk pl pm pu pv pw bp px bb bk">struct {<br />    __uint(type, BPF_MAP_TYPE_HASH);<br />    __uint(max_entries, MAX_TASK_ENTRIES);<br />    __uint(key_size, sizeof(u32));<br />    __uint(value_size, sizeof(u64));<br />} runq_lat SEC(".maps");SEC("tp_btf/sched_wakeup")<br />int tp_sched_wakeup(u64 *ctx)<br />{<br />    struct task_struct *task = (void *)ctx[0];<br />    u32 pid = task-&gt;pid;<br />    u64 ts = bpf_ktime_get_ns();bpf_map_update_elem(&amp;runq_lat, &amp;pid, &amp;ts, BPF_NOEXIST);<br />    return 0;<br />}</pre><p id="5852" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Conversely, the <strong class="my gv">sched_switch</strong> hook is triggered when the CPU switches between processes. This hook provides pointers to the process currently utilizing the CPU and the process about to take over. We use the upcoming task's process ID (PID) to fetch the timestamp from the eBPF map. This timestamp represents when the process entered the queue, which we had previously stored. We then calculate the run queue latency by simply subtracting the timestamps.</p><pre class="pi pj pk pl pm pu pv pw bp px bb bk">SEC("tp_btf/sched_switch")<br />int tp_sched_switch(u64 *ctx)<br />{<br />    struct task_struct *prev = (struct task_struct *)ctx[1];<br />    struct task_struct *next = (struct task_struct *)ctx[2];<br />    u32 prev_pid = prev-&gt;pid;<br />    u32 next_pid = next-&gt;pid;// fetch timestamp of when the next task was enqueued<br />    u64 *tsp = bpf_map_lookup_elem(&amp;runq_lat, &amp;next_pid);<br />    if (tsp == NULL) {<br />        return 0; // missed enqueue<br />    }// calculate runq latency before deleting the stored timestamp<br />    u64 now = bpf_ktime_get_ns();<br />    u64 runq_lat = now - *tsp;// delete pid from enqueued map<br />    bpf_map_delete_elem(&amp;runq_lat, &amp;next_pid);<br />    ....</pre><p id="58f0" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">One of the advantages of eBPF is its ability to provide pointers to the actual kernel data structures representing processes or threads, also known as tasks in kernel terminology. This feature enables access to a wealth of information stored about a process. We required the process's cgroup ID to associate it with a container for our specific use case. However, the cgroup information in the struct is safeguarded by an<a class="af nv" href="https://elixir.bootlin.com/linux/v6.6.16/source/include/linux/sched.h#L1225" rel="noopener ugc nofollow" target="_blank"> RCU (Read Copy Update) lock</a>.</p><p id="0a15" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">To safely access this RCU-protected information, we can leverage <a class="af nv" href="https://docs.kernel.org/bpf/kfuncs.html" rel="noopener ugc nofollow" target="_blank">kfuncs</a> in eBPF. kfuncs are kernel functions that can be called from eBPF programs. There are kfuncs available to lock and unlock RCU read-side critical sections. These functions ensure that our eBPF program remains safe and efficient while retrieving the cgroup ID from the task struct.</p><pre class="pi pj pk pl pm pu pv pw bp px bb bk">void bpf_rcu_read_lock(void) __ksym;<br />void bpf_rcu_read_unlock(void) __ksym;u64 get_task_cgroup_id(struct task_struct *task)<br />{<br />    struct css_set *cgroups;<br />    u64 cgroup_id;<br />    bpf_rcu_read_lock();<br />    cgroups = task-&gt;cgroups;<br />    cgroup_id = cgroups-&gt;dfl_cgrp-&gt;kn-&gt;id;<br />    bpf_rcu_read_unlock();<br />    return cgroup_id;<br />}</pre><p id="ba4a" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Having the data ready, we must package it and send it to userspace. For this purpose, we chose the eBPF <a class="af nv" href="https://nakryiko.com/posts/bpf-ringbuf/" rel="noopener ugc nofollow" target="_blank">ring buffer</a>. It is efficient, high-performing, and user-friendly. It can handle variable-length data records and allows data reading without necessitating extra memory copying or syscalls. However, the sheer amount of data points was causing the userspace program to use too much CPU, so we implemented a rate limiter in eBPF to sample the data effectively.</p><pre class="pi pj pk pl pm pu pv pw bp px bb bk">struct {<br />    __uint(type, BPF_MAP_TYPE_RINGBUF);<br />    __uint(max_entries, 256 * 1024);<br />} events SEC(".maps");struct {<br />    __uint(type, BPF_MAP_TYPE_PERCPU_HASH);<br />    __uint(max_entries, MAX_TASK_ENTRIES);<br />    __uint(key_size, sizeof(u64));<br />    __uint(value_size, sizeof(u64));<br />} cgroup_id_to_last_event_ts SEC(".maps");struct runq_event {<br />    u64 prev_cgroup_id;<br />    u64 cgroup_id;<br />    u64 runq_lat;<br />    u64 ts;<br />};SEC("tp_btf/sched_switch")<br />int tp_sched_switch(u64 *ctx)<br />{<br />    // ....<br />    // The previous code<br />    // ....u64 prev_cgroup_id = get_task_cgroup_id(prev);<br />    u64 cgroup_id = get_task_cgroup_id(next);// per-cgroup-id-per-CPU rate-limiting <br />    // to balance observability with performance overhead<br />    u64 *last_ts = <br />        bpf_map_lookup_elem(&amp;cgroup_id_to_last_event_ts, &amp;cgroup_id);<br />    u64 last_ts_val = last_ts == NULL ? 0 : *last_ts;// check the rate limit for the cgroup_id in consideration<br />    // before doing more work<br />    if (now - last_ts_val &lt; RATE_LIMIT_NS) {<br />        // Rate limit exceeded, drop the event<br />        return 0;<br />    }struct runq_event *event;<br />    event = bpf_ringbuf_reserve(&amp;events, sizeof(*event), 0);if (event) {<br />        event-&gt;prev_cgroup_id = prev_cgroup_id;<br />        event-&gt;cgroup_id = cgroup_id;<br />        event-&gt;runq_lat = runq_lat;<br />        event-&gt;ts = now;<br />        bpf_ringbuf_submit(event, 0);<br />        // Update the last event timestamp for the current cgroup_id<br />        bpf_map_update_elem(&amp;cgroup_id_to_last_event_ts, &amp;cgroup_id,<br />            &amp;now, BPF_ANY);}return 0;<br />}</pre><p id="cd03" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Our userspace application, developed in Go, processes events from the ring buffer to emit metrics to our metrics backend, <a class="af nv" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/introducing-atlas-netflixs-primary-telemetry-platform-bd31f4d8ed9a">Atlas</a>. Each event includes a run queue latency sample with a cgroup ID, which we associate with running containers on the host. We categorize it as a system service if no such association is found. When a cgroup ID correlates with a container, we emit a percentile timer Atlas metric (<code class="cx qd qe qf pv b">runq.latency</code>) for that container. We also increment a counter metric (<code class="cx qd qe qf pv b">sched.switch.out</code>) to monitor preemptions occurring for the container's processes. Access to the prev_cgroup_id of the preempted process allows us to tag the metric with the cause of the preemption, whether it's due to a process within the same container (or cgroup), a process in another container, or a system service.</p><p id="1bb4" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">It's important to highlight that both the <code class="cx qd qe qf pv b">runq.latency</code> metric and the <code class="cx qd qe qf pv b">sched.switch.out</code> metrics are needed to determine if a container is affected by noisy neighbors, which is the goal we aim to achieve — relying solely on the runq.latency metric can lead to misconceptions. For example, if a container is at or over its cgroup CPU limit, the scheduler will throttle it, resulting in an apparent spike in run queue latency due to delays in the queue. If we were only to consider this metric, we might incorrectly attribute the performance degradation to noisy neighbors when it's actually because the container is hitting its CPU request limits. However, simultaneous spikes in both metrics, mainly when the cause is a different container or system process, clearly indicate a noisy neighbor issue.</p><h1 id="b5da" class="nw nx gu bf ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os ot bk">A Noisy Neighbor Story</h1><p id="8be9" class="pw-post-body-paragraph mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt gn bk">Below is the <code class="cx qd qe qf pv b">runq.latency</code> metric for a server running a single container with ample CPU overhead. The 99th percentile averages 83.4µs (microseconds), serving as our baseline. Although there are some spikes reaching 400µs, the latency remains within acceptable parameters.</p></div></div><div class="oz"><div class="ab cb"><div class="ly pa lz pb ma pc cf pd cg pe ci bh"><figure class="pi pj pk pl pm oz pn po paragraph-image"><div role="button" tabindex="0" class="pp pq fj pr bh ps"><div class="pf pg qg"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*_DcYxRgeDwX5i07IrdTZyA.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*_DcYxRgeDwX5i07IrdTZyA.png" /><img alt="" class="bh md pt c" width="1000" height="420" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="qh ff qi pf pg qj qk bf b bg z du">container1’s 99th percentile runq.latency averages 83µs (microseconds), with spikes up to 400µs, without adjacent containers. This serves as our baseline for a container not contending for CPU on a host.</figcaption></figure></div></div></div><div class="ab cb"><div class="ci bh fz ga gb gc"><p id="d6c6" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">At 10:35, launching <code class="cx qd qe qf pv b">container2</code>, which fully utilized all CPUs on the host, caused a significant 131-millisecond spike (131,000 microseconds) in <code class="cx qd qe qf pv b">container1</code>'s P99 run queue latency. This spike would be noticeable in the userspace application if it were serving HTTP traffic. If userspace app owners reported an unexplained latency spike, we could quickly identify the noisy neighbor issue through run queue latency metrics.</p></div></div><div class="oz"><div class="ab cb"><div class="ly pa lz pb ma pc cf pd cg pe ci bh"><figure class="pi pj pk pl pm oz pn po paragraph-image"><div role="button" tabindex="0" class="pp pq fj pr bh ps"><div class="pf pg ql"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*DJrwEbrWPOxVMS0JP7uE9A.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*DJrwEbrWPOxVMS0JP7uE9A.png" /><img alt="" class="bh md pt c" width="1000" height="459" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="qh ff qi pf pg qj qk bf b bg z du">Launching container2 at 10:35, which maxes out all CPUs on the host, <strong class="bf ny">caused a 131-millisecond spike in container1’s P99 run queue latency</strong> due to increased preemptions by system processes. This indicates a noisy neighbor issue, where system services compete for CPU time with containers.</figcaption></figure></div></div></div><div class="ab cb"><div class="ci bh fz ga gb gc"><p id="1ce4" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">The <code class="cx qd qe qf pv b">sched.switch.out</code> metric indicates that the spike was due to increased preemptions by system processes, highlighting a noisy neighbor issue where system services compete with containers for CPU time. Our metrics show that the noisy neighbors were actually system processes, likely triggered by <code class="cx qd qe qf pv b">container2</code> consuming all available CPU capacity.</p><h1 id="37af" class="nw nx gu bf ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os ot bk">Optimizing eBPF Code</h1><p id="5964" class="pw-post-body-paragraph mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt gn bk">We developed an open-source eBPF process monitor called <a class="af nv" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/announcing-bpftop-streamlining-ebpf-performance-optimization-6a727c1ae2e5">bpftop</a> to measure the overhead of eBPF code in this hot kernel path. Our estimates suggest that the instrumentation adds less than 600 nanoseconds to each sched_* hook. We conducted a performance analysis on a Java service running in a container, and the instrumentation did not introduce significant overhead. The performance variance with the run queue profiling code active versus inactive was not measurable in milliseconds.</p><p id="2b6f" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">During our research on how eBPF statistics are measured in the kernel, we identified an opportunity to improve its calculation. We submitted this <a class="af nv" href="https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=ce09cbdd988887662546a1175bcfdfc6c8fdd150" rel="noopener ugc nofollow" target="_blank">patch</a>, which was included in the Linux kernel 6.10 release.</p></div></div><div class="oz"><div class="ab cb"><div class="ly pa lz pb ma pc cf pd cg pe ci bh"><figure class="pi pj pk pl pm oz pn po paragraph-image"><div role="button" tabindex="0" class="pp pq fj pr bh ps"><div class="pf pg qm"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*YD6hkXce9a70AgvSHstgWA.gif" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*YD6hkXce9a70AgvSHstgWA.gif" /><img alt="" class="bh md pt c" width="1000" height="208" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure></div></div></div><div class="ab cb"><div class="ci bh fz ga gb gc"><p id="3e54" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Through trial and error and using bpftop, we identified several optimizations that helped maintain low overhead for this code:</p><ul class=""><li id="bdd6" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt qn qo qp bk">We found that BPF_MAP_TYPE_HASH was the most performant for storing enqueued timestamps. Using BPF_MAP_TYPE_TASK_STORAGE resulted in nearly a twofold performance decline. BPF_MAP_TYPE_PERCPU_HASH was slightly less performant than BPF_MAP_TYPE_HASH, which was unexpected and requires further investigation.</li><li id="0f38" class="mw mx gu my b mz qq nb nc nd qr nf ng nh qs nj nk nl qt nn no np qu nr ns nt qn qo qp bk">The BPF_CORE_READ helper adds 20–30 nanoseconds per invocation. In the case of raw tracepoints, specifically those that are "BTF-enabled" (tp_btf/*), it is safe and more efficient to access the task struct members directly. Andrii Nakryiko recommends this approach in this <a class="af nv" href="https://nakryiko.com/posts/bpf-core-reference-guide/#btf-enabled-bpf-program-types-with-direct-memory-reads" rel="noopener ugc nofollow" target="_blank">blog post</a>.</li><li id="0f4d" class="mw mx gu my b mz qq nb nc nd qr nf ng nh qs nj nk nl qt nn no np qu nr ns nt qn qo qp bk">BPF_MAP_TYPE_LRU_HASH maps are 40–50 nanoseconds slower per operation than regular hash maps. Due to space concerns from PID churn, we initially used them for enqueued timestamps. We have since increased the map size, mitigating this risk.</li><li id="c6da" class="mw mx gu my b mz qq nb nc nd qr nf ng nh qs nj nk nl qt nn no np qu nr ns nt qn qo qp bk">The sched_switch, sched_wakeup, and sched_wakeup_new are all triggered for kernel tasks, which are identifiable by their PID of 0. We found monitoring these tasks unnecessary, so we implemented several early exit conditions and conditional logic to prevent executing costly operations, such as accessing BPF maps, when dealing with a kernel task. Notably, kernel tasks operate through the scheduler queue like any regular process.</li></ul><h1 id="d4d7" class="nw nx gu bf ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os ot bk">Conclusion</h1><p id="6194" class="pw-post-body-paragraph mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt gn bk">Our findings highlight the value of low-overhead continuous instrumentation of the Linux kernel with eBPF. We have integrated these metrics into customer dashboards, enabling actionable insights and guiding multitenancy performance discussions. We can also now use these metrics to refine CPU isolation strategies to minimize the impact of noisy neighbors. Additionally, thanks to these metrics, we've gained deeper insights into the Linux scheduler.</p><p id="4565" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">This project has also deepened our understanding of eBPF technology and underscored the importance of tools like bpftop for optimizing eBPF code. As eBPF adoption increases, we foresee more infrastructure observability and business logic shifting to it. One promising project in this space is <a class="af nv" href="https://github.com/sched-ext/scx" rel="noopener ugc nofollow" target="_blank">sched_ext</a>, potentially revolutionizing how scheduling decisions are made and tailored to specific workload needs.</p></div></div></div>]]></description>
      <link>https://netflixtechblog.com/noisy-neighbor-detection-with-ebpf-64b1f4b3bbdd</link>
      <guid>https://netflixtechblog.com/noisy-neighbor-detection-with-ebpf-64b1f4b3bbdd</guid>
      <pubDate>Tue, 10 Sep 2024 20:00:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Recommending for Long-Term Member Satisfaction at Netflix]]></title>
      <description><![CDATA[<div><div></div><p id="989b" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">By <a class="af nu" href="https://www.linkedin.com/in/jiangwei-pan-66a62a13/" rel="noopener ugc nofollow" target="_blank">Jiangwei Pan</a>, <a class="af nu" href="https://www.linkedin.com/in/thegarytang/" rel="noopener ugc nofollow" target="_blank">Gary Tang</a>, <a class="af nu" href="https://www.linkedin.com/in/henry-kang-wang-06701716/" rel="noopener ugc nofollow" target="_blank">Henry Wang</a>, and <a class="af nu" href="https://www.linkedin.com/in/jbasilico/" rel="noopener ugc nofollow" target="_blank">Justin Basilico</a></p><h1 id="1b73" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Introduction</h1><p id="9df8" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">Our mission at Netflix is to entertain the world. Our personalization algorithms play a crucial role in delivering on this mission for all members by recommending the right shows, movies, and games at the right time. This goal extends beyond immediate engagement; we aim to create an experience that brings lasting enjoyment to our members. Traditional recommender systems often optimize for short-term metrics like clicks or engagement, which may not fully capture long-term satisfaction. We strive to recommend content that not only engages members in the moment but also enhances their long-term satisfaction, which increases the value they get from Netflix, and thus they’ll be more likely to continue to be a member.</p><h1 id="36a3" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Recommendations as Contextual Bandit</h1><p id="46d8" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">One simple way we can view recommendations is as a contextual bandit problem. When a member visits, that becomes a context for our system and it selects an action of what recommendations to show, and then the member provides various types of feedback. These feedback signals can be immediate (skips, plays, thumbs up/down, or adding items to their playlist) or delayed (completing a show or renewing their subscription). We can define reward functions to reflect the quality of the recommendations from these feedback signals and then train a contextual bandit policy on historical data to maximize the expected reward.</p><figure class="pb pc pd pe pf pg oy oz paragraph-image"><div role="button" tabindex="0" class="ph pi fj pj bh pk"><div class="oy oz pa"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*Y8QDcyallv_mh7ylPzXqkA.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*Y8QDcyallv_mh7ylPzXqkA.png" /><img alt="" class="bh md pl c" width="700" height="389" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><h1 id="df36" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Improving Recommendations: Models and Objectives</h1><p id="a852" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">There are many ways that a recommendation model can be improved. They may come from more informative input features, more data, different architectures, more parameters, and so forth. In this post, we focus on a less-discussed aspect about improving the recommender objective by defining a reward function that tries to better reflect long-term member satisfaction.</p><h1 id="b2b2" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Retention as Reward?</h1><p id="e108" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">Member retention might seem like an obvious reward for optimizing long-term satisfaction because members should stay if they’re satisfied, however it has several drawbacks:</p><ul class=""><li id="9d43" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt pm pn po bk"><strong class="my gv">Noisy</strong>: Retention can be influenced by numerous external factors, such as seasonal trends, marketing campaigns, or personal circumstances unrelated to the service.</li><li id="901d" class="mw mx gu my b mz pp nb nc nd pq nf ng nh pr nj nk nl ps nn no np pt nr ns nt pm pn po bk"><strong class="my gv">Low Sensitivity</strong>: Retention is only sensitive for members on the verge of canceling their subscription, not capturing the full spectrum of member satisfaction.</li><li id="64c1" class="mw mx gu my b mz pp nb nc nd pq nf ng nh pr nj nk nl ps nn no np pt nr ns nt pm pn po bk"><strong class="my gv">Hard to Attribute</strong>: Members might cancel only after a series of bad recommendations.</li><li id="e86d" class="mw mx gu my b mz pp nb nc nd pq nf ng nh pr nj nk nl ps nn no np pt nr ns nt pm pn po bk"><strong class="my gv">Slow to Measure</strong>: We only get one signal per account per month.</li></ul><p id="26e7" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Due to these challenges, optimizing for retention alone is impractical.</p><h1 id="c903" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Proxy Rewards</h1><p id="216f" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">Instead, we can train our bandit policy to optimize a proxy reward function that is highly aligned with long-term member satisfaction while being sensitive to individual recommendations. The proxy reward <em class="pu">r(user, item)</em> is a function of user interaction with the recommended item. For example, if we recommend “One Piece” and a member plays then subsequently completes and gives it a thumbs-up, a simple proxy reward might be defined as <em class="pu">r(user, item) = f(play, complete, thumb)</em>.</p><figure class="pb pc pd pe pf pg oy oz paragraph-image"><div role="button" tabindex="0" class="ph pi fj pj bh pk"><div class="oy oz pv"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*xfSMqEoF0I2_qOPu%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*xfSMqEoF0I2_qOPu%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*xfSMqEoF0I2_qOPu%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*xfSMqEoF0I2_qOPu%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*xfSMqEoF0I2_qOPu%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*xfSMqEoF0I2_qOPu%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*xfSMqEoF0I2_qOPu%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*xfSMqEoF0I2_qOPu 640w, https://miro.medium.com/v2/resize:fit:720/0*xfSMqEoF0I2_qOPu 720w, https://miro.medium.com/v2/resize:fit:750/0*xfSMqEoF0I2_qOPu 750w, https://miro.medium.com/v2/resize:fit:786/0*xfSMqEoF0I2_qOPu 786w, https://miro.medium.com/v2/resize:fit:828/0*xfSMqEoF0I2_qOPu 828w, https://miro.medium.com/v2/resize:fit:1100/0*xfSMqEoF0I2_qOPu 1100w, https://miro.medium.com/v2/resize:fit:1400/0*xfSMqEoF0I2_qOPu 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pl c" width="700" height="186" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><h2 id="d77d" class="pw nw gu bf nx px py dy ob pz qa ea of nh qb qc qd nl qe qf qg np qh qi qj qk bk">Click-through rate (CTR)</h2><p id="1f77" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">Click-through rate (CTR), or in our case play-through rate, can be viewed as a simple proxy reward where <em class="pu">r(user, item) </em>= 1 if the user clicks a recommendation and 0 otherwise. CTR is a common feedback signal that generally reflects user preference expectations. It is a simple yet strong baseline for many recommendation applications. In some cases, such as ads personalization where the click is the target action, CTR may even be a reasonable reward for production models. However, in most cases, over-optimizing CTR can lead to promoting clickbaity items, which may harm long-term satisfaction.</p><h2 id="29ad" class="pw nw gu bf nx px py dy ob pz qa ea of nh qb qc qd nl qe qf qg np qh qi qj qk bk">Beyond CTR</h2><p id="f461" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">To align the proxy reward function more closely with long-term satisfaction, we need to look beyond simple interactions, consider all types of user actions, and understand their true implications on user satisfaction.</p><p id="65da" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">We give a few examples in the Netflix context:</p><ul class=""><li id="911d" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt pm pn po bk"><strong class="my gv">Fast season completion </strong>✅: Completing a season of a recommended TV show in one day is a strong sign of enjoyment and long-term satisfaction.</li><li id="a04d" class="mw mx gu my b mz pp nb nc nd pq nf ng nh pr nj nk nl ps nn no np pt nr ns nt pm pn po bk"><strong class="my gv">Thumbs-down after completion </strong>❌: Completing a TV show in several weeks followed by a thumbs-down indicates low satisfaction despite significant time spent.</li><li id="680e" class="mw mx gu my b mz pp nb nc nd pq nf ng nh pr nj nk nl ps nn no np pt nr ns nt pm pn po bk"><strong class="my gv">Playing a movie for just 10 minutes </strong>❓: In this case, the user’s satisfaction is ambiguous. The brief engagement might indicate that the user decided to abandon the movie, or it could simply mean the user was interrupted and plans to finish the movie later, perhaps the next day.</li><li id="c0e6" class="mw mx gu my b mz pp nb nc nd pq nf ng nh pr nj nk nl ps nn no np pt nr ns nt pm pn po bk"><strong class="my gv">Discovering new genres </strong>✅ ✅: Watching more Korean or game shows after “Squid Game” suggests the user is discovering something new. This discovery was likely even more valuable since it led to a variety of engagements in a new area for a member.</li></ul><h1 id="179d" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Reward Engineering</h1><p id="305f" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">Reward engineering is the iterative process of refining the proxy reward function to align with long-term member satisfaction. It is similar to feature engineering, except that it can be derived from data that isn’t available at serving time. Reward engineering involves four stages: hypothesis formation, defining a new proxy reward, training a new bandit policy, and A/B testing. Below is a simple example.</p><figure class="pb pc pd pe pf pg oy oz paragraph-image"><div role="button" tabindex="0" class="ph pi fj pj bh pk"><div class="oy oz pv"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*YRi8chIaj_OlV-Fd%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*YRi8chIaj_OlV-Fd%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*YRi8chIaj_OlV-Fd%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*YRi8chIaj_OlV-Fd%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*YRi8chIaj_OlV-Fd%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*YRi8chIaj_OlV-Fd%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*YRi8chIaj_OlV-Fd%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*YRi8chIaj_OlV-Fd 640w, https://miro.medium.com/v2/resize:fit:720/0*YRi8chIaj_OlV-Fd 720w, https://miro.medium.com/v2/resize:fit:750/0*YRi8chIaj_OlV-Fd 750w, https://miro.medium.com/v2/resize:fit:786/0*YRi8chIaj_OlV-Fd 786w, https://miro.medium.com/v2/resize:fit:828/0*YRi8chIaj_OlV-Fd 828w, https://miro.medium.com/v2/resize:fit:1100/0*YRi8chIaj_OlV-Fd 1100w, https://miro.medium.com/v2/resize:fit:1400/0*YRi8chIaj_OlV-Fd 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pl c" width="700" height="378" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><h1 id="018c" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Challenge: Delayed Feedback</h1><p id="093d" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">User feedback used in the proxy reward function is often delayed or missing. For example, a member may decide to play a recommended show for just a few minutes on the first day and take several weeks to fully complete the show. This completion feedback is therefore delayed. Additionally, some user feedback may never occur; while we may wish otherwise, not all members provide a thumbs-up or thumbs-down after completing a show, leaving us uncertain about their level of enjoyment.</p><figure class="pb pc pd pe pf pg oy oz paragraph-image"><div role="button" tabindex="0" class="ph pi fj pj bh pk"><div class="oy oz pv"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*cfammHCaAxkEjJhL%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*cfammHCaAxkEjJhL%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*cfammHCaAxkEjJhL%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*cfammHCaAxkEjJhL%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*cfammHCaAxkEjJhL%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*cfammHCaAxkEjJhL%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*cfammHCaAxkEjJhL%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*cfammHCaAxkEjJhL 640w, https://miro.medium.com/v2/resize:fit:720/0*cfammHCaAxkEjJhL 720w, https://miro.medium.com/v2/resize:fit:750/0*cfammHCaAxkEjJhL 750w, https://miro.medium.com/v2/resize:fit:786/0*cfammHCaAxkEjJhL 786w, https://miro.medium.com/v2/resize:fit:828/0*cfammHCaAxkEjJhL 828w, https://miro.medium.com/v2/resize:fit:1100/0*cfammHCaAxkEjJhL 1100w, https://miro.medium.com/v2/resize:fit:1400/0*cfammHCaAxkEjJhL 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pl c" width="700" height="264" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="d3ce" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">We could try and wait to give a longer window to observe feedback, but how long should we wait for delayed feedback before computing the proxy rewards? If we wait too long (e.g., weeks), we miss the opportunity to update the bandit policy with the latest data. In a highly dynamic environment like Netflix, a stale bandit policy can degrade the user experience and be particularly bad at recommending newer items.</p><h2 id="5011" class="pw nw gu bf nx px py dy ob pz qa ea of nh qb qc qd nl qe qf qg np qh qi qj qk bk">Solution: predict missing feedback</h2><p id="6574" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">We aim to update the bandit policy shortly after making a recommendation while also defining the proxy reward function based on all user feedback, including delayed feedback. Since delayed feedback has not been observed at the time of policy training, we can predict it. This prediction occurs for each training example with delayed feedback, using already observed feedback and other relevant information up to the training time as input features. Thus, the prediction also gets better as time progresses.</p><figure class="pb pc pd pe pf pg oy oz paragraph-image"><div role="button" tabindex="0" class="ph pi fj pj bh pk"><div class="oy oz pv"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*-dmyaQqosWyMq-UU%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*-dmyaQqosWyMq-UU%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*-dmyaQqosWyMq-UU%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*-dmyaQqosWyMq-UU%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*-dmyaQqosWyMq-UU%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*-dmyaQqosWyMq-UU%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*-dmyaQqosWyMq-UU%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*-dmyaQqosWyMq-UU 640w, https://miro.medium.com/v2/resize:fit:720/0*-dmyaQqosWyMq-UU 720w, https://miro.medium.com/v2/resize:fit:750/0*-dmyaQqosWyMq-UU 750w, https://miro.medium.com/v2/resize:fit:786/0*-dmyaQqosWyMq-UU 786w, https://miro.medium.com/v2/resize:fit:828/0*-dmyaQqosWyMq-UU 828w, https://miro.medium.com/v2/resize:fit:1100/0*-dmyaQqosWyMq-UU 1100w, https://miro.medium.com/v2/resize:fit:1400/0*-dmyaQqosWyMq-UU 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md pl c" width="700" height="309" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><p id="c5ce" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">The proxy reward is then calculated for each training example using both observed and predicted feedback. These training examples are used to update the bandit policy.</p><p id="d041" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">But aren’t we still only relying on observed feedback in the proxy reward function? Yes, because delayed feedback is predicted based on observed feedback. However, it is simpler to reason about rewards using all feedback directly. For instance, the delayed thumbs-up prediction model may be a complex neural network that takes into account all observed feedback (e.g., short-term play patterns). It’s more straightforward to define the proxy reward as a simple function of the thumbs-up feedback rather than a complex function of short-term interaction patterns. It can also be used to adjust for potential biases in how feedback is provided.</p><p id="c920" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">The reward engineering diagram is updated with an optional delayed feedback prediction step.</p><figure class="pb pc pd pe pf pg oy oz paragraph-image"><div role="button" tabindex="0" class="ph pi fj pj bh pk"><div class="oy oz ql"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*Rnu7B69daM-JY13CdtM6IQ.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*Rnu7B69daM-JY13CdtM6IQ.png" /><img alt="" class="bh md pl c" width="700" height="371" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure><h2 id="dc5a" class="pw nw gu bf nx px py dy ob pz qa ea of nh qb qc qd nl qe qf qg np qh qi qj qk bk">Two types of ML models</h2><p id="8005" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">It’s worth noting that this approach employs two types of ML models:</p><ul class=""><li id="8fd2" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt pm pn po bk"><strong class="my gv">Delayed Feedback Prediction Models</strong>: These models predict <em class="pu">p(final feedback | observed feedbacks)</em>. The predictions are used to define and compute proxy rewards for bandit policy training examples. As a result, these models are used offline during the bandit policy training.</li><li id="4114" class="mw mx gu my b mz pp nb nc nd pq nf ng nh pr nj nk nl ps nn no np pt nr ns nt pm pn po bk"><strong class="my gv">Bandit Policy Models</strong>: These models are used in the bandit policy <em class="pu">π(item | user; r)</em> to generate recommendations online and in real-time.</li></ul><h1 id="b98d" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Challenge: Online-Offline Metric Disparity</h1><p id="7da3" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">Improved input features or neural network architectures often lead to better offline model metrics (e.g., AUC for classification models). However, when these improved models are subjected to A/B testing, we often observe flat or even negative online metrics, which can quantify long-term member satisfaction.</p><p id="3112" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">This online-offline metric disparity usually occurs when the proxy reward used in the recommendation policy is not fully aligned with long-term member satisfaction. In such cases, a model may achieve higher proxy rewards (offline metrics) but result in worse long-term member satisfaction (online metrics).</p><p id="d753" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Nevertheless, the model improvement is genuine. One approach to resolve this is to further refine the proxy reward definition to align better with the improved model. When this tuning results in positive online metrics, the model improvement can be effectively productized. See [1] for more discussions on this challenge.</p><h1 id="576c" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">Summary and Open Questions</h1><p id="24ec" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">In this post, we provided an overview of our reward engineering efforts to align Netflix recommendations with long-term member satisfaction. While retention remains our north star, it is not easy to optimize directly. Therefore, our efforts focus on defining a proxy reward that is aligned with long-term satisfaction and sensitive to individual recommendations. Finally, we discussed the unique challenge of delayed user feedback at Netflix and proposed an approach that has proven effective for us. Refer to [2] for an earlier overview of the reward innovation efforts at Netflix.</p><p id="f30d" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">As we continue to improve our recommendations, several open questions remain:</p><ul class=""><li id="b35e" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt pm pn po bk">Can we learn a good proxy reward function automatically by correlating behavior with retention?</li><li id="5372" class="mw mx gu my b mz pp nb nc nd pq nf ng nh pr nj nk nl ps nn no np pt nr ns nt pm pn po bk">How long should we wait for delayed feedback before using its predicted value in policy training?</li><li id="9605" class="mw mx gu my b mz pp nb nc nd pq nf ng nh pr nj nk nl ps nn no np pt nr ns nt pm pn po bk">How can we leverage Reinforcement Learning to further align the policy with long-term satisfaction?</li></ul><h1 id="74f9" class="nv nw gu bf nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bk">References</h1><p id="d6c0" class="pw-post-body-paragraph mw mx gu my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gn bk">[1] <a class="af nu" href="https://ojs.aaai.org/aimagazine/index.php/aimagazine/article/view/18140" rel="noopener ugc nofollow" target="_blank">Deep learning for recommender systems: A Netflix case study</a>. AI Magazine 2021. Harald Steck, Linas Baltrunas, Ehtsham Elahi, Dawen Liang, Yves Raimond, Justin Basilico.</p><p id="76f7" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">[2] <a class="af nu" href="https://web.archive.org/web/20231011142826id_/https://dl.acm.org/doi/pdf/10.1145/3604915.3608873" rel="noopener ugc nofollow" target="_blank">Reward innovation for long-term member satisfaction</a>. RecSys 2023. Gary Tang, Jiangwei Pan, Henry Wang, Justin Basilico.</p></div>]]></description>
      <link>https://netflixtechblog.com/recommending-for-long-term-member-satisfaction-at-netflix-ac15cada49ef</link>
      <guid>https://netflixtechblog.com/recommending-for-long-term-member-satisfaction-at-netflix-ac15cada49ef</guid>
      <pubDate>Thu, 29 Aug 2024 03:01:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Improve Your Next Experiment by Learning Better Proxy Metrics From Past Experiments]]></title>
      <description><![CDATA[<div class="ab cb"><div class="ci bh fz ga gb gc"><div><div></div><p id="5ed4" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><em class="nu">By </em><a class="af nv" href="https://www.linkedin.com/in/aurelien-bibaut/" rel="noopener ugc nofollow" target="_blank"><em class="nu">Aurélien Bibaut</em></a><em class="nu">, </em><a class="af nv" href="https://www.linkedin.com/in/winston-chou-6491b0168/" rel="noopener ugc nofollow" target="_blank"><em class="nu">Winston Chou</em></a><em class="nu">, </em><a class="af nv" href="https://www.linkedin.com/in/simon-ejdemyr-22b920123/" rel="noopener ugc nofollow" target="_blank"><em class="nu">Simon Ejdemyr</em></a><em class="nu">, and </em><a class="af nv" href="https://www.linkedin.com/in/kallus/" rel="noopener ugc nofollow" target="_blank"><em class="nu">Nathan Kallus</em></a></p></div></div><div class="nw"><div class="ab cb"><div class="ly nx lz ny ma nz cf oa cg ob ci bh"><figure class="of og oh oi oj nw ok ol paragraph-image"><div role="button" tabindex="0" class="om on fj oo bh op"><div class="oc od oe"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*m_lKjIe460GlWr5JseoQzw.jpeg" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*m_lKjIe460GlWr5JseoQzw.jpeg" /><img alt="" class="bh md oq c" width="1000" height="385" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div></figure></div></div></div><div class="ab cb"><div class="ci bh fz ga gb gc"><p id="a71b" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">We are excited to share <a class="af nv" href="https://arxiv.org/pdf/2402.17637" rel="noopener ugc nofollow" target="_blank">our work</a> on how to learn good proxy metrics from historical experiments at <a class="af nv" href="https://kdd2024.kdd.org/" rel="noopener ugc nofollow" target="_blank">KDD 2024</a>. This work addresses a fundamental question for technology companies and academic researchers alike: how do we establish that a treatment that improves short-term (statistically sensitive) outcomes also improves long-term (statistically insensitive) outcomes? Or, faced with multiple short-term outcomes, how do we optimally trade them off for long-term benefit?</p><p id="2b17" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">For example, in an A/B test, you may observe that a product change improves the click-through rate. However, the test does not provide enough signal to measure a change in long-term retention, leaving you in the dark as to whether this treatment makes users more satisfied with your service. The click-through rate is a <em class="nu">proxy metric</em> (<em class="nu">S</em>, for surrogate, in our paper) while retention is a downstream <em class="nu">business outcome </em>or <em class="nu">north star metric </em>(<em class="nu">Y</em>). We may even have several proxy metrics, such as other types of clicks or the length of engagement after click. Taken together, these form a <em class="nu">vector</em> of proxy metrics.</p><p id="ca1c" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">The goal of our work is to understand the true relationship between the proxy metric(s) and the north star metric — so that we can assess a proxy’s ability to stand in for the north star metric, learn how to combine multiple metrics into a single best one, and better explore and compare different proxies.</p><p id="8738" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Several intuitive approaches to understanding this relationship have surprising pitfalls:</p><ul class=""><li id="9fd6" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt or os ot bk"><strong class="my gv">Looking only at user-level correlations between the proxy <em class="nu">S </em>and north star <em class="nu">Y</em>.</strong> Continuing the example from above, you may find that users with a higher click-through rate also tend to have a higher retention. But this does not mean that a <em class="nu">product change </em>that improves the click-through rate will also improve retention (in fact, promoting clickbait may have the opposite effect). This is because, as any introductory causal inference class will tell you, there are many confounders between <em class="nu">S </em>and <em class="nu">Y</em> — many of which you can never reliably observe and control for.</li><li id="1221" class="mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt or os ot bk"><strong class="my gv">Looking naively at treatment effect correlations between <em class="nu">S </em>and <em class="nu">Y.</em></strong> Suppose you are lucky enough to have many historical A/B tests. Further imagine the ordinary least squares (OLS) regression line through a scatter plot of <em class="nu">Y </em>on <em class="nu">S</em> in which each point represents the (<em class="nu">S</em>,<em class="nu">Y</em>)-treatment effect from a previous test. Even if you find that this line has a positive slope, you unfortunately <em class="nu">cannot</em> conclude that product changes that improve <em class="nu">S </em>will also improve <em class="nu">Y</em>. The reason for this is correlated measurement error — if <em class="nu">S</em> and <em class="nu">Y</em> are positively correlated in the population, then treatment arms that happen to have more users with high <em class="nu">S</em> will also have more users with high <em class="nu">Y</em>.</li></ul><p id="31fc" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Between these naive approaches, we find that the second one is the easier trap to fall into. This is because the dangers of the first approach are well-known, whereas covariances between <em class="nu">estimated</em> treatment effects can appear misleadingly causal. In reality, these covariances can be severely biased compared to what we actually care about: covariances between <em class="nu">true</em> treatment effects. In the extreme — such as when the negative effects of clickbait are substantial but clickiness and retention are highly correlated at the user level — the true relationship between <em class="nu">S </em>and <em class="nu">Y </em>can be negative even if the OLS slope is positive. Only more data per experiment could diminish this bias — using more experiments as data points will only yield more precise estimates of the badly biased slope. At first glance, this would appear to imperil any hope of using existing experiments to detect the relationship.</p><figure class="of og oh oi oj nw oc od paragraph-image"><div role="button" tabindex="0" class="om on fj oo bh op"><div class="oc od oz"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*o0Br8UYxvXPga-Sh%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*o0Br8UYxvXPga-Sh%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*o0Br8UYxvXPga-Sh%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*o0Br8UYxvXPga-Sh%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*o0Br8UYxvXPga-Sh%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*o0Br8UYxvXPga-Sh%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*o0Br8UYxvXPga-Sh%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*o0Br8UYxvXPga-Sh 640w, https://miro.medium.com/v2/resize:fit:720/0*o0Br8UYxvXPga-Sh 720w, https://miro.medium.com/v2/resize:fit:750/0*o0Br8UYxvXPga-Sh 750w, https://miro.medium.com/v2/resize:fit:786/0*o0Br8UYxvXPga-Sh 786w, https://miro.medium.com/v2/resize:fit:828/0*o0Br8UYxvXPga-Sh 828w, https://miro.medium.com/v2/resize:fit:1100/0*o0Br8UYxvXPga-Sh 1100w, https://miro.medium.com/v2/resize:fit:1400/0*o0Br8UYxvXPga-Sh 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /><img alt="" class="bh md oq c" width="700" height="241" role="presentation" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" /></picture></div></div><figcaption class="pa ff pb oc od pc pd bf b bg z du"><em class="pe">This figure shows a hypothetical treatment effect covariance matrix between S and Y (white line; negative correlation), a unit-level sampling covariance matrix creating correlated measurement errors between these metrics (black line; positive correlation), and the covariance matrix of estimated treatment effects which is a weighted combination of the first two (orange line; no correlation).</em></figcaption></figure><p id="8f12" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">To overcome this bias, we propose better ways to leverage historical experiments, inspired by techniques from the literature on weak instrumental variables. More specifically, we show that three estimators are consistent for the true proxy/north-star relationship under different constraints (the <a class="af nv" href="https://arxiv.org/pdf/2402.17637" rel="noopener ugc nofollow" target="_blank">paper</a> provides more details and should be helpful for practitioners interested in choosing the best estimator for their setting):</p><ul class=""><li id="7bee" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt or os ot bk">A <strong class="my gv">Total Covariance (TC) </strong>estimator allows us to estimate the OLS slope from a scatter plot of <em class="nu">true </em>treatment effects by subtracting the scaled measurement error covariance from the covariance of estimated treatment effects. Under the assumption that the correlated measurement error is the same across experiments (homogeneous covariances), the bias of this estimator is inversely proportional to the total number of units across all experiments, as opposed to the number of members per experiment.</li><li id="15f5" class="mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt or os ot bk"><strong class="my gv">Jackknife Instrumental Variables Estimation (JIVE)</strong> converges to the same OLS slope as the TC estimator but does not require the assumption of homogeneous covariances. JIVE eliminates correlated measurement error by removing each observation’s data from the computation of its instrumented surrogate values.</li><li id="a4aa" class="mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt or os ot bk">A <strong class="my gv">Limited Information Maximum Likelihood (LIML) </strong>estimator is statistically efficient as long as there are no direct effects between the treatment and <em class="nu">Y</em> (that is, <em class="nu">S</em> fully mediates all treatment effects on <em class="nu">Y</em>). We find that LIML is highly sensitive to this assumption and recommend TC or JIVE for most applications.</li></ul><p id="e658" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">Our methods yield linear structural models of treatment effects that are easy to interpret. As such, they are well-suited to the decentralized and rapidly-evolving practice of experimentation at Netflix, which runs <a class="af nv" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/experimentation-is-a-major-focus-of-data-science-across-netflix-f67923f8e985">thousands of experiments per year</a> on many diverse parts of the business. Each area of experimentation is staffed by independent Data Science and Engineering teams. While every team ultimately cares about the same north star metrics (e.g., long-term revenue), it is highly impractical for most teams to measure these in short-term A/B tests. Therefore, each has also developed proxies that are more sensitive and directly relevant to their work (e.g., user engagement or latency). To complicate matters more, teams are constantly innovating on these secondary metrics to find the right balance of sensitivity and long-term impact.</p><p id="7ad9" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">In this decentralized environment, linear models of treatment effects are a highly useful tool for coordinating efforts around proxy metrics and aligning them towards the north star:</p><ol class=""><li id="6178" class="mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt pf os ot bk"><strong class="my gv">Managing metric tradeoffs.</strong> Because experiments in one area can affect metrics in another area, there is a need to measure all secondary metrics in all tests, but also to understand the relative impact of these metrics on the north star. This is so we can inform decision-making when one metric trades off against another metric.</li><li id="96f2" class="mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt pf os ot bk"><strong class="my gv">Informing metrics innovation.</strong> To minimize wasted effort on metric development, it is also important to understand how metrics correlate with the north star “net of” existing metrics.</li><li id="5c65" class="mw mx gu my b mz ou nb nc nd ov nf ng nh ow nj nk nl ox nn no np oy nr ns nt pf os ot bk"><strong class="my gv">Enabling teams to work independently.</strong> Lastly, teams need simple tools in order to iterate on their own metrics. Teams may come up with dozens of variations of secondary metrics, and slow, complicated tools for evaluating these variations are unlikely to be adopted. Conversely, our models are easy and fast to fit, and are actively used to develop proxy metrics at Netflix.</li></ol><p id="0d35" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk">We are thrilled about the research and implementation of these methods at Netflix — while also continuing to strive for <strong class="my gv"><em class="nu">great and always better</em></strong>, per our <a class="af nv" href="https://jobs.netflix.com/culture" rel="noopener ugc nofollow" target="_blank">culture</a>. For example, we still have some way to go to develop a more flexible data architecture to streamline the application of these methods within Netflix. Interested in helping us? See our <a class="af nv" href="https://jobs.netflix.com/" rel="noopener ugc nofollow" target="_blank">open job postings</a>!</p><p id="f30b" class="pw-post-body-paragraph mw mx gu my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gn bk"><em class="nu">For feedback on this blog post and for supporting and making this work better, we thank Apoorva Lal, Martin Tingley, Patric Glynn, Richard McDowell, Travis Brooks, and Ayal Chen-Zion.</em></p></div></div></div>]]></description>
      <link>https://netflixtechblog.com/improve-your-next-experiment-by-learning-better-proxy-metrics-from-past-experiments-64c786c2a3ac</link>
      <guid>https://netflixtechblog.com/improve-your-next-experiment-by-learning-better-proxy-metrics-from-past-experiments-64c786c2a3ac</guid>
      <pubDate>Mon, 26 Aug 2024 17:46:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Investigation of a Cross-regional Network Performance Issue]]></title>
      <description><![CDATA[<div><div></div><p id="680a" class="pw-post-body-paragraph mw mx gt my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj"><a class="af nu" href="https://www.linkedin.com/in/hechaoli/" rel="noopener ugc nofollow" target="_blank">Hechao Li</a>, <a class="af nu" href="https://www.linkedin.com/in/rogercruz/" rel="noopener ugc nofollow" target="_blank">Roger Cruz</a></p><h1 id="a44a" class="nv nw gt be nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bj">Cloud Networking Topology</h1><p id="8b0b" class="pw-post-body-paragraph mw mx gt my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gm bj">Netflix operates a highly efficient cloud computing infrastructure that supports a wide array of applications essential for our SVOD (Subscription Video on Demand), live streaming and gaming services. Utilizing Amazon AWS, our infrastructure is hosted across multiple geographic regions worldwide. This global distribution allows our applications to deliver content more effectively by serving traffic closer to our customers. Like any distributed system, our applications occasionally require data synchronization between regions to maintain seamless service delivery.</p><p id="1fe6" class="pw-post-body-paragraph mw mx gt my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj">The following diagram shows a simplified cloud network topology for cross-region traffic.</p><figure class="pb pc pd pe pf pg oy oz paragraph-image"><div role="button" tabindex="0" class="ph pi fi pj bg pk"><div class="oy oz pa"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*RpHklRseVBeBJG6u%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*RpHklRseVBeBJG6u%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*RpHklRseVBeBJG6u%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*RpHklRseVBeBJG6u%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*RpHklRseVBeBJG6u%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*RpHklRseVBeBJG6u%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*RpHklRseVBeBJG6u%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*RpHklRseVBeBJG6u 640w, https://miro.medium.com/v2/resize:fit:720/0*RpHklRseVBeBJG6u 720w, https://miro.medium.com/v2/resize:fit:750/0*RpHklRseVBeBJG6u 750w, https://miro.medium.com/v2/resize:fit:786/0*RpHklRseVBeBJG6u 786w, https://miro.medium.com/v2/resize:fit:828/0*RpHklRseVBeBJG6u 828w, https://miro.medium.com/v2/resize:fit:1100/0*RpHklRseVBeBJG6u 1100w, https://miro.medium.com/v2/resize:fit:1400/0*RpHklRseVBeBJG6u 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /></picture></div></div></figure><h1 id="f7b1" class="nv nw gt be nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bj">The Problem At First Glance</h1><p id="f312" class="pw-post-body-paragraph mw mx gt my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gm bj">Our Cloud Network Engineering on-call team received a request to address a network issue affecting an application with cross-region traffic. Initially, it appeared that the application was experiencing timeouts, likely due to suboptimal network performance. As we all know, the longer the network path, the more devices the packets traverse, increasing the likelihood of issues. For this incident, <strong class="my gu">the client application is located in an internal subnet in the US region while the server application is located in an external subnet in a European region</strong>. Therefore, it is natural to blame the network since packets need to travel long distances through the internet.</p><p id="d4f7" class="pw-post-body-paragraph mw mx gt my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj">As network engineers, our initial reaction when the network is blamed is typically, “No, it can’t be the network,” and our task is to prove it. Given that there were no recent changes to the network infrastructure and no reported AWS issues impacting other applications, the on-call engineer suspected a noisy neighbor issue and sought assistance from the Host Network Engineering team.</p><h1 id="aa0a" class="nv nw gt be nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bj">Blame the Neighbors</h1><p id="de5a" class="pw-post-body-paragraph mw mx gt my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gm bj">In this context, a noisy neighbor issue occurs when a container shares a host with other network-intensive containers. <strong class="my gu">These noisy neighbors consume excessive network resources, causing other containers on the same host to suffer from degraded network performance. </strong>Despite each container having bandwidth limitations, oversubscription can still lead to such issues.</p><p id="f159" class="pw-post-body-paragraph mw mx gt my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj">Upon investigating other containers on the same host — most of which were part of the same application — we quickly eliminated the possibility of noisy neighbors. <strong class="my gu">The network throughput for both the problematic container and all others was significantly below the set bandwidth limits.</strong> We attempted to resolve the issue by removing these bandwidth limits, allowing the application to utilize as much bandwidth as necessary. However, the problem persisted.</p><h1 id="8f96" class="nv nw gt be nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bj">Blame the Network</h1><p id="cdc9" class="pw-post-body-paragraph mw mx gt my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gm bj">We observed some <strong class="my gu">TCP packets in the network marked with the RST flag</strong>, a flag indicating that a connection should be immediately terminated. Although the frequency of these packets was not alarmingly high, the presence of any RST packets still raised suspicion on the network. To determine whether this was indeed a network-induced issue, we conducted a tcpdump on the client. In the packet capture file, we spotted one TCP stream that was closed after exactly 30 seconds.</p><p id="e4ad" class="pw-post-body-paragraph mw mx gt my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj">SYN at 18:47:06</p><figure class="pb pc pd pe pf pg oy oz paragraph-image"><div role="button" tabindex="0" class="ph pi fi pj bg pk"><div class="oy oz pm"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*ZLnTrJNuCBe4tUry%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*ZLnTrJNuCBe4tUry%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*ZLnTrJNuCBe4tUry%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*ZLnTrJNuCBe4tUry%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*ZLnTrJNuCBe4tUry%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*ZLnTrJNuCBe4tUry%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*ZLnTrJNuCBe4tUry%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*ZLnTrJNuCBe4tUry 640w, https://miro.medium.com/v2/resize:fit:720/0*ZLnTrJNuCBe4tUry 720w, https://miro.medium.com/v2/resize:fit:750/0*ZLnTrJNuCBe4tUry 750w, https://miro.medium.com/v2/resize:fit:786/0*ZLnTrJNuCBe4tUry 786w, https://miro.medium.com/v2/resize:fit:828/0*ZLnTrJNuCBe4tUry 828w, https://miro.medium.com/v2/resize:fit:1100/0*ZLnTrJNuCBe4tUry 1100w, https://miro.medium.com/v2/resize:fit:1400/0*ZLnTrJNuCBe4tUry 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /></picture></div></div></figure><p id="6af6" class="pw-post-body-paragraph mw mx gt my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj">After the 3-way handshake (SYN,SYN-ACK,ACK), the traffic started flowing normally. Nothing strange until FIN at 18:47:36 (30 seconds later)</p><figure class="pb pc pd pe pf pg oy oz paragraph-image"><div role="button" tabindex="0" class="ph pi fi pj bg pk"><div class="oy oz pm"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*0-aCcRviD0JHcngn%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*0-aCcRviD0JHcngn%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*0-aCcRviD0JHcngn%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*0-aCcRviD0JHcngn%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*0-aCcRviD0JHcngn%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*0-aCcRviD0JHcngn%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*0-aCcRviD0JHcngn%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*0-aCcRviD0JHcngn 640w, https://miro.medium.com/v2/resize:fit:720/0*0-aCcRviD0JHcngn 720w, https://miro.medium.com/v2/resize:fit:750/0*0-aCcRviD0JHcngn 750w, https://miro.medium.com/v2/resize:fit:786/0*0-aCcRviD0JHcngn 786w, https://miro.medium.com/v2/resize:fit:828/0*0-aCcRviD0JHcngn 828w, https://miro.medium.com/v2/resize:fit:1100/0*0-aCcRviD0JHcngn 1100w, https://miro.medium.com/v2/resize:fit:1400/0*0-aCcRviD0JHcngn 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /></picture></div></div></figure><p id="f0da" class="pw-post-body-paragraph mw mx gt my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj">The packet capture results clearly indicated that <strong class="my gu">it was the client application that initiated the connection termination by sending a FIN packet</strong>. Following this, the server continued to send data; however, since the client had already decided to close the connection, it responded with RST packets to all subsequent data from the server.</p><p id="c1f7" class="pw-post-body-paragraph mw mx gt my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj">To ensure that the client wasn’t closing the connection due to packet loss, we also conducted a packet capture on the server side to verify that all packets sent by the server were received. This task was complicated by the fact that the packets passed through a NAT gateway (NGW), which meant that on the server side, the client’s IP and port appeared as those of the NGW, differing from those seen on the client side. Consequently, to accurately match TCP streams, <strong class="my gu">we needed to identify the TCP stream on the client side, locate the raw TCP sequence number, and then use this number as a filter on the server side to find the corresponding TCP stream.</strong></p><p id="92bd" class="pw-post-body-paragraph mw mx gt my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj">With packet capture results from both the client and server sides, we confirmed that <strong class="my gu">all packets sent by the server were correctly received before the client sent a FIN</strong>.</p><p id="1f6d" class="pw-post-body-paragraph mw mx gt my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj">Now, from the network point of view, the story is clear. The client initiated the connection requesting data from the server. The server kept sending data to the client with no problem. However, at a certain point, <strong class="my gu">despite the server still having data to send, the client chose to terminate the reception of data</strong>. This led us to suspect that the issue might be related to the client application itself.</p><h1 id="6807" class="nv nw gt be nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bj">Blame the Application</h1><p id="e481" class="pw-post-body-paragraph mw mx gt my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gm bj">In order to fully understand the problem, we now need to understand how the application works. As shown in the diagram below, the application runs in the us-east-1 region. <strong class="my gu">It reads data from cross-region servers and writes the data to consumers within the same region.</strong> The client runs as containers, whereas the servers are EC2 instances.</p><p id="e82c" class="pw-post-body-paragraph mw mx gt my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj"><strong class="my gu">Notably, the cross-region read was problematic </strong>while the write path was smooth. Most importantly, there is a 30-second application-level timeout for reading the data. The application (client) errors out if it fails to read an initial batch of data from the servers within 30 seconds. When we increased this timeout to 60 seconds, everything worked as expected. <strong class="my gu">This explains why the client initiated a FIN — because it lost patience waiting for the server to transfer data</strong>.</p><figure class="pb pc pd pe pf pg oy oz paragraph-image"><div role="button" tabindex="0" class="ph pi fi pj bg pk"><div class="oy oz pn"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*zNmSGl1_5vtOHETn%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*zNmSGl1_5vtOHETn%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*zNmSGl1_5vtOHETn%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*zNmSGl1_5vtOHETn%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*zNmSGl1_5vtOHETn%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*zNmSGl1_5vtOHETn%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*zNmSGl1_5vtOHETn%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*zNmSGl1_5vtOHETn 640w, https://miro.medium.com/v2/resize:fit:720/0*zNmSGl1_5vtOHETn 720w, https://miro.medium.com/v2/resize:fit:750/0*zNmSGl1_5vtOHETn 750w, https://miro.medium.com/v2/resize:fit:786/0*zNmSGl1_5vtOHETn 786w, https://miro.medium.com/v2/resize:fit:828/0*zNmSGl1_5vtOHETn 828w, https://miro.medium.com/v2/resize:fit:1100/0*zNmSGl1_5vtOHETn 1100w, https://miro.medium.com/v2/resize:fit:1400/0*zNmSGl1_5vtOHETn 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /></picture></div></div></figure><p id="da9f" class="pw-post-body-paragraph mw mx gt my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj">Could it be that the server was updated to send data more slowly? Could it be that the client application was updated to receive data more slowly? Could it be that the data volume became too large to be completely sent out within 30 seconds? Sadly, <strong class="my gu">we received negative answers for all 3 questions from the application owner.</strong> The server had been operating without changes for over a year, there were no significant updates in the latest rollout of the client, and the data volume had remained consistent.</p><h1 id="76c8" class="nv nw gt be nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bj">Blame the Kernel</h1><p id="efa1" class="pw-post-body-paragraph mw mx gt my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gm bj">If both the network and the application weren’t changed recently, then what changed? In fact, we discovered that the issue coincided with a recent <strong class="my gu">Linux kernel upgrade from version 6.5.13 to 6.6.10</strong>. To test this hypothesis, we rolled back the kernel upgrade and it did restore normal operation to the application.</p><p id="9ee2" class="pw-post-body-paragraph mw mx gt my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj">Honestly speaking, at that time I didn’t believe it was a kernel bug because I assumed the TCP implementation in the kernel should be solid and stable (Spoiler alert: How wrong was I!). But we were also out of ideas from other angles.</p><p id="f536" class="pw-post-body-paragraph mw mx gt my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj">There were about 14k commits between the good and bad kernel versions. Engineers on the team methodically and diligently bisected between the two versions. When the bisecting was narrowed to a couple of commits, <strong class="my gu">a change with “tcp” in its commit message caught our attention. The final bisecting confirmed that </strong><a class="af nu" href="https://lore.kernel.org/netdev/20230717152917.751987-1-edumazet@google.com/T/" rel="noopener ugc nofollow" target="_blank"><strong class="my gu">this commit</strong></a><strong class="my gu"> was our culprit</strong>.</p><p id="cc01" class="pw-post-body-paragraph mw mx gt my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj">Interestingly, while reviewing the email history related to this commit, we found that <a class="af nu" href="https://github.com/eventlet/eventlet/issues/821" rel="noopener ugc nofollow" target="_blank">another user had reported a Python test failure following the same kernel upgrade</a>. Although their solution was not directly applicable to our situation, it suggested that <strong class="my gu">a simpler test might also reproduce our problem</strong>. Using <em class="po">strace</em>, we observed that the application configured the following socket options when communicating with the server:</p><pre class="pb pc pd pe pf pp pq pr bo ps ba bj">[pid 1699] setsockopt(917, SOL_IPV6, IPV6_V6ONLY, [0], 4) = 0<br />[pid 1699] setsockopt(917, SOL_SOCKET, SO_KEEPALIVE, [1], 4) = 0<br />[pid 1699] setsockopt(917, SOL_SOCKET, SO_SNDBUF, [131072], 4) = 0<br />[pid 1699] setsockopt(917, SOL_SOCKET, SO_RCVBUF, [65536], 4) = 0<br />[pid 1699] setsockopt(917, SOL_TCP, TCP_NODELAY, [1], 4) = 0</pre><p id="3e18" class="pw-post-body-paragraph mw mx gt my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj">We then developed a minimal client-server C application that transfers a file from the server to the client, with the client configuring the same set of socket options. During testing, we used a 10M file, which represents the volume of data typically transferred within 30 seconds before the client issues a FIN. <strong class="my gu">On the old kernel, this cross-region transfer completed in 22 seconds, whereas on the new kernel, it took 39 seconds to finish.</strong></p><h1 id="ed1b" class="nv nw gt be nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bj">The Root Cause</h1><p id="1490" class="pw-post-body-paragraph mw mx gt my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gm bj">With the help of the minimal reproduction setup, we were ultimately able to pinpoint the root cause of the problem. In order to understand the root cause, it’s essential to have a grasp of the TCP receive window.</p><h2 id="6a6e" class="py nw gt be nx pz qa dx ob qb qc dz of nh qd qe qf nl qg qh qi np qj qk ql qm bj">TCP Receive Window</h2><p id="a997" class="pw-post-body-paragraph mw mx gt my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gm bj">Simply put, <strong class="my gu">the TCP receive window is how the receiver tells the sender “This is how many bytes you can send me without me ACKing any of them”</strong>. Assuming the sender is the server and the receiver is the client, then we have:</p><figure class="pb pc pd pe pf pg oy oz paragraph-image"><div role="button" tabindex="0" class="ph pi fi pj bg pk"><div class="oy oz qn"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*98gJP81W46nhdonq%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*98gJP81W46nhdonq%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*98gJP81W46nhdonq%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*98gJP81W46nhdonq%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*98gJP81W46nhdonq%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*98gJP81W46nhdonq%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*98gJP81W46nhdonq%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*98gJP81W46nhdonq 640w, https://miro.medium.com/v2/resize:fit:720/0*98gJP81W46nhdonq 720w, https://miro.medium.com/v2/resize:fit:750/0*98gJP81W46nhdonq 750w, https://miro.medium.com/v2/resize:fit:786/0*98gJP81W46nhdonq 786w, https://miro.medium.com/v2/resize:fit:828/0*98gJP81W46nhdonq 828w, https://miro.medium.com/v2/resize:fit:1100/0*98gJP81W46nhdonq 1100w, https://miro.medium.com/v2/resize:fit:1400/0*98gJP81W46nhdonq 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /></picture></div></div></figure><h2 id="f567" class="py nw gt be nx pz qa dx ob qb qc dz of nh qd qe qf nl qg qh qi np qj qk ql qm bj">The Window Size</h2><p id="0d83" class="pw-post-body-paragraph mw mx gt my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gm bj">Now that we know the TCP receive window size could affect the throughput, the question is, how is the window size calculated? As an application writer, you can’t decide the window size, however, you can decide how much memory you want to use for buffering received data. This is configured using <strong class="my gu"><em class="po">SO_RCVBUF</em> socket option</strong> we saw in the <em class="po">strace</em> result above. However, note that the value of this option means how much <strong class="my gu">application data</strong> can be queued in the receive buffer. In <a class="af nu" href="https://man7.org/linux/man-pages/man7/socket.7.html" rel="noopener ugc nofollow" target="_blank">man 7 socket</a>, there is</p><blockquote class="qo qp qq"><p id="a73f" class="mw mx po my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj">SO_RCVBUF</p><p id="ed59" class="mw mx po my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj">Sets or gets the maximum socket receive buffer in bytes.<br /> The kernel doubles this value (to allow space for<br /> bookkeeping overhead) when it is set using setsockopt(2),<br /> and this doubled value is returned by getsockopt(2). The<br /> default value is set by the<br /> /proc/sys/net/core/rmem_default file, and the maximum<br /> allowed value is set by the /proc/sys/net/core/rmem_max<br /> file. The minimum (doubled) value for this option is 256.</p></blockquote><p id="71af" class="pw-post-body-paragraph mw mx gt my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj">This means, when the user gives a value X, then <a class="af nu" href="https://elixir.bootlin.com/linux/v6.9-rc1/source/net/core/sock.c#L976" rel="noopener ugc nofollow" target="_blank">the kernel stores 2X in the variable sk-&gt;sk_rcvbuf</a>. In other words, <strong class="my gu">the kernel assumes that the bookkeeping overhead is as much as the actual data (i.e. 50% of the sk_rcvbuf)</strong>.</p><h2 id="7337" class="py nw gt be nx pz qa dx ob qb qc dz of nh qd qe qf nl qg qh qi np qj qk ql qm bj">sysctl_tcp_adv_win_scale</h2><p id="9a3d" class="pw-post-body-paragraph mw mx gt my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gm bj">However, the assumption above may not be true because the actual overhead really depends on a lot of factors such as Maximum Transmission Unit (MTU). Therefore, <strong class="my gu">the kernel provided this <em class="po">sysctl_tcp_adv_win_scale</em> which you can use to tell the kernel what the actual overhead is</strong>. (I believe 99% of people also don’t know how to set this parameter correctly and I’m definitely one of them. You’re the kernel, if you don’t know the overhead, how can you expect me to know?).</p><p id="4e58" class="pw-post-body-paragraph mw mx gt my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj">According to <a class="af nu" href="https://docs.kernel.org/networking/ip-sysctl.html" rel="noopener ugc nofollow" target="_blank">the <em class="po">sysctl</em> doc</a>,</p><blockquote class="qo qp qq"><p id="23e2" class="mw mx po my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj"><em class="gt">tcp_adv_win_scale — INTEGER</em></p><p id="a18b" class="mw mx po my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj"><em class="gt">Obsolete since linux-6.6 Count buffering overhead as bytes/2^tcp_adv_win_scale (if tcp_adv_win_scale &gt; 0) or bytes-bytes/2^(-tcp_adv_win_scale), if it is &lt;= 0.</em></p><p id="0786" class="mw mx po my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj"><em class="gt">Possible values are [-31, 31], inclusive.</em></p><p id="2f88" class="mw mx po my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj"><em class="gt">Default: 1</em></p></blockquote><p id="ebc0" class="pw-post-body-paragraph mw mx gt my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj">For 99% of people, we’re just using the default value 1, which in turn means the overhead is calculated by <em class="po">rcvbuf/2^tcp_adv_win_scale = 1/2 * rcvbuf</em>. This matches the assumption when setting the <em class="po">SO_RCVBUF</em> value.</p><p id="289d" class="pw-post-body-paragraph mw mx gt my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj">Let’s recap. Assume you set <em class="po">SO_RCVBUF</em> to 65536, which is the value set by the application as shown in the <em class="po">setsockopt</em> syscall. Then we have:</p><ul class=""><li id="e79f" class="mw mx gt my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt qr qs qt bj">SO_RCVBUF = 65536</li><li id="72c0" class="mw mx gt my b mz qu nb nc nd qv nf ng nh qw nj nk nl qx nn no np qy nr ns nt qr qs qt bj">rcvbuf = 2 * 65536 = 131072</li><li id="e031" class="mw mx gt my b mz qu nb nc nd qv nf ng nh qw nj nk nl qx nn no np qy nr ns nt qr qs qt bj">overhead = rcvbuf / 2 = 131072 / 2 = 65536</li><li id="7e11" class="mw mx gt my b mz qu nb nc nd qv nf ng nh qw nj nk nl qx nn no np qy nr ns nt qr qs qt bj">receive window size = rcvbuf — overhead = 131072–65536 = 65536</li></ul><p id="6c6d" class="pw-post-body-paragraph mw mx gt my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj">(Note, this calculation is simplified. The real calculation is more complex.)</p><p id="4e0b" class="pw-post-body-paragraph mw mx gt my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj">In short, the receive window size before the kernel upgrade was 65536. With this window size, the application was able to transfer 10M data within 30 seconds.</p><h2 id="0ac0" class="py nw gt be nx pz qa dx ob qb qc dz of nh qd qe qf nl qg qh qi np qj qk ql qm bj">The Change</h2><p id="8abe" class="pw-post-body-paragraph mw mx gt my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gm bj"><a class="af nu" href="https://lore.kernel.org/netdev/20230717152917.751987-1-edumazet@google.com/T/" rel="noopener ugc nofollow" target="_blank">This commit</a> obsoleted <em class="po">sysctl_tcp_adv_win_scale</em> and introduced a <em class="po">scaling_ratio</em> that can more accurately calculate the overhead or window size, which is the right thing to do. With the change, the window size is now <em class="po">rcvbuf * scaling_ratio</em>.</p><p id="9bdf" class="pw-post-body-paragraph mw mx gt my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj">So how is <em class="po">scaling_ratio</em> calculated? It is calculated using <strong class="my gu"><em class="po">skb-&gt;len/skb-&gt;truesize</em></strong> where <em class="po">skb-&gt;len</em> is the length of the tcp data length in an <em class="po">skb</em> and <em class="po">truesize</em> is the total size of the <em class="po">skb</em>. <strong class="my gu">This is surely a more accurate ratio based on real data rather than a hardcoded 50%.</strong> Now, here is the next question: during the TCP handshake <strong class="my gu">before any data is transferred, how do we decide the initial <em class="po">scaling_ratio</em>? </strong>The answer is, a magic and conservative ratio was chosen with the value being roughly 0.25.</p><p id="f7cb" class="pw-post-body-paragraph mw mx gt my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj">Now we have:</p><ul class=""><li id="5f22" class="mw mx gt my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt qr qs qt bj">SO_RCVBUF = 65536</li><li id="3e62" class="mw mx gt my b mz qu nb nc nd qv nf ng nh qw nj nk nl qx nn no np qy nr ns nt qr qs qt bj">rcvbuf = 2 * 65536 = 131072</li><li id="88c1" class="mw mx gt my b mz qu nb nc nd qv nf ng nh qw nj nk nl qx nn no np qy nr ns nt qr qs qt bj">receive window size = rcvbuf * 0.25 = 131072 * 0.25 = 32768</li></ul><p id="79ea" class="pw-post-body-paragraph mw mx gt my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj">In short, <strong class="my gu">the receive window size halved after the kernel upgrade. Hence the throughput was cut in half</strong>,<strong class="my gu"> causing the data transfer time to double.</strong></p><p id="de39" class="pw-post-body-paragraph mw mx gt my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj">Naturally, you may ask, I understand that the initial window size is small, but <strong class="my gu">why doesn’t the window grow when we have a more accurate ratio of the payload later</strong> (i.e. <em class="po">skb-&gt;len/skb-&gt;truesize</em>)? With some debugging, we eventually found out that the <em class="po">scaling_ratio</em> does <a class="af nu" href="https://elixir.bootlin.com/linux/v6.7.9/source/net/ipv4/tcp_input.c#L248" rel="noopener ugc nofollow" target="_blank">get updated to a more accurate <em class="po">skb-&gt;len/skb-&gt;truesize</em></a>, which in our case is around 0.66. However, another variable, <em class="po">window_clamp</em>, is not updated accordingly. <em class="po">window_clamp</em> is the <a class="af nu" href="https://elixir.bootlin.com/linux/v6.7.9/source/include/linux/tcp.h#L256" rel="noopener ugc nofollow" target="_blank">maximum receive window allowed to be advertised</a>, which is also initialized to <em class="po">0.25 * rcvbuf </em>using the initial <em class="po">scaling_ratio</em>. As a result, <strong class="my gu">the receive window size is capped at this value and can’t grow bigger</strong>.</p><h1 id="a8f3" class="nv nw gt be nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bj">The Fix</h1><p id="e049" class="pw-post-body-paragraph mw mx gt my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gm bj">In theory, the fix is to update <em class="po">window_clamp</em> along with <em class="po">scaling_ratio</em>. However, in order to have a simple fix that doesn’t introduce other unexpected behaviors, <a class="af nu" href="https://git.kernel.org/pub/scm/linux/kernel/git/netdev/net-next.git/commit/?id=697a6c8cec03" rel="noopener ugc nofollow" target="_blank">our final fix was to increase the initial <em class="po">scaling_ratio</em> from 25% to 50%</a>. This will make the receive window size backward compatible with the original default <em class="po">sysctl_tcp_adv_win_scale</em>.</p><p id="fb5c" class="pw-post-body-paragraph mw mx gt my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj">Meanwhile, notice that the problem is not only caused by the changed kernel behavior but also by the fact that the application sets <em class="po">SO_RCVBUF</em> and has a 30-second application-level timeout. In fact, the application is Kafka Connect and both settings are the default configurations (<a class="af nu" href="https://kafka.apache.org/documentation/#connectconfigs_receive.buffer.bytes" rel="noopener ugc nofollow" target="_blank"><em class="po">receive.buffer.bytes=64k</em></a> and <a class="af nu" href="https://kafka.apache.org/documentation/#consumerconfigs_request.timeout.ms" rel="noopener ugc nofollow" target="_blank"><em class="po">request.timeout.ms=30s</em></a>). We also<a class="af nu" href="https://issues.apache.org/jira/browse/KAFKA-16496" rel="noopener ugc nofollow" target="_blank"> created a kafka ticket to change receive.buffer.bytes to -1</a> to allow Linux to auto tune the receive window.</p><h1 id="53cb" class="nv nw gt be nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bj">Conclusion</h1><p id="0152" class="pw-post-body-paragraph mw mx gt my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gm bj">This was a very interesting debugging exercise that covered many layers of Netflix’s stack and infrastructure. While it technically wasn’t the “network” to blame, this time it turned out the culprit was the software components that make up the network (i.e. the TCP implementation in the kernel).</p><p id="d2c0" class="pw-post-body-paragraph mw mx gt my b mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns nt gm bj">If tackling such technical challenges excites you, consider joining our Cloud Infrastructure Engineering teams. Explore opportunities by visiting <a class="af nu" href="https://jobs.netflix.com/" rel="noopener ugc nofollow" target="_blank">Netflix Jobs</a> and searching for Cloud Engineering positions.</p><h1 id="eb85" class="nv nw gt be nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or os bj">Acknowledgments</h1><p id="c3e7" class="pw-post-body-paragraph mw mx gt my b mz ot nb nc nd ou nf ng nh ov nj nk nl ow nn no np ox nr ns nt gm bj">Special thanks to our stunning colleagues <a class="af nu" href="https://www.linkedin.com/in/alok-tiagi-99205015/" rel="noopener ugc nofollow" target="_blank">Alok Tiagi</a>, <a class="af nu" href="https://www.linkedin.com/in/artemtkachuk/" rel="noopener ugc nofollow" target="_blank">Artem Tkachuk</a>, <a class="af nu" href="https://www.linkedin.com/in/jethanadams/" rel="noopener ugc nofollow" target="_blank">Ethan Adams</a>, <a class="af nu" href="https://www.linkedin.com/in/jorge-rodriguez-12b5595/" rel="noopener ugc nofollow" target="_blank">Jorge Rodriguez</a>, <a class="af nu" href="https://www.linkedin.com/in/nickmahilani/" rel="noopener ugc nofollow" target="_blank">Nick Mahilani</a>, <a class="af nu" href="https://tycho.pizza/" rel="noopener ugc nofollow" target="_blank">Tycho Andersen</a> and <a class="af nu" href="https://www.linkedin.com/in/vinay-rayini/" rel="noopener ugc nofollow" target="_blank">Vinay Rayini</a> for investigating and mitigating this issue. We would also like to thank Linux kernel network expert <a class="af nu" href="https://www.linkedin.com/in/eric-dumazet-ba252942/" rel="noopener ugc nofollow" target="_blank">Eric Dumazet</a> for reviewing and applying the patch.</p></div>]]></description>
      <link>https://netflixtechblog.com/investigation-of-a-cross-regional-network-performance-issue-422d6218fdf1</link>
      <guid>https://netflixtechblog.com/investigation-of-a-cross-regional-network-performance-issue-422d6218fdf1</guid>
      <pubDate>Tue, 06 Aug 2024 00:18:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Java 21 Virtual Threads - Dude, Where’s My Lock?]]></title>
      <description><![CDATA[<div class="ab ca"><div class="ch bg fy fz ga gb"><div><div><h2 id="6f13" class="pw-subtitle-paragraph hq gs gt be b hr hs ht hu hv hw hx hy hz ia ib ic id ie if cp dt">Getting real with virtual threads</h2><div></div><p id="5713" class="pw-post-body-paragraph nh ni gt nj b hr nk nl nm hu nn no np nq nr ns nt nu nv nw nx ny nz oa ob oc gm bj">By <a class="af od" href="https://www.linkedin.com/in/vfilanovsky/" rel="noopener ugc nofollow" target="_blank">Vadim Filanovsky</a>, <a class="af od" href="https://www.linkedin.com/in/mike-huang-a552781/" rel="noopener ugc nofollow" target="_blank">Mike Huang</a>, <a class="af od" href="https://www.linkedin.com/in/danny-thomas-a623413/" rel="noopener ugc nofollow" target="_blank">Danny Thomas</a> and <a class="af od" href="https://www.linkedin.com/in/martinchalupa/" rel="noopener ugc nofollow" target="_blank">Martin Chalupa</a></p><h1 id="2d29" class="oe of gt be og oh oi ht oj ok ol hw om on oo op oq or os ot ou ov ow ox oy oz bj">Intro</h1><p id="2e2e" class="pw-post-body-paragraph nh ni gt nj b hr pa nl nm hu pb no np nq pc ns nt nu pd nw nx ny pe oa ob oc gm bj">Netflix has an extensive history of using Java as our primary programming language across our vast fleet of microservices. As we pick up newer versions of Java, our JVM Ecosystem team seeks out new language features that can improve the ergonomics and performance of our systems. In a <a class="af od" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/bending-pause-times-to-your-will-with-generational-zgc-256629c9386b">recent article</a>, we detailed how our workloads benefited from switching to generational ZGC as our default garbage collector when we migrated to Java 21. Virtual threads is another feature we are excited to adopt as part of this migration.</p><p id="3630" class="pw-post-body-paragraph nh ni gt nj b hr nk nl nm hu nn no np nq nr ns nt nu nv nw nx ny nz oa ob oc gm bj">For those new to virtual threads, <a class="af od" href="https://docs.oracle.com/en/java/javase/21/core/virtual-threads.html" rel="noopener ugc nofollow" target="_blank">they are described</a> as “lightweight threads that dramatically reduce the effort of writing, maintaining, and observing high-throughput concurrent applications.” Their power comes from their ability to be suspended and resumed automatically via continuations when blocking operations occur, thus freeing the underlying operating system threads to be reused for other operations. Leveraging virtual threads can unlock higher performance when utilized in the appropriate context.</p><p id="3d96" class="pw-post-body-paragraph nh ni gt nj b hr nk nl nm hu nn no np nq nr ns nt nu nv nw nx ny nz oa ob oc gm bj">In this article we discuss one of the peculiar cases that we encountered along our path to deploying virtual threads on Java 21.</p><h1 id="17c0" class="oe of gt be og oh oi ht oj ok ol hw om on oo op oq or os ot ou ov ow ox oy oz bj">The problem</h1><p id="3256" class="pw-post-body-paragraph nh ni gt nj b hr pa nl nm hu pb no np nq pc ns nt nu pd nw nx ny pe oa ob oc gm bj">Netflix engineers raised several independent reports of intermittent timeouts and hung instances to the Performance Engineering and JVM Ecosystem teams. Upon closer examination, we noticed a set of common traits and symptoms. In all cases, the apps affected ran on Java 21 with SpringBoot 3 and embedded Tomcat serving traffic on REST endpoints. The instances that experienced the issue simply stopped serving traffic even though the JVM on those instances remained up and running. One clear symptom characterizing the onset of this issue is a persistent increase in the number of sockets in <code class="cw pf pg ph pi b">closeWait</code> state as illustrated by the graph below:</p><figure class="pm pn po pp pq pr pj pk paragraph-image"><div role="button" tabindex="0" class="ps pt fi pu bg pv"><div class="pj pk pl"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*b5oZiN2Ew96GEeZ9oIIhPA.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*b5oZiN2Ew96GEeZ9oIIhPA.png" /></picture></div></div></figure><h1 id="f9a6" class="oe of gt be og oh oi ht oj ok ol hw om on oo op oq or os ot ou ov ow ox oy oz bj">Collected diagnostics</h1><p id="4115" class="pw-post-body-paragraph nh ni gt nj b hr pa nl nm hu pb no np nq pc ns nt nu pd nw nx ny pe oa ob oc gm bj">Sockets remaining in <code class="cw pf pg ph pi b">closeWait</code> state indicate that the remote peer closed the socket, but it was never closed on the local instance, presumably because the application failed to do so. This can often indicate that the application is hanging in an abnormal state, in which case application thread dumps may reveal additional insight.</p><p id="b6ef" class="pw-post-body-paragraph nh ni gt nj b hr nk nl nm hu nn no np nq nr ns nt nu nv nw nx ny nz oa ob oc gm bj">In order to troubleshoot this issue, we first leveraged our <a class="af od" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/improved-alerting-with-atlas-streaming-eval-e691c60dc61e">alerts system</a> to catch an instance in this state. Since we periodically collect and persist thread dumps for all JVM workloads, we can often retroactively piece together the behavior by examining these thread dumps from an instance. However, we were surprised to find that all our thread dumps show a perfectly idle JVM with no clear activity. Reviewing recent changes revealed that these impacted services enabled virtual threads, and we knew that virtual thread call stacks do not show up in <code class="cw pf pg ph pi b">jstack</code>-generated thread dumps. To obtain a more complete thread dump containing the state of the virtual threads, we used the “<code class="cw pf pg ph pi b">jcmd Thread.dump_to_file</code>” command instead. As a last-ditch effort to introspect the state of JVM, we also collected a heap dump from the instance.</p><h1 id="25f0" class="oe of gt be og oh oi ht oj ok ol hw om on oo op oq or os ot ou ov ow ox oy oz bj">Analysis</h1><p id="165a" class="pw-post-body-paragraph nh ni gt nj b hr pa nl nm hu pb no np nq pc ns nt nu pd nw nx ny pe oa ob oc gm bj">Thread dumps revealed thousands of “blank” virtual threads:</p><pre class="pm pn po pp pq px pi py bo pz ba bj">#119821 "" virtual#119820 "" virtual#119823 "" virtual#120847 "" virtual#119822 "" virtual<br />...</pre><p id="1d0f" class="pw-post-body-paragraph nh ni gt nj b hr nk nl nm hu nn no np nq nr ns nt nu nv nw nx ny nz oa ob oc gm bj">These are the VTs (virtual threads) for which a thread object is created, but has not started running, and as such, has no stack trace. In fact, there were approximately the same number of blank VTs as the number of sockets in closeWait state. To make sense of what we were seeing, we need to first understand how VTs operate.</p><p id="f352" class="pw-post-body-paragraph nh ni gt nj b hr nk nl nm hu nn no np nq nr ns nt nu nv nw nx ny nz oa ob oc gm bj">A virtual thread is not mapped 1:1 to a dedicated OS-level thread. Rather, we can think of it as a task that is scheduled to a fork-join thread pool. When a virtual thread enters a blocking call, like waiting for a <code class="cw pf pg ph pi b">Future</code>, it relinquishes the OS thread it occupies and simply remains in memory until it is ready to resume. In the meantime, the OS thread can be reassigned to execute other VTs in the same fork-join pool. This allows us to multiplex a lot of VTs to just a handful of underlying OS threads. In JVM terminology, the underlying OS thread is referred to as the “carrier thread” to which a virtual thread can be “mounted” while it executes and “unmounted” while it waits. A great in-depth description of virtual thread is available in <a class="af od" href="https://openjdk.org/jeps/444" rel="noopener ugc nofollow" target="_blank">JEP 444</a>.</p><p id="fa4e" class="pw-post-body-paragraph nh ni gt nj b hr nk nl nm hu nn no np nq nr ns nt nu nv nw nx ny nz oa ob oc gm bj">In our environment, we utilize a blocking model for Tomcat, which in effect holds a worker thread for the lifespan of a request. By enabling virtual threads, Tomcat switches to virtual execution. Each incoming request creates a new virtual thread that is simply scheduled as a task on a <a class="af od" href="https://github.com/apache/tomcat/blob/10.1.24/java/org/apache/tomcat/util/threads/VirtualThreadExecutor.java" rel="noopener ugc nofollow" target="_blank">Virtual Thread Executor</a>. We can see Tomcat creates a <code class="cw pf pg ph pi b">VirtualThreadExecutor</code> <a class="af od" href="https://github.com/apache/tomcat/blob/10.1.24/java/org/apache/tomcat/util/net/AbstractEndpoint.java#L1070-L1071" rel="noopener ugc nofollow" target="_blank">here</a>.</p><p id="d97c" class="pw-post-body-paragraph nh ni gt nj b hr nk nl nm hu nn no np nq nr ns nt nu nv nw nx ny nz oa ob oc gm bj">Tying this information back to our problem, the symptoms correspond to a state when Tomcat keeps creating a new web worker VT for each incoming request, but there are no available OS threads to mount them onto.</p><h1 id="520c" class="oe of gt be og oh oi ht oj ok ol hw om on oo op oq or os ot ou ov ow ox oy oz bj">Why is Tomcat stuck?</h1><p id="31ec" class="pw-post-body-paragraph nh ni gt nj b hr pa nl nm hu pb no np nq pc ns nt nu pd nw nx ny pe oa ob oc gm bj">What happened to our OS threads and what are they busy with? As <a class="af od" href="https://docs.oracle.com/en/java/javase/21/core/virtual-threads.html#GUID-04C03FFC-066D-4857-85B9-E5A27A875AF9" rel="noopener ugc nofollow" target="_blank">described here</a>, a VT will be pinned to the underlying OS thread if it performs a blocking operation while inside a <code class="cw pf pg ph pi b">synchronized</code> block or method. This is exactly what is happening here. Here is a relevant snippet from a thread dump obtained from the stuck instance:</p><pre class="pm pn po pp pq px pi py bo pz ba bj">#119515 "" virtual<br />      java.base/jdk.internal.misc.Unsafe.park(Native Method)<br />      java.base/java.lang.VirtualThread.parkOnCarrierThread(VirtualThread.java:661)<br />      java.base/java.lang.VirtualThread.park(VirtualThread.java:593)<br />      java.base/java.lang.System$2.parkVirtualThread(System.java:2643)<br />      java.base/jdk.internal.misc.VirtualThreads.park(VirtualThreads.java:54)<br />      java.base/java.util.concurrent.locks.LockSupport.park(LockSupport.java:219)<br />      java.base/java.util.concurrent.locks.AbstractQueuedSynchronizer.acquire(AbstractQueuedSynchronizer.java:754)<br />      java.base/java.util.concurrent.locks.AbstractQueuedSynchronizer.acquire(AbstractQueuedSynchronizer.java:990)<br />      java.base/java.util.concurrent.locks.ReentrantLock$Sync.lock(ReentrantLock.java:153)<br />      java.base/java.util.concurrent.locks.ReentrantLock.lock(ReentrantLock.java:322)<br />      zipkin2.reporter.internal.CountBoundedQueue.offer(CountBoundedQueue.java:54)<br />      zipkin2.reporter.internal.AsyncReporter$BoundedAsyncReporter.report(AsyncReporter.java:230)<br />      zipkin2.reporter.brave.AsyncZipkinSpanHandler.end(AsyncZipkinSpanHandler.java:214)<br />      brave.internal.handler.NoopAwareSpanHandler$CompositeSpanHandler.end(NoopAwareSpanHandler.java:98)<br />      brave.internal.handler.NoopAwareSpanHandler.end(NoopAwareSpanHandler.java:48)<br />      brave.internal.recorder.PendingSpans.finish(PendingSpans.java:116)<br />      brave.RealSpan.finish(RealSpan.java:134)<br />      brave.RealSpan.finish(RealSpan.java:129)<br />      io.micrometer.tracing.brave.bridge.BraveSpan.end(BraveSpan.java:117)<br />      io.micrometer.tracing.annotation.AbstractMethodInvocationProcessor.after(AbstractMethodInvocationProcessor.java:67)<br />      io.micrometer.tracing.annotation.ImperativeMethodInvocationProcessor.proceedUnderSynchronousSpan(ImperativeMethodInvocationProcessor.java:98)<br />      io.micrometer.tracing.annotation.ImperativeMethodInvocationProcessor.process(ImperativeMethodInvocationProcessor.java:73)<br />      io.micrometer.tracing.annotation.SpanAspect.newSpanMethod(SpanAspect.java:59)<br />      java.base/jdk.internal.reflect.DirectMethodHandleAccessor.invoke(DirectMethodHandleAccessor.java:103)<br />      java.base/java.lang.reflect.Method.invoke(Method.java:580)<br />      org.springframework.aop.aspectj.AbstractAspectJAdvice.invokeAdviceMethodWithGivenArgs(AbstractAspectJAdvice.java:637)<br />...</pre><p id="627b" class="pw-post-body-paragraph nh ni gt nj b hr nk nl nm hu nn no np nq nr ns nt nu nv nw nx ny nz oa ob oc gm bj">In this stack trace, we enter the synchronization in <code class="cw pf pg ph pi b">brave.RealSpan.finish(<a class="af od" href="https://github.com/openzipkin/brave/blob/6.0.3/brave/src/main/java/brave/RealSpan.java#L134" rel="noopener ugc nofollow" target="_blank">RealSpan.java:134</a>)</code>. This virtual thread is effectively pinned — it is mounted to an actual OS thread even while it waits to acquire a reentrant lock. There are 3 VTs in this exact state and another VT identified as “<code class="cw pf pg ph pi b">&lt;redacted&gt; @DefaultExecutor - 46542</code>” that also follows the same code path. These 4 virtual threads are pinned while waiting to acquire a lock. Because the app is deployed on an instance with 4 vCPUs, <a class="af od" href="https://github.com/openjdk/jdk21u/blob/jdk-21.0.3-ga/src/java.base/share/classes/java/lang/VirtualThread.java#L1102-L1134" rel="noopener ugc nofollow" target="_blank">the fork-join pool that underpins VT execution</a> also contains 4 OS threads. Now that we have exhausted all of them, no other virtual thread can make any progress. This explains why Tomcat stopped processing the requests and why the number of sockets in <code class="cw pf pg ph pi b">closeWait</code> state keeps climbing. Indeed, Tomcat accepts a connection on a socket, creates a request along with a virtual thread, and passes this request/thread to the executor for processing. However, the newly created VT cannot be scheduled because all of the OS threads in the fork-join pool are pinned and never released. So these newly created VTs are stuck in the queue, while still holding the socket.</p><h1 id="caf7" class="oe of gt be og oh oi ht oj ok ol hw om on oo op oq or os ot ou ov ow ox oy oz bj">Who has the lock?</h1><p id="f2d5" class="pw-post-body-paragraph nh ni gt nj b hr pa nl nm hu pb no np nq pc ns nt nu pd nw nx ny pe oa ob oc gm bj">Now that we know VTs are waiting to acquire a lock, the next question is: Who holds the lock? Answering this question is key to understanding what triggered this condition in the first place. Usually a thread dump indicates who holds the lock with either “<code class="cw pf pg ph pi b">- locked &lt;0x…&gt; (at …)</code>” or “<code class="cw pf pg ph pi b">Locked ownable synchronizers</code>,” but neither of these show up in our thread dumps. As a matter of fact, no locking/parking/waiting information is included in the <code class="cw pf pg ph pi b">jcmd</code>-generated thread dumps. This is a limitation in Java 21 and will be addressed in the future releases. Carefully combing through the thread dump reveals that there are a total of 6 threads contending for the same <code class="cw pf pg ph pi b">ReentrantLock</code> and associated <code class="cw pf pg ph pi b">Condition</code>. Four of these six threads are detailed in the previous section. Here is another thread:</p><pre class="pm pn po pp pq px pi py bo pz ba bj">#119516 "" virtual<br />      java.base/java.lang.VirtualThread.park(VirtualThread.java:582)<br />      java.base/java.lang.System$2.parkVirtualThread(System.java:2643)<br />      java.base/jdk.internal.misc.VirtualThreads.park(VirtualThreads.java:54)<br />      java.base/java.util.concurrent.locks.LockSupport.park(LockSupport.java:219)<br />      java.base/java.util.concurrent.locks.AbstractQueuedSynchronizer.acquire(AbstractQueuedSynchronizer.java:754)<br />      java.base/java.util.concurrent.locks.AbstractQueuedSynchronizer.acquire(AbstractQueuedSynchronizer.java:990)<br />      java.base/java.util.concurrent.locks.ReentrantLock$Sync.lock(ReentrantLock.java:153)<br />      java.base/java.util.concurrent.locks.ReentrantLock.lock(ReentrantLock.java:322)<br />      zipkin2.reporter.internal.CountBoundedQueue.offer(CountBoundedQueue.java:54)<br />      zipkin2.reporter.internal.AsyncReporter$BoundedAsyncReporter.report(AsyncReporter.java:230)<br />      zipkin2.reporter.brave.AsyncZipkinSpanHandler.end(AsyncZipkinSpanHandler.java:214)<br />      brave.internal.handler.NoopAwareSpanHandler$CompositeSpanHandler.end(NoopAwareSpanHandler.java:98)<br />      brave.internal.handler.NoopAwareSpanHandler.end(NoopAwareSpanHandler.java:48)<br />      brave.internal.recorder.PendingSpans.finish(PendingSpans.java:116)<br />      brave.RealScopedSpan.finish(RealScopedSpan.java:64)<br />      ...</pre><p id="fc4e" class="pw-post-body-paragraph nh ni gt nj b hr nk nl nm hu nn no np nq nr ns nt nu nv nw nx ny nz oa ob oc gm bj">Note that while this thread seemingly goes through the same code path for finishing a span, it does not go through a <code class="cw pf pg ph pi b">synchronized</code> block. Finally here is the 6th thread:</p><pre class="pm pn po pp pq px pi py bo pz ba bj">#107 "AsyncReporter &lt;redacted&gt;"<br />      java.base/jdk.internal.misc.Unsafe.park(Native Method)<br />      java.base/java.util.concurrent.locks.LockSupport.park(LockSupport.java:221)<br />      java.base/java.util.concurrent.locks.AbstractQueuedSynchronizer.acquire(AbstractQueuedSynchronizer.java:754)<br />      java.base/java.util.concurrent.locks.AbstractQueuedSynchronizer$ConditionObject.awaitNanos(AbstractQueuedSynchronizer.java:1761)<br />      zipkin2.reporter.internal.CountBoundedQueue.drainTo(CountBoundedQueue.java:81)<br />      zipkin2.reporter.internal.AsyncReporter$BoundedAsyncReporter.flush(AsyncReporter.java:241)<br />      zipkin2.reporter.internal.AsyncReporter$Flusher.run(AsyncReporter.java:352)<br />      java.base/java.lang.Thread.run(Thread.java:1583)</pre><p id="1f0e" class="pw-post-body-paragraph nh ni gt nj b hr nk nl nm hu nn no np nq nr ns nt nu nv nw nx ny nz oa ob oc gm bj">This is actually a normal platform thread, not a virtual thread. Paying particular attention to the line numbers in this stack trace, it is peculiar that the thread seems to be blocked within the internal <code class="cw pf pg ph pi b">acquire()</code> method <em class="qf">after</em> <a class="af od" href="https://github.com/openjdk/jdk21u/blob/jdk-21.0.3-ga/src/java.base/share/classes/java/util/concurrent/locks/AbstractQueuedSynchronizer.java#L1761" rel="noopener ugc nofollow" target="_blank">completing the wait</a>. In other words, this calling thread owned the lock upon entering <code class="cw pf pg ph pi b">awaitNanos()</code>. We know the lock was explicitly acquired <a class="af od" href="https://github.com/openzipkin/zipkin-reporter-java/blob/3.4.0/core/src/main/java/zipkin2/reporter/internal/CountBoundedQueue.java#L76" rel="noopener ugc nofollow" target="_blank">here</a>. However, by the time the wait completed, it could not reacquire the lock. Summarizing our thread dump analysis:</p><figure class="pm pn po pp pq pr"><div class="qg jq l fi"></div></figure><p id="b6e7" class="pw-post-body-paragraph nh ni gt nj b hr nk nl nm hu nn no np nq nr ns nt nu nv nw nx ny nz oa ob oc gm bj">There are 5 virtual threads and 1 regular thread waiting for the lock. Out of those 5 VTs, 4 of them are pinned to the OS threads in the fork-join pool. There’s still no information on who owns the lock. As there’s nothing more we can glean from the thread dump, our next logical step is to peek into the heap dump and introspect the state of the lock.</p><h1 id="096b" class="oe of gt be og oh oi ht oj ok ol hw om on oo op oq or os ot ou ov ow ox oy oz bj">Inspecting the lock</h1><p id="a6d9" class="pw-post-body-paragraph nh ni gt nj b hr pa nl nm hu pb no np nq pc ns nt nu pd nw nx ny pe oa ob oc gm bj">Finding the lock in the heap dump was relatively straightforward. Using the excellent <a class="af od" href="https://eclipse.dev/mat/" rel="noopener ugc nofollow" target="_blank">Eclipse MAT</a> tool, we examined the objects on the stack of the <code class="cw pf pg ph pi b">AsyncReporter</code> non-virtual thread to identify the lock object. Reasoning about the current state of the lock was perhaps the trickiest part of our investigation. Most of the relevant code can be found in the <a class="af od" href="https://github.com/openjdk/jdk21u/blob/jdk-21.0.3-ga/src/java.base/share/classes/java/util/concurrent/locks/AbstractQueuedSynchronizer.java" rel="noopener ugc nofollow" target="_blank">AbstractQueuedSynchronizer.java</a>. While we don’t claim to fully understand the inner workings of it, we reverse-engineered enough of it to match against what we see in the heap dump. This diagram illustrates our findings:</p></div></div><div class="pr"><div class="ab ca"><div class="mj qj mk qk ml ql ce qm cf qn ch bg"><figure class="pm pn po pp pq pr qp qq paragraph-image"><div role="button" tabindex="0" class="ps pt fi pu bg pv"><div class="pj pk qo"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/1*6AOJeVdbhmStpb9CRj30nw.png" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/1*6AOJeVdbhmStpb9CRj30nw.png" /></picture></div></div></figure></div></div></div><div class="ab ca"><div class="ch bg fy fz ga gb"><p id="11f3" class="pw-post-body-paragraph nh ni gt nj b hr nk nl nm hu nn no np nq nr ns nt nu nv nw nx ny nz oa ob oc gm bj">First off, the <code class="cw pf pg ph pi b">exclusiveOwnerThread</code> field is <code class="cw pf pg ph pi b">null</code> (2), signifying that no one owns the lock. We have an “empty” <code class="cw pf pg ph pi b">ExclusiveNode</code> (3) at the head of the list (<code class="cw pf pg ph pi b">waiter</code> is <code class="cw pf pg ph pi b">null</code> and <code class="cw pf pg ph pi b">status</code> is cleared) followed by another <code class="cw pf pg ph pi b">ExclusiveNode</code> with <code class="cw pf pg ph pi b">waiter</code> pointing to one of the virtual threads contending for the lock — <code class="cw pf pg ph pi b">#119516</code> (4). The only place we found that clears the <code class="cw pf pg ph pi b">exclusiveOwnerThread</code> field is within the <code class="cw pf pg ph pi b">ReentrantLock.Sync.tryRelease()</code> method (<a class="af od" href="https://github.com/openjdk/jdk21u/blob/jdk-21.0.3-ga/src/java.base/share/classes/java/util/concurrent/locks/ReentrantLock.java#L178" rel="noopener ugc nofollow" target="_blank">source link</a>). There we also set <code class="cw pf pg ph pi b">state = 0</code> matching the state that we see in the heap dump (1).</p><p id="b2bc" class="pw-post-body-paragraph nh ni gt nj b hr nk nl nm hu nn no np nq nr ns nt nu nv nw nx ny nz oa ob oc gm bj">With this in mind, we traced the <a class="af od" href="https://github.com/openjdk/jdk21u/blob/jdk-21.0.3-ga/src/java.base/share/classes/java/util/concurrent/locks/AbstractQueuedSynchronizer.java#L1058-L1064" rel="noopener ugc nofollow" target="_blank">code path</a> to <code class="cw pf pg ph pi b">release()</code> the lock. After successfully calling <code class="cw pf pg ph pi b">tryRelease()</code>, the lock-holding thread attempts to <a class="af od" href="https://github.com/openjdk/jdk21u/blob/jdk-21.0.3-ga/src/java.base/share/classes/java/util/concurrent/locks/AbstractQueuedSynchronizer.java#L641-L647" rel="noopener ugc nofollow" target="_blank">signal the next waiter</a> in the list. At this point, the lock-holding thread is still at the head of the list, even though ownership of the lock is <em class="qf">effectively released</em>. The <em class="qf">next </em>node in the list points to the thread that is <em class="qf">about to acquire the lock</em>.</p><p id="a4a6" class="pw-post-body-paragraph nh ni gt nj b hr nk nl nm hu nn no np nq nr ns nt nu nv nw nx ny nz oa ob oc gm bj">To understand how this signaling works, let’s look at the <a class="af od" href="https://github.com/openjdk/jdk21u/blob/jdk-21.0.3-ga/src/java.base/share/classes/java/util/concurrent/locks/AbstractQueuedSynchronizer.java#L670-L765" rel="noopener ugc nofollow" target="_blank">lock acquire path</a> in the <code class="cw pf pg ph pi b">AbstractQueuedSynchronizer.acquire()</code> method. Grossly oversimplifying, it’s an infinite loop, where threads attempt to acquire the lock and then park if the attempt was unsuccessful:</p><pre class="pm pn po pp pq px pi py bo pz ba bj">while(true) {<br />   if (tryAcquire()) {<br />      return; // lock acquired<br />   }<br />   park();<br />}</pre><p id="b3b3" class="pw-post-body-paragraph nh ni gt nj b hr nk nl nm hu nn no np nq nr ns nt nu nv nw nx ny nz oa ob oc gm bj">When the lock-holding thread releases the lock and signals to unpark the next waiter thread, the unparked thread iterates through this loop again, giving it another opportunity to acquire the lock. Indeed, our thread dump indicates that all of our waiter threads are parked on <a class="af od" href="https://github.com/openjdk/jdk21u/blob/jdk-21.0.3-ga/src/java.base/share/classes/java/util/concurrent/locks/AbstractQueuedSynchronizer.java#L754" rel="noopener ugc nofollow" target="_blank">line 754</a>. Once unparked, the thread that managed to acquire the lock should end up in <a class="af od" href="https://github.com/openjdk/jdk21u/blob/jdk-21.0.3-ga/src/java.base/share/classes/java/util/concurrent/locks/AbstractQueuedSynchronizer.java#L716-L723" rel="noopener ugc nofollow" target="_blank">this code block</a>, effectively resetting the head of the list and clearing the reference to the waiter.</p><p id="8fcd" class="pw-post-body-paragraph nh ni gt nj b hr nk nl nm hu nn no np nq nr ns nt nu nv nw nx ny nz oa ob oc gm bj">To restate this more concisely, the lock-owning thread is referenced by the head node of the list. Releasing the lock notifies the next node in the list while acquiring the lock resets the head of the list to the current node. This means that what we see in the heap dump reflects the state when one thread has already released the lock but the next thread has yet to acquire it. It’s a weird in-between state that should be transient, but our JVM is stuck here. We know thread <code class="cw pf pg ph pi b">#119516</code> was notified and is about to acquire the lock because of the <code class="cw pf pg ph pi b">ExclusiveNode</code> state we identified at the head of the list. However, thread dumps show that thread <code class="cw pf pg ph pi b">#119516</code> continues to wait, just like other threads contending for the same lock. How can we reconcile what we see between the thread and heap dumps?</p><h1 id="2378" class="oe of gt be og oh oi ht oj ok ol hw om on oo op oq or os ot ou ov ow ox oy oz bj">The lock with no place to run</h1><p id="f1cc" class="pw-post-body-paragraph nh ni gt nj b hr pa nl nm hu pb no np nq pc ns nt nu pd nw nx ny pe oa ob oc gm bj">Knowing that thread <code class="cw pf pg ph pi b">#119516</code> was actually notified, we went back to the thread dump to re-examine the state of the threads. Recall that we have 6 total threads waiting for the lock with 4 of the virtual threads each pinned to an OS thread. These 4 will not yield their OS thread until they acquire the lock and proceed out of the <code class="cw pf pg ph pi b">synchronized</code> block. <code class="cw pf pg ph pi b">#107 “AsyncReporter &lt;redacted&gt;”</code> is a regular platform thread, so nothing should prevent it from proceeding if it acquires the lock. This leaves us with the last thread: <code class="cw pf pg ph pi b">#119516</code>. It is a VT, but it is not pinned to an OS thread. Even if it’s notified to be unparked, it cannot proceed because there are no more OS threads left in the fork-join pool to schedule it onto. That’s exactly what happens here — although <code class="cw pf pg ph pi b">#119516</code> is signaled to unpark itself, it cannot leave the parked state because the fork-join pool is occupied by the 4 other VTs waiting to acquire the same lock. None of those pinned VTs can proceed until they acquire the lock. It’s a variation of the <a class="af od" href="https://en.wikipedia.org/wiki/Deadlock" rel="noopener ugc nofollow" target="_blank">classic deadlock problem</a>, but instead of 2 locks we have one lock and a semaphore with 4 permits as represented by the fork-join pool.</p><p id="5cdd" class="pw-post-body-paragraph nh ni gt nj b hr nk nl nm hu nn no np nq nr ns nt nu nv nw nx ny nz oa ob oc gm bj">Now that we know exactly what happened, it was easy to come up with a <a class="af od" href="https://gist.github.com/DanielThomas/0b099c5f208d7deed8a83bf5fc03179e" rel="noopener ugc nofollow" target="_blank">reproducible test case</a>.</p><h1 id="ee54" class="oe of gt be og oh oi ht oj ok ol hw om on oo op oq or os ot ou ov ow ox oy oz bj">Conclusion</h1><p id="6f88" class="pw-post-body-paragraph nh ni gt nj b hr pa nl nm hu pb no np nq pc ns nt nu pd nw nx ny pe oa ob oc gm bj">Virtual threads are expected to improve performance by reducing overhead related to thread creation and context switching. Despite some sharp edges as of Java 21, virtual threads largely deliver on their promise. In our quest for more performant Java applications, we see further virtual thread adoption as a key towards unlocking that goal. We look forward to Java 23 and beyond, which brings a wealth of upgrades and hopefully addresses the integration between virtual threads and locking primitives.</p><p id="30c0" class="pw-post-body-paragraph nh ni gt nj b hr nk nl nm hu nn no np nq nr ns nt nu nv nw nx ny nz oa ob oc gm bj">This exploration highlights just one type of issue that performance engineers solve at Netflix. We hope this glimpse into our problem-solving approach proves valuable to others in their future investigations.</p></div></div></div></div>]]></description>
      <link>https://netflixtechblog.com/java-21-virtual-threads-dude-wheres-my-lock-3052540e231d</link>
      <guid>https://netflixtechblog.com/java-21-virtual-threads-dude-wheres-my-lock-3052540e231d</guid>
      <pubDate>Mon, 29 Jul 2024 20:04:00 +0200</pubDate>
    </item>
    <item>
      <title><![CDATA[Maestro: Netflix’s Workflow Orchestrator]]></title>
      <description><![CDATA[<div><div></div><p id="0ec7" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj">By <a class="af nt" href="https://www.linkedin.com/in/jheua/" rel="noopener ugc nofollow" target="_blank">Jun He</a>, <a class="af nt" href="https://www.linkedin.com/in/natalliadzenisenka/" rel="noopener ugc nofollow" target="_blank">Natallia Dzenisenka</a>, <a class="af nt" href="https://www.linkedin.com/in/praneethy91/" rel="noopener ugc nofollow" target="_blank">Praneeth Yenugutala</a>, <a class="af nt" href="https://www.linkedin.com/in/yingyi-zhang-a0a164111/" rel="noopener ugc nofollow" target="_blank">Yingyi Zhang</a>, and <a class="af nt" href="https://www.linkedin.com/in/anjali-norwood-9521a16" rel="noopener ugc nofollow" target="_blank">Anjali Norwood</a></p><h1 id="4de7" class="nu nv gt be nw nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or bj">TL;DR</h1><p id="3080" class="pw-post-body-paragraph mv mw gt mx b my os na nb nc ot ne nf ng ou ni nj nk ov nm nn no ow nq nr ns gm bj">We are thrilled to announce that the Maestro source code is now open to the public! Please visit the <a class="af nt" href="https://github.com/Netflix/maestro" rel="noopener ugc nofollow" target="_blank">Maestro GitHub repository</a> to get started. If you find it useful, please <a class="af nt" href="https://github.com/Netflix/maestro" rel="noopener ugc nofollow" target="_blank">give us a star</a>.</p><h2 id="2bef" class="ox nv gt be nw oy oz dx oa pa pb dz oe ng pc pd pe nk pf pg ph no pi pj pk pl bj">What is Maestro</h2><p id="3756" class="pw-post-body-paragraph mv mw gt mx b my os na nb nc ot ne nf ng ou ni nj nk ov nm nn no ow nq nr ns gm bj">Maestro is a general-purpose, horizontally scalable workflow orchestrator designed to manage large-scale workflows such as data pipelines and machine learning model training pipelines. It oversees the entire lifecycle of a workflow, from start to finish, including retries, queuing, task distribution to compute engines, etc.. Users can package their business logic in various formats such as Docker images, notebooks, bash script, SQL, Python, and more. Unlike traditional workflow orchestrators that only support Directed Acyclic Graphs (DAGs), Maestro supports both acyclic and cyclic workflows and also includes multiple reusable patterns, including foreach loops, subworkflow, and conditional branch, etc.</p><h2 id="6c09" class="ox nv gt be nw oy oz dx oa pa pb dz oe ng pc pd pe nk pf pg ph no pi pj pk pl bj">Our Journey with Maestro</h2><p id="f334" class="pw-post-body-paragraph mv mw gt mx b my os na nb nc ot ne nf ng ou ni nj nk ov nm nn no ow nq nr ns gm bj">Since we first introduced Maestro in <a class="af nt" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/orchestrating-data-ml-workflows-at-scale-with-netflix-maestro-aaa2b41b800c">this blog post</a>, we have successfully migrated hundreds of thousands of workflows to it on behalf of users with minimal interruption. The transition was seamless, and Maestro has met our design goals by handling our ever-growing workloads. Over the past year, we’ve seen a remarkable 87.5% increase in executed jobs. Maestro now launches thousands of workflow instances and runs half a million jobs daily on average, and has completed around 2 million jobs on particularly busy days.</p><h2 id="fd3e" class="ox nv gt be nw oy oz dx oa pa pb dz oe ng pc pd pe nk pf pg ph no pi pj pk pl bj">Scalability and Versatility</h2><p id="5f06" class="pw-post-body-paragraph mv mw gt mx b my os na nb nc ot ne nf ng ou ni nj nk ov nm nn no ow nq nr ns gm bj">Maestro is a fully managed workflow orchestrator that provides Workflow-as-a-Service to thousands of end users, applications, and services at Netflix. It supports a wide range of workflow use cases, including ETL pipelines, ML workflows, AB test pipelines, pipelines to move data between different storages, etc. Maestro’s horizontal scalability ensures it can manage both a large number of workflows and a large number of jobs within a single workflow.</p><p id="c621" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj">At Netflix, workflows are intricately connected. Splitting them into smaller groups and managing them across different clusters adds unnecessary complexity and degrades the user experience. This approach also requires additional mechanisms to coordinate these fragmented workflows. Since Netflix’s data tables are housed in a single data warehouse, we believe a single orchestrator should handle all workflows accessing it.</p><p id="986e" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj">Join us on this exciting journey by exploring the <a class="af nt" href="https://github.com/Netflix/maestro" rel="noopener ugc nofollow" target="_blank">Maestro GitHub repository</a> and contributing to its ongoing development. Your support and feedback are invaluable as we continue to improve the Maestro project.</p><h1 id="7110" class="nu nv gt be nw nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or bj">Introducing Maestro</h1><p id="4941" class="pw-post-body-paragraph mv mw gt mx b my os na nb nc ot ne nf ng ou ni nj nk ov nm nn no ow nq nr ns gm bj">Netflix Maestro offers a comprehensive set of features designed to meet the diverse needs of both engineers and non-engineers. It includes the common functions and reusable patterns applicable to various use cases in a loosely coupled way.</p><p id="1921" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj">A workflow definition is defined in a JSON format. Maestro combines user-supplied fields with those managed by Maestro to form a flexible and powerful orchestration definition. An example can be found in the <a class="af nt" href="https://github.com/Netflix/maestro/wiki/Workflow-definition-example" rel="noopener ugc nofollow" target="_blank">Maestro repository wiki</a>.</p><p id="2ac0" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj">A Maestro workflow definition comprises two main sections: properties and versioned workflow including its metadata. Properties include author and owner information, and execution settings. Maestro preserves key properties across workflow versions, such as author and owner information, run strategy, and concurrency settings. This consistency simplifies management and aids in trouble-shootings. If the ownership of the current workflow changes, the new owner can claim the ownership of the workflows without creating a new workflow version. Users can also enable the triggering or alerting features for a given workflow over the properties.</p><p id="656f" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj">Versioned workflow includes attributes like a unique identifier, name, description, tags, timeout settings, and criticality levels (low, medium, high) for prioritization. Each workflow change creates a new version, enabling tracking and easy reversion, with the active or the latest version used by default. A workflow consists of steps, which are the nodes in the workflow graph defined by users. Steps can represent jobs, another workflow using subworkflow step, or a loop using foreach step. Steps consist of unique identifiers, step types, tags, input and output step parameters, step dependencies, retry policies, and failure mode, step outputs, etc. Maestro supports configurable retry policies based on error types to enhance step resilience.</p><p id="cf60" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj">This high-level overview of Netflix Maestro’s workflow definition and properties highlights its flexibility to define complex workflows. Next, we dive into some of the useful features in the following sections.</p><h2 id="5be6" class="ox nv gt be nw oy oz dx oa pa pb dz oe ng pc pd pe nk pf pg ph no pi pj pk pl bj">Workflow Run Strategy</h2><p id="ae40" class="pw-post-body-paragraph mv mw gt mx b my os na nb nc ot ne nf ng ou ni nj nk ov nm nn no ow nq nr ns gm bj">Users want to automate data pipelines while retaining control over the execution order. This is crucial when workflows cannot run in parallel or must halt current executions when new ones occur. Maestro uses predefined run strategies to decide whether a workflow instance should run or not. Here is the list of predefined run strategies Maestro offers.</p><p id="a617" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj"><strong class="mx gu">Sequential Run Strategy</strong><br />This is the default strategy used by maestro, which runs workflows one at a time based on a First-In-First-Out (FIFO) order. With this run strategy, Maestro runs workflows in the order they are triggered. Note that an execution does not depend on the previous states. Once a workflow instance reaches one of the terminal states, whether succeeded or not, Maestro will start the next one in the queue.</p><p id="9d17" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj"><strong class="mx gu">Strict Sequential Run Strategy<br /></strong>With this run strategy, Maestro will run workflows in the order they are triggered but block execution if there’s a blocking error in the workflow instance history. Newly triggered workflow instances are queued until the error is resolved by manually restarting the failed instances or marking the failed ones unblocked.</p><figure class="pp pq pr ps pt pu pm pn paragraph-image"><div role="button" tabindex="0" class="pv pw fi px bg py"><div class="pm pn po"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*NvdLiYWhhWb0tvL-%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*NvdLiYWhhWb0tvL-%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*NvdLiYWhhWb0tvL-%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*NvdLiYWhhWb0tvL-%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*NvdLiYWhhWb0tvL-%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*NvdLiYWhhWb0tvL-%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*NvdLiYWhhWb0tvL-%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*NvdLiYWhhWb0tvL- 640w, https://miro.medium.com/v2/resize:fit:720/0*NvdLiYWhhWb0tvL- 720w, https://miro.medium.com/v2/resize:fit:750/0*NvdLiYWhhWb0tvL- 750w, https://miro.medium.com/v2/resize:fit:786/0*NvdLiYWhhWb0tvL- 786w, https://miro.medium.com/v2/resize:fit:828/0*NvdLiYWhhWb0tvL- 828w, https://miro.medium.com/v2/resize:fit:1100/0*NvdLiYWhhWb0tvL- 1100w, https://miro.medium.com/v2/resize:fit:1400/0*NvdLiYWhhWb0tvL- 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /></picture></div></div></figure><p id="8320" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj">In the above example, run5 fails at 5AM, then later runs are queued but do not run. When someone manually marks run5 unblocked or restarts it, then the workflow execution will resume. This run strategy is useful for time insensitive but business critical workflows. This gives the workflow owners the option to review the failures at a later time and unblock the executions after verifying the correctness.</p><p id="b23e" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj"><strong class="mx gu">First-only Run Strategy</strong><br />With this run strategy, Maestro ensures that the running workflow is complete before queueing a new workflow instance. If a new workflow instance is queued while the current one is still running, Maestro will remove the queued instance. Maestro will execute a new workflow instance only if there is no workflow instance currently running, effectively turning off queuing with this run strategy. This approach helps to avoid idempotency issues by not queuing new workflow instances.</p><p id="eaa2" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj"><strong class="mx gu">Last-only Run Strategy</strong><br />With this run strategy, Maestro ensures the running workflow is the latest triggered one and keeps only the last instance. If a new workflow instance is queued while there is an existing workflow instance already running, Maestro will stop the running instance and execute the newly triggered one. This is useful if a workflow is designed to always process the latest data, such as processing the latest snapshot of an entire table each time.</p><p id="7523" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj"><strong class="mx gu">Parallel with Concurrency Limit Run Strategy</strong><br />With this run strategy, Maestro will run multiple triggered workflow instances in parallel, constrained by a predefined concurrency limit. This helps to fan out and distribute the execution, enabling the processing of large amounts of data within the time limit. A common use case for this strategy is for backfilling the old data.</p><h2 id="3d04" class="ox nv gt be nw oy oz dx oa pa pb dz oe ng pc pd pe nk pf pg ph no pi pj pk pl bj">Parameters and Expression Language Support</h2><p id="7591" class="pw-post-body-paragraph mv mw gt mx b my os na nb nc ot ne nf ng ou ni nj nk ov nm nn no ow nq nr ns gm bj">In Maestro, parameters play an important role. Maestro supports dynamic parameters with code injection, which is super useful and powerful. This feature significantly enhances the flexibility and dynamism of workflows, allowing using parameters to control execution logic and enable state sharing between workflows and their steps, as well as between upstream and downstream steps. Together with other Maestro features, it makes the defining of workflows dynamic and enables users to define parameterized workflows for complex use cases.</p><p id="efd3" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj">However, code injection introduces significant security and safety concerns. For example, users might unintentionally write an infinite loop that creates an array and appends items to it, eventually crashing the server with out-of-memory (OOM) issues. While one approach could be to ask users to embed the injected code within their business logic instead of the workflow definition, this would impose additional work on users and tightly couple their business logic with the workflow. In certain cases, this approach blocks users to design some complex parameterized workflows.</p><p id="ef02" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj">To mitigate these risks and assist users to build parameterized workflows, we developed our own customized expression language parser, a simple, secure, and safe expression language (SEL). SEL supports code injection while incorporating validations during syntax tree parsing to protect the system. It leverages the Java Security Manager to restrict access, ensuring a secure and controlled environment for code execution.</p><p id="3a2f" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj"><strong class="mx gu">Simple, Secure, and Safe Expression Language (SEL)<br /></strong>SEL is a homemade simple, secure, and safe expression language (SEL) to address the risks associated with code injection within Maestro parameterized workflows. It is a simple expression language and the grammar and syntax follow JLS (<a class="af nt" href="https://docs.oracle.com/javase/specs/" rel="noopener ugc nofollow" target="_blank">Java Language Specifications</a>). SEL supports a subset of JLS, focusing on Maestro use cases. For example, it supports data types for all Maestro parameter types, raising errors, datetime handling, and many predefined utility methods. SEL also includes additional runtime checks, such as loop iteration limits, array size checks, object memory size limits and so on, to enhance security and reliability. For more details about SEL, please refer to the <a class="af nt" href="https://github.com/Netflix/maestro/blob/main/netflix-sel/docs/index.md#welcome-to-sel" rel="noopener ugc nofollow" target="_blank">Maestro GitHub documentation</a>.</p><p id="0162" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj"><strong class="mx gu">Output Parameters</strong><br />To further enhance parameter support, Maestro allows for callable step execution, which returns output parameters from user execution back to the system. The output data is transmitted to Maestro via its REST API, ensuring that the step runtime does not have direct access to the Maestro database. This approach significantly reduces security concerns.</p><p id="e293" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj"><strong class="mx gu">Parameterized Workflows</strong><br />Thanks to the powerful parameter support, users can easily create parameterized workflows in addition to static ones. Users enjoy defining parameterized workflows because they are easy to manage and troubleshoot while being powerful enough to solve complex use cases.</p><ul class=""><li id="ecc7" class="mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns qa qb qc bj">Static workflows are simple and easy to use but come with limitations. Often, users have to duplicate the same workflow multiple times to accommodate minor changes. Additionally, workflow and jobs cannot share the states without using parameters.</li><li id="26fe" class="mv mw gt mx b my qd na nb nc qe ne nf ng qf ni nj nk qg nm nn no qh nq nr ns qa qb qc bj">On the other hand, completely dynamic workflows can be challenging to manage and support. They are difficult to debug or troubleshoot and hard to be reused by others.</li><li id="2cc0" class="mv mw gt mx b my qd na nb nc qe ne nf ng qf ni nj nk qg nm nn no qh nq nr ns qa qb qc bj">Parameterized workflows strike a balance by being initialized step by step at runtime based on user defined parameters. This approach provides great flexibility for users to control the execution at runtime while remaining easy to manage and understand.</li></ul><p id="76e0" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj">As we described in <a class="af nt" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/orchestrating-data-ml-workflows-at-scale-with-netflix-maestro-aaa2b41b800c#360e">the previous Maestro blog post</a>, parameter support enables the creation of complex parameterized workflows, such as backfill data pipelines.</p><h2 id="03a6" class="ox nv gt be nw oy oz dx oa pa pb dz oe ng pc pd pe nk pf pg ph no pi pj pk pl bj">Workflow Execution Patterns</h2><p id="7289" class="pw-post-body-paragraph mv mw gt mx b my os na nb nc ot ne nf ng ou ni nj nk ov nm nn no ow nq nr ns gm bj">Maestro provides multiple useful building blocks that allow users to easily define dataflow patterns or other workflow patterns. It provides support for common patterns directly within the Maestro engine. Direct engine support not only enables us to optimize these patterns but also ensures a consistent approach to implementing them. Next, we will talk about the three major building blocks that Maestro provides.</p><p id="f47a" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj"><strong class="mx gu">Foreach Support</strong><br />In Maestro, the foreach pattern is modeled as a dedicated step within the original workflow definition. Each iteration of the foreach loop is internally treated as a separate workflow instance, which scales similarly as any other Maestro workflow based on the step executions (i.e. a sub-graph) defined within the foreach definition block. The execution of sub-graph within a foreach step is delegated to a separate workflow instance. Foreach step then monitors and collects the status of these foreach workflow instances, each managing the execution of a single iteration. For more details, please refer to <a class="af nt" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/orchestrating-data-ml-workflows-at-scale-with-netflix-maestro-aaa2b41b800c#360e">our previous Maestro blog post</a>.</p><p id="d95e" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj">The foreach pattern is frequently used to repeatedly run the same jobs with different parameters, such as data backfilling or machine learning model tuning. It would be tedious and time consuming to request users to explicitly define each iteration in the workflow definition (potentially hundreds of thousands of iterations). Additionally, users would need to create new workflows if the foreach range changes, further complicating the process.</p><p id="5686" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj"><strong class="mx gu">Conditional Branch Support</strong><br />The conditional branch feature allows subsequent steps to run only if specific conditions in the upstream step are met. These conditions are defined using the SEL expression language, which is evaluated at runtime. Combined with other building blocks, users can build powerful workflows, e.g. doing some remediation if the audit check step fails and then run the job again.</p><p id="9e17" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj"><strong class="mx gu">Subworkflow Support<br /></strong>The subworkflow feature allows a workflow step to run another workflow, enabling the sharing of common functions across multiple workflows. This effectively enables “workflow as a function” and allows users to build a graph of workflows. For example, we have observed complex workflows consisting of hundreds of subworkflows to process data across hundreds tables, where subworkflows are provided by multiple teams.</p><p id="7d57" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj">These patterns can be combined together to build composite patterns for complex workflow use cases. For instance, we can loop over a set of subworkflows or run nested foreach loops. One example that Maestro users developed is an auto-recovery workflow that utilizes both conditional branch and subworkflow features to handle errors and retry jobs automatically.</p><figure class="pp pq pr ps pt pu pm pn paragraph-image"><div role="button" tabindex="0" class="pv pw fi px bg py"><div class="pm pn po"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*d7XTqfPjAkuCBv6C%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*d7XTqfPjAkuCBv6C%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*d7XTqfPjAkuCBv6C%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*d7XTqfPjAkuCBv6C%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*d7XTqfPjAkuCBv6C%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*d7XTqfPjAkuCBv6C%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*d7XTqfPjAkuCBv6C%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*d7XTqfPjAkuCBv6C 640w, https://miro.medium.com/v2/resize:fit:720/0*d7XTqfPjAkuCBv6C 720w, https://miro.medium.com/v2/resize:fit:750/0*d7XTqfPjAkuCBv6C 750w, https://miro.medium.com/v2/resize:fit:786/0*d7XTqfPjAkuCBv6C 786w, https://miro.medium.com/v2/resize:fit:828/0*d7XTqfPjAkuCBv6C 828w, https://miro.medium.com/v2/resize:fit:1100/0*d7XTqfPjAkuCBv6C 1100w, https://miro.medium.com/v2/resize:fit:1400/0*d7XTqfPjAkuCBv6C 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /></picture></div></div></figure><p id="add9" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj">In this example, subworkflow `job1` runs another workflow consisting of extract-transform-load (ETL) and audit jobs. Next, a status check job leverages the Maestro parameter and SEL support to retrieve the status of the previous job. Based on this status, it can decide whether to complete the workflow or to run a recovery job to address any data issues. After resolving the issue, it then executes subworkflow `job2`, which runs the same workflow as subworkflow `job1`.</p><h2 id="dd55" class="ox nv gt be nw oy oz dx oa pa pb dz oe ng pc pd pe nk pf pg ph no pi pj pk pl bj">Step Runtime and Step Parameter</h2><p id="8600" class="pw-post-body-paragraph mv mw gt mx b my os na nb nc ot ne nf ng ou ni nj nk ov nm nn no ow nq nr ns gm bj"><strong class="mx gu">Step Runtime Interface<br /></strong>In Maestro, we use step runtime to describe a job at execution time. The step runtime interface defines two pieces of information:</p><ol class=""><li id="2de4" class="mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns qi qb qc bj">A set of basic APIs to control the behavior of a step instance at execution runtime.</li><li id="7fbd" class="mv mw gt mx b my qd na nb nc qe ne nf ng qf ni nj nk qg nm nn no qh nq nr ns qi qb qc bj">Some simple data structures to track step runtime state and execution result.</li></ol><p id="e253" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj">Maestro offers a few step runtime implementations such as foreach step runtime, subworkflow step runtime (mentioned in previous section). Each implementation defines its own logic for start, execute and terminate operations. At runtime, these operations control the way to initialize a step instance, perform the business logic and terminate the execution under certain conditions (i.e. manual intervention by users).</p><p id="17d4" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj">Also, Maestro step runtime internally keeps track of runtime state as well as the execution result of the step. The runtime state is used to determine the next state transition of the step and tell if it has failed or terminated. The execution result hosts both step artifacts and the timeline of step execution history, which are accessible by subsequent steps.</p><p id="ddb1" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj"><strong class="mx gu">Step Parameter Merging<br /></strong>To control step behavior in a dynamic way, Maestro supports both runtime parameters and tags injection in step runtime. This makes a Maestro step more flexible to absorb runtime changes (i.e. overridden parameters) before actually being started. Maestro internally maintains a step parameter map that is initially empty and is updated by merging step parameters in the order below:</p><ul class=""><li id="8853" class="mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns qa qb qc bj"><strong class="mx gu">Default General Parameters</strong>: Parameters merging starts from default parameters that in general every step should have. For example, workflow_instance_id, step_instance_uuid, step_attempt_id and step_id are required parameters for each maestro step. They are internally reserved by maestro and cannot be passed by users.</li><li id="4d0d" class="mv mw gt mx b my qd na nb nc qe ne nf ng qf ni nj nk qg nm nn no qh nq nr ns qa qb qc bj"><strong class="mx gu">Injected Parameters</strong>: Maestro then merges injected parameters (if present) into the parameter map. The injected parameters come from step runtime, which are dynamically generated based on step schema. Each type of step can have its own schema with specific parameters associated with this step. The step schema can evolve independently with no need to update Maestro code.</li><li id="54c3" class="mv mw gt mx b my qd na nb nc qe ne nf ng qf ni nj nk qg nm nn no qh nq nr ns qa qb qc bj"><strong class="mx gu">Default Typed Parameters</strong>: After injecting runtime parameters, Maestro tries to merge default parameters that are related to a specific type of step. For example, foreach step has loop_params and loop_index default parameters which are internally set by maestro and used for foreach step only.</li><li id="0a3b" class="mv mw gt mx b my qd na nb nc qe ne nf ng qf ni nj nk qg nm nn no qh nq nr ns qa qb qc bj"><strong class="mx gu">Workflow and Step Info Parameters</strong>: These parameters contain information about step and the workflow it belongs to. This can be identity information, i.e. workflow_id and will be merged to step parameter map if present.</li><li id="0292" class="mv mw gt mx b my qd na nb nc qe ne nf ng qf ni nj nk qg nm nn no qh nq nr ns qa qb qc bj"><strong class="mx gu">Undefined New Parameters</strong>: When starting or restarting a maestro workflow instance, users can specify new step parameters that are not present in initial step definition. ParamsManager merges these parameters to ensure they are available at execution time.</li><li id="d548" class="mv mw gt mx b my qd na nb nc qe ne nf ng qf ni nj nk qg nm nn no qh nq nr ns qa qb qc bj"><strong class="mx gu">Step Definition Parameters</strong>: These step parameters are defined by users at definition time and get merged if they are not empty.</li><li id="3f46" class="mv mw gt mx b my qd na nb nc qe ne nf ng qf ni nj nk qg nm nn no qh nq nr ns qa qb qc bj"><strong class="mx gu">Run and Restart Parameters</strong>: When starting or restarting a maestro workflow instance, users can override defined parameters by providing run or restart parameters. These two types of parameters are merged at the end so that step runtime can see the most recent and accurate parameter space.</li></ul><p id="7dbf" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj">The parameters merging logic can be visualized in the diagram below.</p><figure class="pp pq pr ps pt pu pm pn paragraph-image"><div role="button" tabindex="0" class="pv pw fi px bg py"><div class="pm pn po"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*bARelX8reZTdmFgr%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*bARelX8reZTdmFgr%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*bARelX8reZTdmFgr%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*bARelX8reZTdmFgr%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*bARelX8reZTdmFgr%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*bARelX8reZTdmFgr%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*bARelX8reZTdmFgr%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*bARelX8reZTdmFgr 640w, https://miro.medium.com/v2/resize:fit:720/0*bARelX8reZTdmFgr 720w, https://miro.medium.com/v2/resize:fit:750/0*bARelX8reZTdmFgr 750w, https://miro.medium.com/v2/resize:fit:786/0*bARelX8reZTdmFgr 786w, https://miro.medium.com/v2/resize:fit:828/0*bARelX8reZTdmFgr 828w, https://miro.medium.com/v2/resize:fit:1100/0*bARelX8reZTdmFgr 1100w, https://miro.medium.com/v2/resize:fit:1400/0*bARelX8reZTdmFgr 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /></picture></div></div></figure><h2 id="d5a8" class="ox nv gt be nw oy oz dx oa pa pb dz oe ng pc pd pe nk pf pg ph no pi pj pk pl bj">Step Dependencies and Signals</h2><p id="9493" class="pw-post-body-paragraph mv mw gt mx b my os na nb nc ot ne nf ng ou ni nj nk ov nm nn no ow nq nr ns gm bj">Steps in the Maestro execution workflow graph can express execution dependencies using step dependencies. A step dependency specifies the data-related conditions required by a step to start execution. These conditions are usually defined based on signals, which are pieces of messages carrying information such as parameter values and can be published through step outputs or external systems like SNS or Kafka messages.</p><p id="5a67" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj">Signals in Maestro serve both signal trigger pattern and signal dependencies (a publisher-subscriber) pattern. One step can publish an output signal (<a class="af nt" href="https://github.com/Netflix/maestro/blob/main/maestro-common/src/testFixtures/resources/fixtures/instances/sample-step-instance-failed.json#L151-L215" rel="noopener ugc nofollow" target="_blank">a sample example</a>) that can unblock the execution of multiple other steps that depend on it. A <a class="af nt" href="https://github.com/Netflix/maestro/blob/main/maestro-common/src/main/java/com/netflix/maestro/models/definition/SignalOutputsDefinition.java" rel="noopener ugc nofollow" target="_blank">signal definition</a> includes a list of mapped parameters, allowing Maestro to perform “signal matching” on a subset of fields. Additionally, Maestro supports <a class="af nt" href="https://github.com/Netflix/maestro/blob/main/maestro-common/src/main/java/com/netflix/maestro/models/parameter/SignalOperator.java" rel="noopener ugc nofollow" target="_blank">signal operators</a> like &lt;, &gt;, etc., on signal parameter values.</p><p id="652d" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj">Netflix has built various abstractions on top of the concept of signals. For instance, a ETL workflow can update a table with data and send signals that unblock steps in downstream workflows dependent on that data. Maestro supports “signal lineage,” which allows users to navigate all historical instances of signals and the workflow steps that match (i.e. publishing or consuming) those signals. Signal triggering guarantees exactly-once execution for the workflow subscribing a signal or a set of joined signals. This approach is efficient, as it conserves resources by only executing the workflow or step when the specified conditions in the signals are met. A signal service is implemented for those advanced abstractions. Please refer to the <a class="af nt" rel="noopener ugc nofollow" target="_blank" href="https://netflixtechblog.com/orchestrating-data-ml-workflows-at-scale-with-netflix-maestro-aaa2b41b800c#1fdf">Maestro blog</a> for further details on it.</p><h2 id="28e5" class="ox nv gt be nw oy oz dx oa pa pb dz oe ng pc pd pe nk pf pg ph no pi pj pk pl bj">Breakpoint</h2><p id="8764" class="pw-post-body-paragraph mv mw gt mx b my os na nb nc ot ne nf ng ou ni nj nk ov nm nn no ow nq nr ns gm bj">Maestro allows users to set breakpoints on workflow steps, functioning similarly to code-level breakpoints in an IDE. When a workflow instance executes and reaches a step with a breakpoint, that step enters a “paused” state. This halts the workflow graph’s progression until a user manually resumes from the breakpoint. If multiple instances of a workflow step are paused at a breakpoint, resuming one instance will only affect that specific instance, leaving the others in a paused state. Deleting the breakpoint will cause all paused step instances to resume.</p><p id="cace" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj">This feature is particularly useful during the initial development of a workflow, allowing users to inspect step executions and output data. It is also beneficial when running a step multiple times in a “foreach” pattern with various input parameters. Setting a single breakpoint on a step will cause all iterations of the foreach loop to pause at that step for debugging purposes. Additionally, the breakpoint feature allows human intervention during the workflow execution and can also be used for other purposes, e.g. supporting mutating step states while the workflow is running.</p><h2 id="7ab1" class="ox nv gt be nw oy oz dx oa pa pb dz oe ng pc pd pe nk pf pg ph no pi pj pk pl bj">Timeline</h2><p id="ed4f" class="pw-post-body-paragraph mv mw gt mx b my os na nb nc ot ne nf ng ou ni nj nk ov nm nn no ow nq nr ns gm bj">Maestro includes a step execution timeline, capturing all significant events such as execution state machine changes and the reasoning behind them. This feature is useful for debugging, providing insights into the status of a step. For example, it logs transitions such as “Created” and “Evaluating params”, etc. An example of a timeline is included <a class="af nt" href="https://github.com/Netflix/maestro/blob/main/maestro-common/src/testFixtures/resources/fixtures/instances/sample-step-instance-failed.json#L137-L150" rel="noopener ugc nofollow" target="_blank">here</a> for reference. The implemented step runtimes can add the timeline events into the timeline to surface the execution information to the end users.</p><h2 id="461d" class="ox nv gt be nw oy oz dx oa pa pb dz oe ng pc pd pe nk pf pg ph no pi pj pk pl bj">Retry Policies</h2><p id="33ba" class="pw-post-body-paragraph mv mw gt mx b my os na nb nc ot ne nf ng ou ni nj nk ov nm nn no ow nq nr ns gm bj">Maestro supports retry policies for steps that reach a terminal state due to failure. Users can specify the number of retries and configure retry policies, including delays between retries and exponential backoff strategies, in addition to fixed interval retries. Maestro distinguishes between two types of retries: “platform” and “user.” Platform retries address platform-level errors unrelated to user logic, while user retries are for user-defined conditions. Each type can have its own set of retry policies.</p><p id="46fa" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj">Automatic retries are beneficial for handling transient errors that can be resolved without user intervention. Maestro provides the flexibility to set retries to zero for non-idempotent steps to avoid retry. This feature ensures that users have control over how retries are managed based on their specific requirements.</p><h2 id="ef2f" class="ox nv gt be nw oy oz dx oa pa pb dz oe ng pc pd pe nk pf pg ph no pi pj pk pl bj">Aggregated View</h2><p id="4c03" class="pw-post-body-paragraph mv mw gt mx b my os na nb nc ot ne nf ng ou ni nj nk ov nm nn no ow nq nr ns gm bj">Because a workflow instance can have multiple runs, it is important for users to see an aggregated state of all steps in the workflow instance. Aggregated view is computed by merging base aggregated view with current runs instance step statuses. For example, as you can see on the figure below simulating a simple case, there is a first run, where step1 and step2 succeeded, step3 failed, and step4 and step5 have not started. When the user restarts the run, the run starts from step3 in run 2 with step1 and step2 skipped which succeeded in the previous run. After all steps succeed, the aggregated view shows the run states for all steps.</p><figure class="pp pq pr ps pt pu pm pn paragraph-image"><div role="button" tabindex="0" class="pv pw fi px bg py"><div class="pm pn qj"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*UC1Sj36z5IfvDz9X%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*UC1Sj36z5IfvDz9X%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*UC1Sj36z5IfvDz9X%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*UC1Sj36z5IfvDz9X%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*UC1Sj36z5IfvDz9X%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*UC1Sj36z5IfvDz9X%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*UC1Sj36z5IfvDz9X%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*UC1Sj36z5IfvDz9X 640w, https://miro.medium.com/v2/resize:fit:720/0*UC1Sj36z5IfvDz9X 720w, https://miro.medium.com/v2/resize:fit:750/0*UC1Sj36z5IfvDz9X 750w, https://miro.medium.com/v2/resize:fit:786/0*UC1Sj36z5IfvDz9X 786w, https://miro.medium.com/v2/resize:fit:828/0*UC1Sj36z5IfvDz9X 828w, https://miro.medium.com/v2/resize:fit:1100/0*UC1Sj36z5IfvDz9X 1100w, https://miro.medium.com/v2/resize:fit:1400/0*UC1Sj36z5IfvDz9X 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /></picture></div></div></figure><h2 id="0eeb" class="ox nv gt be nw oy oz dx oa pa pb dz oe ng pc pd pe nk pf pg ph no pi pj pk pl bj">Rollup</h2><p id="6431" class="pw-post-body-paragraph mv mw gt mx b my os na nb nc ot ne nf ng ou ni nj nk ov nm nn no ow nq nr ns gm bj">Rollup provides a high-level summary of a workflow instance, detailing the status of each step and the count of steps in each status. It flattens steps across the current instance and any nested non-inline workflows like subworkflows or foreach steps. For instance, if a successful workflow has three steps, one of which is a subworkflow corresponding to a five-step workflow, the rollup will indicate that seven steps succeeded. Only leaf steps are counted in the rollup, as other steps serve merely as pointers to concrete workflows.</p><p id="4d8c" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj">Rollup also retains references to any non-successful steps, offering a clear overview of step statuses and facilitating easy navigation to problematic steps, even within nested workflows. The aggregated rollup for a workflow instance is calculated by combining the current run’s runtime data with a base rollup. The current state is derived from the statuses of active steps, including aggregated rollups for foreach and subworkflow steps. The base rollup is established when the workflow instance begins and includes statuses of inline steps (excluding foreach and subworkflows) from the previous run that are not part of the current run.</p><p id="957b" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj">For subworkflow steps, the rollup simply reflects the rollup of the subworkflow instance. For foreach steps, the rollup combines the base rollup of the foreach step with the current state rollup. The base is derived from the previous run’s aggregated rollup, excluding the iterations to be restarted in the new run. The current state is periodically updated by aggregating rollups of running iterations until all iterations reach a terminal state.</p><p id="2610" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj">Due to these processes, the rollup model is eventually consistent. While the figure below illustrates a straightforward example of rollup, the calculations can become complex and recursive, especially with multiple levels of nested foreaches and subworkflows.</p><figure class="pp pq pr ps pt pu pm pn paragraph-image"><div role="button" tabindex="0" class="pv pw fi px bg py"><div class="pm pn po"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*ISib6wPtCLtAbOuU%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*ISib6wPtCLtAbOuU%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*ISib6wPtCLtAbOuU%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*ISib6wPtCLtAbOuU%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*ISib6wPtCLtAbOuU%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*ISib6wPtCLtAbOuU%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*ISib6wPtCLtAbOuU%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*ISib6wPtCLtAbOuU 640w, https://miro.medium.com/v2/resize:fit:720/0*ISib6wPtCLtAbOuU 720w, https://miro.medium.com/v2/resize:fit:750/0*ISib6wPtCLtAbOuU 750w, https://miro.medium.com/v2/resize:fit:786/0*ISib6wPtCLtAbOuU 786w, https://miro.medium.com/v2/resize:fit:828/0*ISib6wPtCLtAbOuU 828w, https://miro.medium.com/v2/resize:fit:1100/0*ISib6wPtCLtAbOuU 1100w, https://miro.medium.com/v2/resize:fit:1400/0*ISib6wPtCLtAbOuU 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /></picture></div></div></figure><h2 id="fb0a" class="ox nv gt be nw oy oz dx oa pa pb dz oe ng pc pd pe nk pf pg ph no pi pj pk pl bj">Maestro Event Publishing</h2><p id="8676" class="pw-post-body-paragraph mv mw gt mx b my os na nb nc ot ne nf ng ou ni nj nk ov nm nn no ow nq nr ns gm bj">When workflow definition, workflow instance or step instance is changed, Maestro generates an event, processes it internally and publishes the processed event to external system(s). Maestro has both internal and external events. The internal event tracks changes within the life cycle of workflow, workflow instance or step instance. It is published to an internal queue and processed within Maestro. After internal events are processed, some of them will be transformed into external event and sent out to the external queue (i.e. SNS, Kafka). The external event carries maestro status change information for downstream services. The event publishing flow is illustrated in the diagram below:</p><figure class="pp pq pr ps pt pu pm pn paragraph-image"><div role="button" tabindex="0" class="pv pw fi px bg py"><div class="pm pn po"><picture><img src="https://miro.medium.com/v2/resize:fit:640/format:webp/0*n2Kiea-ngDjnKppJ%20640w,%20https://miro.medium.com/v2/resize:fit:720/format:webp/0*n2Kiea-ngDjnKppJ%20720w,%20https://miro.medium.com/v2/resize:fit:750/format:webp/0*n2Kiea-ngDjnKppJ%20750w,%20https://miro.medium.com/v2/resize:fit:786/format:webp/0*n2Kiea-ngDjnKppJ%20786w,%20https://miro.medium.com/v2/resize:fit:828/format:webp/0*n2Kiea-ngDjnKppJ%20828w,%20https://miro.medium.com/v2/resize:fit:1100/format:webp/0*n2Kiea-ngDjnKppJ%201100w,%20https://miro.medium.com/v2/resize:fit:1400/format:webp/0*n2Kiea-ngDjnKppJ%201400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" alt="image" /><source data-testid="og" srcset="https://miro.medium.com/v2/resize:fit:640/0*n2Kiea-ngDjnKppJ 640w, https://miro.medium.com/v2/resize:fit:720/0*n2Kiea-ngDjnKppJ 720w, https://miro.medium.com/v2/resize:fit:750/0*n2Kiea-ngDjnKppJ 750w, https://miro.medium.com/v2/resize:fit:786/0*n2Kiea-ngDjnKppJ 786w, https://miro.medium.com/v2/resize:fit:828/0*n2Kiea-ngDjnKppJ 828w, https://miro.medium.com/v2/resize:fit:1100/0*n2Kiea-ngDjnKppJ 1100w, https://miro.medium.com/v2/resize:fit:1400/0*n2Kiea-ngDjnKppJ 1400w" sizes="(min-resolution: 4dppx) and (max-width: 700px) 50vw, (-webkit-min-device-pixel-ratio: 4) and (max-width: 700px) 50vw, (min-resolution: 3dppx) and (max-width: 700px) 67vw, (-webkit-min-device-pixel-ratio: 3) and (max-width: 700px) 65vw, (min-resolution: 2.5dppx) and (max-width: 700px) 80vw, (-webkit-min-device-pixel-ratio: 2.5) and (max-width: 700px) 80vw, (min-resolution: 2dppx) and (max-width: 700px) 100vw, (-webkit-min-device-pixel-ratio: 2) and (max-width: 700px) 100vw, 700px" /></picture></div></div></figure><p id="8ae3" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj">As shown in the diagram, the Maestro event processor bridges the two aforementioned Maestro events. It listens on the internal queue to get the published <a class="af nt" href="https://github.com/Netflix/maestro/tree/main/maestro-engine/src/main/java/com/netflix/maestro/engine/jobevents" rel="noopener ugc nofollow" target="_blank">internal events</a>. Within the processor, the internal job event is processed based on its type and gets converted to an <a class="af nt" href="https://github.com/Netflix/maestro/tree/main/maestro-common/src/main/java/com/netflix/maestro/models/events" rel="noopener ugc nofollow" target="_blank">external event</a> if needed. The notification publisher at the end emits the external event so that downstream services can consume.</p><p id="da33" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj">The downstream services are mostly event-driven. The Maestro event carries the most useful message for downstream services to capture different changes in Maestro. In general, these changes can be classified into two categories: workflow change and instance status change. The workflow change event is associated with actions at workflow level, i.e definition or properties of a workflow has changed. Meanwhile, instance status change tracks status transition on workflow instance or step instance.</p><h1 id="9ed3" class="nu nv gt be nw nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or bj">Get Started with Maestro</h1><p id="5db4" class="pw-post-body-paragraph mv mw gt mx b my os na nb nc ot ne nf ng ou ni nj nk ov nm nn no ow nq nr ns gm bj">Maestro has been extensively used within Netflix, and today, we are excited to make the Maestro source code publicly available. We hope that the scalability and usability that Maestro offers can expedite workflow development outside Netflix. We invite you to try Maestro, use it within your organization, and contribute to its development.</p><p id="7494" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj">You can find the Maestro code repository at <a class="af nt" href="https://github.com/Netflix/maestro" rel="noopener ugc nofollow" target="_blank">github.com/Netflix/maestro</a>. If you have any questions, thoughts, or comments about Maestro, please feel free to create a <a class="af nt" href="https://github.com/Netflix/maestro/issues" rel="noopener ugc nofollow" target="_blank">GitHub issue</a> in the Maestro repository. We are eager to hear from you.</p><p id="f320" class="pw-post-body-paragraph mv mw gt mx b my mz na nb nc nd ne nf ng nh ni nj nk nl nm nn no np nq nr ns gm bj">We are taking workflow orchestration to the next level and constantly solving new problems and challenges, please stay tuned for updates. If you are passionate about solving large scale orchestration problems, please <a class="af nt" href="https://jobs.netflix.com/search?team=Data+Platform" rel="noopener ugc nofollow" target="_blank">join us</a>.</p><h1 id="5fe2" class="nu nv gt be nw nx ny nz oa ob oc od oe of og oh oi oj ok ol om on oo op oq or bj">Acknowledgements</h1><p id="d8df" class="pw-post-body-paragraph mv mw gt mx b my os na nb nc ot ne nf ng ou ni nj nk ov nm nn no ow nq nr ns gm bj">Thanks to other Maestro team members, <a class="af nt" href="https://www.linkedin.com/in/binbing-hou/" rel="noopener ugc nofollow" target="_blank">Binbing Hou</a>, <a class="af nt" href="http://linkedin.com/in/zhuoran-d-96848b154" rel="noopener ugc nofollow" target="_blank">Zhuoran Dong</a>, <a class="af nt" href="https://www.linkedin.com/in/brittany-truong-a35b54bb" rel="noopener ugc nofollow" target="_blank">Brittany Truong</a>, <a class="af nt" href="https://www.linkedin.com/in/rdeepak2002/" rel="noopener ugc nofollow" target="_blank">Deepak Ramalingam</a>, <a class="af nt" href="http://linkedin.com/in/moctarba" rel="noopener ugc nofollow" target="_blank">Moctar Ba</a>, for their contributions to the Maestro project. Thanks to our Product Manager <a class="af nt" href="https://www.linkedin.com/in/ashpokh/" rel="noopener ugc nofollow" target="_blank">Ashim Pokharel</a> for driving the strategy and requirements. We’d also like to thank <a class="af nt" href="https://www.linkedin.com/in/andrew-seier/" rel="noopener ugc nofollow" target="_blank">Andrew Seier</a>, <a class="af nt" href="https://www.linkedin.com/in/romain-cledat-4a211a5" rel="noopener ugc nofollow" target="_blank">Romain Cledat</a>, <a class="af nt" href="https://www.linkedin.com/in/agorajek/" rel="noopener ugc nofollow" target="_blank">Olek Gorajek</a>, and other stunning colleagues at Netflix for their contributions to the Maestro project. We also thank Prashanth Ramdas, Eva Tse, David Noor, Charles Smith and other leaders of Netflix engineering organizations for their constructive feedback and suggestions on the Maestro project.</p></div>]]></description>
      <link>https://netflixtechblog.com/maestro-netflixs-workflow-orchestrator-ee13a06f9c78</link>
      <guid>https://netflixtechblog.com/maestro-netflixs-workflow-orchestrator-ee13a06f9c78</guid>
      <pubDate>Mon, 22 Jul 2024 19:38:00 +0200</pubDate>
    </item>
  </channel>
</rss>
