Sr. Content Developer at Microsoft, working remotely in PA, TechBash conference organizer, former Microsoft MVP, Husband, Dad and Geek.
159809 stories
·
33 followers

Product Owners Earn Team Loyalty By Showing Up, Learning, And Deciding | Joshua McDonald

1 Share

Joshua McDonald: Product Owners Earn Team Loyalty By Showing Up, Learning, And Deciding

The Great Product Owner: Vulnerable Enough To Learn The Technical Details

Read the full Show Notes and search through the world's largest audio library on Agile and Scrum directly on the Scrum Master Toolbox Podcast website: http://bit.ly/SMTP_ShowNotes.

 

"I've seen developers be fiercely loyal to their product owner who makes that type of effort." - Joshua McDonald

 

Joshua describes the great Product Owner as someone who connects with the team even without a technical background. They do not pretend to know everything. Instead, they get into refinement with curiosity, write stories with the developers, and say, "I think this is what needs to be done, but correct me if I'm wrong." That vulnerability creates a strong partnership. Developers see the effort and often respond with loyalty, to the point where they may refuse to continue important discussions without the Product Owner in the room. The lesson is useful for Scrum Masters coaching Product Owners: deep technical expertise is not the entry ticket. Presence, curiosity, and the willingness to learn with the team are what create trust.

 

Self-reflection Question: How does your Product Owner show the team they are willing to learn the product with them?

The Bad Product Owner: The Dependent Note Taker

Read the full Show Notes and search through the world's largest audio library on Agile and Scrum directly on the Scrum Master Toolbox Podcast website: http://bit.ly/SMTP_ShowNotes.

 

"Sometimes you ask the product owner, you're the bus driver, where do you want us to go? And they say, let me go talk with the product manager." - Joshua McDonald

 

The Product Owner anti-pattern Joshua highlights is the dependent PO: the person who shows up as a note taker rather than a decision maker. This often happens when a Product Manager sits behind the PO and every decision has to be checked elsewhere. The day-to-day symptoms are easy to spot: camera off, muted, disconnected, asking people to repeat questions, missing meetings they scheduled themselves, and leaving the team without timely answers. Joshua's coaching response starts with one-on-ones, support, and curiosity. He asks how they are doing, what support they need, and what signals would help the team see they are engaged. When the product feels too technical, he sits beside them, learns with them, and helps them build confidence instead of leaving them exposed.

 

Self-reflection Question: Where is your Product Owner dependent on someone else for decisions, and how is that affecting the team?

 

[The Scrum Master Toolbox Podcast Recommends]

🔥In the ruthless world of fintech, success isn't just about innovation—it's about coaching!🔥

Angela thought she was just there to coach a team. But now, she's caught in the middle of a corporate espionage drama that could make or break the future of digital banking. Can she help the team regain their mojo and outwit their rivals, or will the competition crush their ambitions? As alliances shift and the pressure builds, one thing becomes clear: this isn't just about the product—it's about the people.

 

🚨 Will Angela's coaching be enough? Find out in Shift: From Product to People—the gripping story of high-stakes innovation and corporate intrigue.

 

Buy Now on Amazon

 

[The Scrum Master Toolbox Podcast Recommends]

 

About Joshua McDonald

 

Joshua is endlessly curious about better ways of working. He helps teams grow by blending experimentation, creativity, and AI with practical coaching. When a meeting feels routine, he's already testing a new approach to make it more valuable. Energetic and inventive, he turns everyday collaboration into opportunities for team growth.

 

You can link with Joshua McDonald on LinkedIn.





Download audio: https://traffic.libsyn.com/secure/scrummastertoolbox/20260821_Joshua_McDonald_F.mp3?dest-id=246429
Read the whole story
alvinashcraft
just a second ago
reply
Pennsylvania, USA
Share this story
Delete

AGL 483: Nick Jonsson

1 Share

About Nick

Executive LonelinessNick Jonsson is widely recognized as the world’s leading authority on Executive Loneliness. His pioneering work has earned him a dedicated Executive Loneliness Wikipedia page and international recognition, establishing him as one of the world’s most respected voices on leadership, vulnerability, mental wellbeing, and authentic connection.

Ranked #28 in the world among the Global Gurus Top 30 Coaching Professionals, Nick is a TEDx speaker, international bestselling author of Executive Loneliness, Professional Certified Coach (PCC), Certified Master Coach (CMC), qualified Counsellor and Psychotherapist, and a certified Team Psychological Safety practitioner.

After more than two decades in international business leadership, Nick experienced burnout, addiction, and profound loneliness while leading from the top. His recovery transformed both his life and his purpose. Today, everything he teaches is grounded in lived experience as well as professional expertise.

Nick works globally with leaders, teams, and organizations through keynote speaking, executive coaching, counselling and psychotherapy, leadership facilitation, and Team Psychological Safety training. He is also a top 2% Ironman triathlete, founder of a men’s support group, suicide prevention volunteer, and host of The Limitless Podcast, featuring conversations with leaders and changemakers from around the world.

His mission is to empower people to build a limitless life through vulnerability and holistic wellbeing.


Today We Talked About

  • Executive Loneliness
  • Asking for help
  • Rock Bottom
  • Talking to others
  • Physical training
  • Relationships
  • Honest converesations
  • Relationships replaced alcohol
  • Share how you feel
  • Fear
  • Don’t overshare

Connect with Nick


Leave me a tip $
Click here to Donate to the show


I hope you enjoyed this show, please head over to Apple Podcasts and subscribe and leave me a rating and review, even one sentence will help spread the word.  Thanks again!





Download audio: https://media.blubrry.com/a_geek_leader_podcast__/mc.blubrry.com/a_geek_leader_podcast__/AGL_483_Nick_Jonsson.mp3?awCollectionId=300549&awEpisodeId=12186429&aw_0_azn.pgenre=Business&aw_0_1st.ri=blubrry&aw_0_azn.pcountry=US&aw_0_azn.planguage=en&cat_exclude=IAB1-8%2CIAB1-9%2CIAB7-41%2CIAB8-5%2CIAB8-18%2CIAB11-4%2CIAB25%2CIAB26&aw_0_cnt.rss=https%3A%2F%2Fwww.ageekleader.com%2Ffeed%2Fpodcast
Read the whole story
alvinashcraft
just a second ago
reply
Pennsylvania, USA
Share this story
Delete

Announcing Azure Web PubSub chat in public preview

1 Share

Chat is becoming a standard interaction model across customer support, collaboration, gaming, marketplaces, healthcare, financial services, and AI-powered applications. But building a production-ready chat system involves much more than opening a WebSocket connection. Teams also need to manage rooms and membership, deliver and order messages, retain conversation history, recover from connection interruptions, and enforce permissions.

Today, we are excited to announce that Azure Web PubSub chat is available in public preview. Azure Web PubSub chat is a managed capability built on Azure Web PubSub. It provides chat-focused client and server APIs so developers can work directly with familiar concepts such as rooms, messages, members, users, and roles, while Azure manages the underlying real-time messaging infrastructure.

Focus on the chat experience, not the messaging plumbing

Azure Web PubSub already helps developers build large-scale, real-time applications. The new chat capability adds a higher-level abstraction for applications whose primary interaction model is conversation. Instead of defining custom events, message schemas, membership logic, persistence workflows, and reconnect behavior, developers can use built-in chat operations to:

  • Create one-to-one or group rooms.
  • Add and remove room members.
  • Send ordered messages in real time.
  • Retrieve persistent room message history.
  • Apply built-in or custom roles and permissions.
  • Reconnect clients and recover messages after temporary connection loss.

These capabilities run on Azure Web PubSub infrastructure and inherit its automatic scaling, geo-replication, security, and compliance foundations.

Chat-native APIs for clients and servers

Azure Web PubSub chat provides two complementary ways to build.

The JavaScript client SDK, available through the `@azure/web-pubsub-chat-client` npm package, connects applications to a chat hub. Clients can create rooms, exchange messages, load history, manage members, and subscribe to chat events.

The Chat REST API supports trusted server-side and administrative workflows, including managing rooms, users, members, messages, roles, and permissions. This separation lets client applications deliver responsive real-time experiences while backend services retain control over identity, moderation, governance, and business rules.

For example, after connecting a `ChatClient`, creating a room and sending a message takes only a few calls:

const room = await client.createRoom("Project Falcon", ["bob", "carol"]); client.on("message", ({ message }) => { console.log(`${message.createdBy}: ${message.content.text}`); }); await client.sendToRoom(room.roomId, "Welcome to the project room!");

Applications can also read message history through an asynchronous iterator, making it straightforward to implement initial conversation loading or incremental history as a user scrolls.

Keep chat data in your Azure Storage account

Persistent chat data remains in an Azure Storage account selected by the application owner. Azure Web PubSub chat uses the Web PubSub resource's managed identity to access that storage, avoiding storage connection strings or keys in the chat configuration.

The stored data includes:

  • Messages and conversation history
  • Rooms and room membership
  • Users
  • Roles and permissions

Your data stays in your storage account, and the service keeps no separate copy. This model gives organizations direct ownership of their persisted chat data while the managed service handles real-time delivery and chat operations.

Built-in access control with room-level flexibility

Chat applications often need different privileges for participants, moderators, room owners, support agents, or automated services. Azure Web PubSub chat includes a role and permission model for these scenarios.

Built-in room roles distinguish between members and operators. Both can send messages, read history, and invite members by default, while operators can also remove users. Developers can define custom user or room roles when an application needs a different permission set. Role administration is performed through the Chat REST API, keeping permission management in trusted server-side code.

Reliable conversations across connections and devices

Users expect chat to continue working when a laptop changes networks, a mobile connection briefly drops, or the same account is open in multiple browser tabs or devices. Azure Web PubSub chat runs on Azure Web PubSub's reliable WebSocket connection. The client reconnects and recovers automatically after an interruption, while the service fans messages out to a user's active connections. Persistent history also allows users to load earlier messages after reconnecting or joining a room later.

Choose the right Web PubSub capability

The new chat capability complements the existing standard Web PubSub hub. Use a chat hub when the application is centered on conversations and benefits from built-in rooms, membership, history, roles, and chat-focused APIs. Use a standard hub when the application needs full control over its protocol or supports a different real-time workload, such as telemetry, device signaling, multiplayer state, notifications, or live dashboards. Both options use Azure Web PubSub, allowing teams to select the abstraction that best fits each real-time scenario.

Get started

To try Azure Web PubSub chat:

  1. Create or open an Azure Web PubSub resource.
  2. Link an Azure Storage account under Persistent Storages.
  3. Add a Chat Hub and associate it with that storage.
  4. Generate a client access URL for testing.
  5. Install the JavaScript client SDK:
npm install azure/web-pubsub-chat-client
  1. Connect a client, create a room, and send your first message.

The portal-generated client URL is intended for experimentation. In production, issue client access URLs from an authenticated backend and derive the chat user ID from the signed-in application identity. Managed identity and Microsoft Entra ID can be used for keyless server authentication. 

Azure Web PubSub chat is available now in public preview. Start with the Azure Web PubSub chat overview, follow the quickstart, or explore the client SDK and REST API. We look forward to seeing the customer conversations, collaboration experiences, and AI-powered applications you build with it.

Read the whole story
alvinashcraft
16 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

Issue 764

1 Share

Comment

Hello, iOS Dev Weekly readers 👋 It’s me Konsta.

I have been feeling especially happy to be part of the iOS community lately. Maybe it is because, even in this AI-heavy moment, I still see Apple hiring iOS engineers and investing in the tools we use every day. I recently came across a few open roles on LinkedIn that seemed connected to Instruments and new AI/ML APIs.

That feels important because many people building iOS apps with AI agents seem to run into the same issue: performance, observability and scalability. AI can be very good at completing one task in isolation, but when everything comes together, it does not always understand the bigger picture of the app. It doesn’t really think like a staff-level engineer. That is where experience, critical thinking, and strong programming foundations still matter.

Whether you are leading a small project or making decisions across a bigger team, those calls still need a human point of view. That is also the thread I saw in this week’s issue. Some of the articles I am sharing today should help you make better decisions, while others should help you understand how things work under the hood.

– Konstantinos Nikoloutsos

You’ve been learning iOS with Kodeco for 16 years. Now for your entire team.

Kodeco has been helping iOS developers learn since 2010, and as the landscape changes, we’re moving with it—helping you make the most of AI, without losing the craft that you love. Our Apple Foundation Models book shipped in July, with an in-depth course on using AI as a development tool coming this autumn. All of Kodeco’s content is included in Kodeco for Teams, which comes out of a training budget rather than your pocket, and provides access for your whole team. We can show you round in 20 minutes. Or take us to your leader.

News

Apple announces changes for apps in the European Union

New EU terms go into effect on October 1 2026. Core Technology Fee is replaced by a 5% Core Technology Commission for apps distributed outside the App Store. My opinion is that this probably will not make sense for every app, and I do not expect everyone to suddenly leave the App Store, but it may make alternative distribution more realistic for developers who already have a direct relationship with their customers.


Looking for maintainers

Reading and contributing to Open-source projects was always a quick and fun way to increase my knowledge and meet with new people. I still remember myself learning about the protocol witness from pointfreeco snapshottesting library, an alternative way of using protocols. In this article Massicotte is sharing a list of open source libraries that are looking for maintainers and I think it is a great opportunity.

Code

ContentBuilder Explained - from 10.9s to 5.38ms in 11-level nested swiftUI

The compiler is unable to type-check this expression in reasonable time. Yes, we’ve all seen this error at least once in our career and always questioned why the compiler cannot be smarter. In this article Fatbobman is explaining how ContentBuilder works and then focuses on what Apple changed to make SwiftUI expressions easier for the compiler to handle. The interesting part for me is that this is not just a new public API, but a change in how SwiftUI can reduce the amount of generic type information the compiler has to reason about. If you have ever split a view into smaller pieces just to make the type checker happy, this one is worth reading.


The State macro in Xcode 27

Xcode 27 quietly rewrites @State as a macro. The performance win is real with no more rerunning your class init on every view re-instantiation. Apple bundled this change on the Xcode version and not your deployment target, so it can break your build even if you’re still targeting iOS 15. For example two patterns stop compiling: initializers that override a declared initial value, and extensions using the synthesized memberwise init. Worth a quick grep before upgrading as Blake Crosley suggests.


iOS 27: StateReporter

Observability is becoming one of the most important parts of building reliable apps. Since we are moving faster than before, especially with AI helping us write code, I think that also means we need to understand problems faster when something goes wrong. In this article Anton Gubarenko explains StateReporting, a new iOS 27 framework that lets you attach app state to diagnostics, so tools like MetricKit and Instruments can show what the user was doing when a performance issue happened. In my current job, we are also looking for ways to improve our observability for the same reason: we do not want to wait days to understand that something is not working as expected.

Tools

A framework to make decisions faster as a Lead Software Engineer

Taking the perfect decision everytime as a leader is what is expected, right? And then you end up over-thinking, what if this is not the best solution? and then you think again and finally you feel you should deliver faster to What if you can follow a framework to help you take decisions faster that you can later iterate on to make better. Mohammad Faani in this article shares a framework he uses on his decision process that I think every leader should take a look.


Using AI while exercising your critical thinking

Have you ever seen the expression “Don’t be a meat proxy”? In our day-to-day, we are spending most of our time trying to solve a problem by speaking with an agent. The dopamine that we are getting sometimes is measured by how many problems we have solved and how fast we are going. How many problems are we solving? How many features are we building? One thing that is really important here is that we should always challenge the agent and try to apply on top of what he is saying. This is what, in my opinion, differentiates a really good engineer from someone that just uses the tool. Bruno Rocha makes the case that AI-assisted development still needs an engineer in the loop. The useful framing here is that your value is not forwarding the answer, but judging whether the answer makes sense.


Trait-ifying our libraries to reduce transitive dependencies

Have you ever wanted to be able to depend on a different flavour of a library? Let’s say you want to depend on a monkay but all of a sudden you end up with the whole jungle.. There must be a way to configure the SPM package to decide, based on the caller’s input how to be structured. Introducing SPM traits, a way for the caller to choose the trait he wants to enable.

One usecase of that is used by Point-Free who they are using SwiftPM package traits to let users opt into only the pieces of their libraries they need, instead of pulling in every transitive dependency. I believe traits is a powerful SPM feature that gives a lot of flexibility to the authors while giving a lot of options to the consumers!


What is a swift package registry?

I still catch myself thinking of Swift packages as Git repositories with tags, so Dave’s explanation on package registries was a useful reset. In this article, Dave Verwer, long readers of this newsletter will already know him, is using the tuist registry to load an spm module and gives some hints why package registries work is needed. My favorite one was that registry-hosted packages ship as immutable, verified source archives.

And finally...

Foldable iPhone rumors keep getting louder.. Users may love the extra screen, but I suspect a few iOS engineers are already feeling nostalgic for the good old days when an iPhone was just one rectangle 😅 Really curious on how something like this can spawn new design ideas.

Read the whole story
alvinashcraft
25 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

Stop Making TUIs

1 Share

Stop Making TUIs

Thomas Ptacek advocates for building real native user interfaces for even the smallest of personal tools, because coding agents have reduced the cost of getting a usable-enough GUI up and running to almost nothing.

I wrote about my vibe-coded bandwidth and GPU monitoring macOS task bar apps back in March, and I'm still using both of those on a daily basis.

I'm not habitually knocking out real UIs for my other projects yet, but I'm running out of excuses!

Thomas:

If you haven’t tried your hand at turning one of your 500 throwaway CLIs into a native app, you’re doing yourself a disservice. Go build a native UI. It’ll probably change the way you think.

Tags: thomas-ptacek, ai, generative-ai, llms, vibe-coding, coding-agents

Read the whole story
alvinashcraft
51 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

A Tale of Two Flink Autoscalers

1 Share

Samuel Yeboah, Francesco Di Chiara and Mingliang Liu

Today, Netflix runs two Flink autoscalers. That is exactly one more than we want. We built the first one in-house years ago, when there was no mature option suited to our platform. The second came from the Apache Flink community, and it can scale workloads our homegrown system was never designed for. We now run both in production and are steadily converging on the open-source one. Along the way we learned some hard lessons about metrics, cost, and the real price of maintaining infrastructure you could instead adopt, and we hope they are useful whether you run a handful of Flink jobs or tens of thousands.

Why autoscaling is not optional at our scale

Netflix has run stream processing on Apache Flink since 2017. As of 2026 we operate more than 30,000 Flink jobs across multiple AWS regions. Most are not deployed by hand; they are generated by our managed platform Data Mesh, so the majority of users never touch a Flink job directly. A smaller but growing set are custom jobs, built and operated by teams across the company for use cases like personalization, Ads, and Live events. They range from single-operator jobs that shuttle records between Kafka topics to stateful pipelines with branches, joins, and terabytes of state, and their load swings with daily cycles, launches, and regional failovers.

Provisioning every one of those jobs for its peak is wasteful; provisioning for the average causes lag during surges. And in our platform a scaling action is not free: by default it means taking a savepoint, stopping the job gracefully, and restarting it at the new size, which for a large stateful job can take minutes. That leaves a genuinely hard question: how do you give each job the resources it needs, when it needs them, without a human in the loop and without breaking anything?

The first autoscaler: watching from outside

Our first answer, built around 2019, was an autoscaler shaped like a stream-processing job. It ran on Mantis, consuming a live feed of cluster-level metrics from Atlas, our telemetry platform including CPU, network, Kafka lag, input-rate, and consume-rate signals for every job. The scaler combined lag-derived catch-up time, CPU/network utilization thresholds, observed performance history, and regression over recent input rate to decide when to scale up or whether a smaller cluster could handle the lookahead window. Because the autoscaler operates independently of the Flink platform, it remains unaffected by issues within Flink itself. Building it as a streaming job also made it easy to scale. Each autoscaler node handled the metrics for a subset of Flink jobs, and we never had to write custom sharding or coordination logic to keep up with a growing Flink fleet. It reliably cut resource usage by 25–45% across thousands of managed pipelines. Check our previous talk at Flink Forward 2020.

But watching from outside has a ceiling. The system reasoned about a whole cluster through coarse container metrics, and it scaled a single knob, the total TaskManager count, so every operator in a job moved together. That fit the simple, single-operator pipelines it was built for, but not the multi-operator, stateful DAGs that teams were increasingly bringing to us for Ads, recommendations, and games. Those were exactly the jobs it could not reason about, and supporting each new case meant more custom logic rather than any general capability.

The autoscaler is only as good as the metrics served by external systems beneath it. Those metrics could miss real trouble: a job could be completely busy without any of it showing up as CPU utilization, leaving the job stuck in a degraded state the scaler had no way to see. Recently a networking migration quietly changed how some traffic was reported, and a subset of the Atlas metrics the scaler relied on stopped capturing everything accurately. The gap stayed invisible until it surfaced in production much later.

It was time to reconsider build versus buy.

The second autoscaler: reasoning from inside

When we started, the Flink community had no mature autoscaler to offer. By the time we re-evaluated, it did: the Apache Flink Autoscaler. Instead of watching containers from outside, it reasons from inside the job.

Figure 1: Architecture of the two Flink autoscalers

Its key idea is to estimate each operator’s true processing rate (TPR): the throughput it could sustain if it were fully busy. Flink reports, per subtask, the fraction of each second spent doing actual work, separate from time spent backpressured or idle. Dividing observed throughput by that busy fraction extrapolates capacity to full utilization: an operator handling 700 records/sec while busy 70% of the time has a TPR of 700 / 0.7 = 1,000 records/sec. Starting from the sources, the autoscaler walks the job graph and uses each operator’s TPR, its input/output ratios, and a target utilization to compute the parallelism every vertex needs so that no operator becomes the bottleneck, rather than resizing the whole cluster as a unit.

Figure 2: Flink job DAG: current → desired parallelism per vertex, based on busyness

The two approaches make a different contract, summarized below.

Table 1: Comparison of the two Flink autoscalers

The decisive difference for us is the last two rows: the OSS autoscaler can scale exactly the stateful, multi-operator jobs our homegrown system could not, and it lets each job carry its own configuration — stabilization periods, thresholds, and other scaling behavior tuned to the workload.. That made it the natural fit for the custom jobs teams had been scaling by hand.

Making it work at Netflix scale

Adopting the algorithm was straightforward; the community had done the hard part. The work for us was running it reliably across our own jobs, and this is where our system differs most from the stock open-source deployment.

Firstly, the OSS autoscaler was originally architected to reside within the Kubernetes Operator for Flink, but our Flink platform runs on its own control plane, not that operator (see our previous talk at Current Conference 2024). Community later made a fantastic decision to keep the core logic as a standalone library. They refactored four generic interfaces that made it easy to plug directly into our internal ecosystem: a context carrying job metadata and REST API info, a state store, an event handler, and a realizer that applies scaling decisions.

That service is a Spring Boot application whose orchestration runs on Temporal, the durable workflow engine. An orchestrator workflow polls our Flink control plane about once a minute for the jobs with autoscaling enabled, and starts one long-running workflow per job. Each per-job workflow pulls that job’s per-vertex metrics from its Flink JobManager, runs the OSS evaluation algorithm, and, when a scaling decision results, hands it to a realizer that actuates the change through our Flink control plane.

Figure 3: The OSS-based Flink Autoscaler architecture with Temporal workflows

The workflow-per-job design was a direct response to pain. We first ran evaluations in a single batch loop over the whole set of jobs, and it was fragile: one slow or misbehaving job could stall metric collection and scaling for every job behind it. Giving each job its own durable workflow isolated that blast radius, so a single problematic job now fails and retries on its own, and the runtime scales out as we onboard more jobs.

Secondly, three engineering gaps stood between “works in community” and “works at Netflix scale”:

  • Metric collection at high parallelism. On big jobs, pulling metrics from the JobManager became a bottleneck, and part of the cause was in Flink’s runtime. To address that, we changed the JobManager to cache transient metric names and clean them up once instead of rescanning on every fetch, and we added server-side filtering so the autoscaler asks only for the metrics it needs. This let the autoscaler work on jobs up to 3,000 Flink subtasks, where it had previously struggled above roughly 1,000. Those are in our internal fork of Flink release, while some are contributed upstream such as FLINK-36172.
  • Preserving forward chaining. Two separate vertices joined by a forward connection must run at the same parallelism, because records are handed over in memory on a fixed local channel. Scale one of them alone and Flink does not fail; it silently converts that edge into a network shuffle. Our fork detects forward-connected subgraphs and scales each as a unit.
  • Respecting sink limits. Some sinks have finite write capacity, so we added detection for async-sink backpressure (also a fork change) to keep the autoscaler from scaling a job up into a sink that cannot absorb more.

Before it actuates anything, the realizer runs a set of safety checks. For example, it refuses to scale a job down in a region being evacuated during a company-wide region failover. It also verifies there is enough disk for the new cluster to hold the job’s checkpoint state, and it adds a small standby buffer for larger clusters.

The road to one autoscaler

Last year, the OSS-based autoscaler achieved general availability for custom jobs at Netflix, yielding promising initial outcomes. For instance, our client telemetry and logging team achieved a 58% reduction in its annualized Flink compute expenditures, saving approximately $1.1 million annually. This efficiency is driven by three key factors. First, whereas static provisioning must always account for peak loads, autoscaling dynamically adapts to daily cycles, capturing the drop in traffic during nights and weekends compared to weekday peaks. Second, rather than relying on teams to manually optimize resources following performance improvements or post-holiday slowdowns, the autoscaler continually adjusts capacity. Finally, adopting uniform container dimensions enables superior bin-packing and more granular scaling increments.

Additionally, scaling down too eagerly is its own trap. Cut too deep and CPU saturates, lag spikes, and the system cannot react instantly because its metric window and stabilization period have to rebuild after each restart. We now run a target utilization of 0.45, below the community default of 0.7, deliberately trading a little efficiency for stability. Fewer and calmer rescales are worth the marginal cost for large stateful jobs.

While our scaler provides fine-grained signals and vertex-level decision units for stateful DAGs, fast rescaling still heavily depends on Flink Core’s state restoration performance. Today, the biggest remaining cost in scaling a stateful job isn’t the scaler’s logic — it’s the restart and state recovery process itself. Flink 2 addresses this through its disaggregated state architecture, keeping state in external storage rather than on local disk, which can sharply reduce how much a rescale or recovery depends on total state size. Having started supporting Flink 2.2 at Netflix, we plan on experimenting with this new state backend to see if it can help eliminate state recovery bottlenecks when scaling large stateful jobs.

Looking ahead, we aim to migrate all internal scaler use cases onto the new one based on OSS autoscaler to simplify our operational surface area.

Key Takeaways

Along the way, three lessons that generalize beyond Flink:

  • Metric choice matters more than algorithm sophistication. Our most useful debugging was rarely about the scaling math; it was about which signal to trust most. Understand your metrics before you tune your algorithm.
  • Set sensible defaults, but leave room to tune. Our managed jobs are similar enough that one good default covers most of them untouched, which is the point of a platform. But forcing a single configuration on every job punishes the ones that do not fit, so we pair defaults with per-job overrides and deliberately hide the knobs that need deep expertise. Most teams should never have to think about the autoscaler.
  • Adopt, then extend. We built in-house because in 2019 nothing mature fit our platform. When a strong community project appeared, the right move was neither to defend our investment forever nor to rip it out overnight, but to adopt it for new workloads, contribute fixes back, and plan a deliberate migration.

Thanks to the Flink and Data Mesh teams for the control-plane changes this work depended on, to the Temporal team and our early pilot teams, and to the Apache Flink autoscaler maintainers whose foundation we built on. Special thanks to Andy Zhang, Calvin Cheung, Daniel Trager, Guil Pires, Mark Cho, Matthew Kornitsky, Nikhil Sulegaon, Sujay Jain, and Tom Lee.


A Tale of Two Flink Autoscalers was originally published in Netflix TechBlog on Medium, where people are continuing the conversation by highlighting and responding to this story.

Read the whole story
alvinashcraft
1 minute ago
reply
Pennsylvania, USA
Share this story
Delete
Next Page of Stories