# Sarmad Saleem - Full Content > Software engineer, builder, and maker. Writing about engineering, tools, systems, and ideas. Source: https://sarmadsaleem.com Format: Markdown --- ## Blog Posts ### YouTube Recommendations URL: https://sarmadsaleem.com/thoughts/yt-recommendations/ Date: 2024-05-01 Tags: curation, resources I try to limit my social media consumption, but YouTube has been a net positive for me. Reason behind that is a purpose-built subscription feed, free of algorithmic noise. It is one of the few corners of the internet I actually trust. Below are the channels I follow, organized by what I get out of them. ## Educational - [Kurzgesagt](https://www.youtube.com/@kurzgesagt) - science and philosophy with incredible animation - [Crash Course](https://www.youtube.com/@crashcourse) - structured courses on almost everything - [Veritasium](https://www.youtube.com/@veritasium) - science and engineering deep dives - [TED-Ed](https://www.youtube.com/@TEDEd) - short animated lessons on a wide range of topics - [Cleo Abram](https://www.youtube.com/@CleoAbram) - optimistic tech explainers - [PolyMatter](https://www.youtube.com/@PolyMatter) - economics and geopolitics, well researched - [Branch Education](https://www.youtube.com/@BranchEducation) - how hardware and technology actually works - [Economics Explained](https://www.youtube.com/@EconomicsExplained) - economic concepts made accessible - [Big Think](https://www.youtube.com/@bigthink) - expert-led talks on science, philosophy, and culture ## Tech - [MKBHD](https://www.youtube.com/@mkbhd) - consumer tech reviews, best in class - [Fireship](https://www.youtube.com/@Fireship) - fast-paced dev tutorials and industry commentary - [Strange Parts](https://www.youtube.com/@StrangeParts) - hardware hacking and manufacturing adventures - [Honeypot](https://www.youtube.com/@Honeypotio) - developer documentaries - [Andrej Karpathy](https://www.youtube.com/@AndrejKarpathy) - neural networks explained from first principles - [Art of the Problem](https://www.youtube.com/playlist?list=PLhr1KZpdzukdeX8mQ2qO73bg6UKQHYsHb) - the history and science of information ## Geopolitics & curiosity - [Johnny Harris](https://www.youtube.com/@johnnyharris) - geopolitics and borders, beautifully produced - [Hoog](https://www.youtube.com/@hoogyoutube) - geopolitical analysis with good depth - [Real Life Lore](https://www.youtube.com/@RealLifeLore) - geography and geopolitics explained visually - [Vox](https://www.youtube.com/@Vox) - explainers on culture, policy, and world events - [Bloomberg](https://www.youtube.com/@business) - business and economic reporting - [TLDR News EU](https://www.youtube.com/@TLDRnewsEU) - European politics, concise and balanced - [The Economist](https://www.youtube.com/@TheEconomist) - global affairs and economic analysis - [ColdFusion](https://www.youtube.com/coldfusion) - technology and business stories with great storytelling - [Urban Stories](https://www.youtube.com/@UrbanStories) - city planning and urban design ## Business & entrepreneurship - [Garry Tan](https://www.youtube.com/@GarryTan) - YC president, startup advice and Silicon Valley insights - [Slidebean](https://www.youtube.com/@slidebean) - startup breakdowns and pitch deck analysis ## Financial freedom - [Nischa](https://www.youtube.com/@nischa) - personal finance and building wealth - [Ben Felix](https://www.youtube.com/@BenFelixCSI) - evidence-based investing, cuts through financial myths - [The Plain Bagel](https://www.youtube.com/@ThePlainBagel) - financial concepts explained clearly without the hype ## Podcasts - [Lex Fridman](https://www.youtube.com/@lexfridman) - long-form conversations with scientists, engineers, and thinkers - [Colin and Samir](https://www.youtube.com/@ColinandSamir) - the creator economy and what makes content work - [Joe Rogan](https://www.youtube.com/@joerogan) - wide-ranging interviews, hit or miss but the hits are great - [Rich Roll](https://www.youtube.com/@richroll) - health, wellness, and endurance-oriented conversations - [Lenny's Podcast](https://www.youtube.com/@LennysPodcast) - product management and growth, highly tactical ## Psychology - [The School of Life](https://www.youtube.com/@theschooloflifetv) - emotional intelligence and relationships - [Daily Stoic](https://www.youtube.com/@DailyStoic) - stoic philosophy applied to modern life - [Naval](https://www.youtube.com/@NavalR) - wealth, happiness, and clear thinking ## Lifestyle - [Matt D'Avella](https://www.youtube.com/@mattdavella) - minimalism, habits, and filmmaking - [Ali Abdaal](https://www.youtube.com/@aliabdaal) - productivity and building a life you enjoy - [Nathaniel Drew](https://www.youtube.com/@nathanieldrew) - thoughtful videos on language, travel, and identity - [No Backup Plan](https://www.youtube.com/@nobackupplan) - life as a creative - [Matthew Encina](https://www.youtube.com/@MatthewEncina) - design, creativity, and career - [Never Too Small](https://www.youtube.com/@nevertoosmall) - compact living and clever apartment design ## Travel - [Attaché](https://www.youtube.com/@Attachetravel) - cinematic city guides - [Thrillist](https://www.youtube.com/@thrillist) - food and travel culture - [Top Jaw](https://www.youtube.com/@topjaw) - food-focused travel - [Little Big World](https://www.youtube.com/@LittleBigWorld) - tilt-shift miniature timelapses of cities - [Kraig Adams](https://www.youtube.com/@kraigadams) - solo hiking and minimal dialogue travel films ## Food - [J. Kenji López-Alt](https://www.youtube.com/@JKenjiLopezAlt) - POV home cooking from a food science perspective - [My Name Is Andong](https://www.youtube.com/@mynameisandong) - recipes with cultural context - [Eater](https://www.youtube.com/@eater) - food culture and restaurant stories ## Cars - [Hagerty](https://www.youtube.com/Hagerty) - classic cars and automotive culture - [AMMO NYC](https://www.youtube.com/@AMMO-NYC) - detailing and car care - [Top Speed Germany](https://www.youtube.com/topspeedgermany) - autobahn runs and speed tests ### Every phone I ever owned URL: https://sarmadsaleem.com/thoughts/every-phone-i-ever-owned/ Date: 2024-04-02 Tags: personal A childhood friend of mine and I took a trip down the memory lane about which mobile phones we used to own. Between us we pieced together an approximate timeline. Nineteen phones in twenty years. Polyphonic ringtones to Face ID. ## Before smartphones The Nokia tune. T9 texting. Snake. Beaming files over infrared. GPRS GPRS to EDGE. Batteries lasted a week because there wasn't much to drain them. Phones were built like bricks and survived being treated like ones too. The Sony Ericssons were a step up with actual cameras and better media. The 6600 ran Symbian, which was the closest thing to a smartphone OS back then. You could install apps, sort of. The E63 was my first QWERTY keyboard after years on a numpad.
Nokia 1100
Nokia 1100
2003
Nokia 6030
Nokia 6030
2004
Sony Ericsson K510
Sony Ericsson K510
2005
Nokia 1112
Nokia 1112
2005
Nokia 6600
Nokia 6600
2005
Sony Ericsson K750
Sony Ericsson K750
2006
Nokia E63
Nokia E63
2009
## The Android wild west Android opened everything up. Root your phone, flash custom ROMs, brick it, recover it, do it again. Four phones in two years. Every manufacturer had a different take and nothing was locked down.
Samsung S2
Samsung S2
2011
LG G2
LG G2
2012
HTC One X
HTC One X
2012
LG Optimus G E970
LG Optimus G
2012
## Finding a lane First taste of iOS fluidity with the iPhone 5. Then a BlackBerry Bold because I missed physical keyboards. Then a Nexus 5 for that stock Android experience. I was bouncing between ecosystems trying to figure out what mattered more, the freedom to tinker or things just working.
Apple iPhone 5
iPhone 5
2014
BlackBerry Bold 9700
BlackBerry Bold 9700
2015
Google Nexus 5
Google Nexus 5
2015
## Locked into Apple Turns out I wanted things to "just work". AirDrop, iMessage, handoff between devices. Apple kept adding small conveniences that made switching feel like more hassle than it was worth. Five iPhones in and I'm fully locked in. Can't be bothered to leave.
Apple iPhone 6
iPhone 6
2016
Apple iPhone 8
iPhone 8
2018
Apple iPhone XR
iPhone XR
2020
Apple iPhone 13
iPhone 13
2022
Apple iPhone 15 Pro Max
iPhone 15 Pro Max
2023
### Berlin Guide URL: https://sarmadsaleem.com/thoughts/berlin-guide/ Date: 2023-10-30 Tags: travel This started as a simple list for colleagues visiting Berlin for work. Over time it grew into something I kept sharing with friends moving to the city. I wish I had this when I first moved here. Berlin is a weird city, but one I've grown to like. ## Crowd-sourced guides These two cover the widest range of topics for living in Berlin: - [All About Berlin](https://allaboutberlin.com/guides) - [Simple Germany](https://www.simplegermany.com/) ## In case of emergency All emergency numbers are listed [here](https://allaboutberlin.com/guides/emergency-numbers). The two you need to know: - 👮‍♂️ `110` for Police - 🚨 `112` for Emergency Services (Ambulance, Fire Brigade) ## There's an app for everything - **Food delivery** - Wolt, Uber Eats - **Quick commerce** - Gorillas, Flink, Getir - **Ride hailing** - FreeNow, Uber, Bolt. More on [getting around Berlin](https://www.berlin.de/en/getting-around/). - **Medicine delivery** - MAYD - **Apartments** - [ImmoScout24](https://www.immobilienscout24.de/) is the default. Competition is fierce, especially inside the ring. Chasing leads on new builds via [Neubau Compass](https://www.neubaukompass.com/berlin/) and reaching out directly often works better. Best bet is usually a recommendation from colleagues or friends. - **Banking** - Deutsche Bank, Commerzbank, C24 for traditional. Revolut and Wise for neobanks. Wise is also the best for remittances. - **Insurance** - Very common in Germany. Health insurance (TK) comes with your job. Beyond that, consider legal, home contents, and liability insurance. [Getsafe](https://www.hellogetsafe.com/) is a good option. - **Tax returns** - Find a tax advisor or use [Taxfix](https://taxfix.de/). ## Sundays are weird Everything is closed. Grocery stores, malls, all of it. Only restaurants stay open. [This guide](https://allaboutberlin.com/guides/open-on-sundays-in-berlin) lists what's actually open on Sundays. ## Neighbourhoods [Berlin neighbourhood map](https://www.tripsavvy.com/berlin-germany-neighborhood-guide-4140486) for the official overview and [Hoodmaps](https://hoodmaps.com/berlin-neighborhood-map) for the crowd-sourced version. ## Events [Visit Berlin event calendar](https://www.visitberlin.de/en/event-calendar-berlin) ## Desi grocery stores [View all on Google Maps](https://maps.app.goo.gl/UgNkAMirXwwnHCAw7?g_st=i) - [Zora Supermarket](https://maps.app.goo.gl/2rSuyzRW5C2wAXF1A) - [Tariq Store](https://maps.app.goo.gl/1FU2b4pyeS5YXETe9?g_st=ic) - [Punjab Food](https://maps.app.goo.gl/GBzWT4nfDJPhNDqdA?g_st=ic) - [Dong Xuan](https://maps.app.goo.gl/DZW5C16GDyv6BEDC7?g_st=ic) - Order online via [Spicy Village](https://www.spicevillage.eu/) ## Halal meat shops [View all on Google Maps](https://maps.app.goo.gl/8uJxFQRvqDMRtUJY9?g_st=i) - [Al Kaiser Supermarkt](https://maps.app.goo.gl/YBUNUkJrcsdGSnWBA?g_st=ic) - [Almaraii Markt](https://maps.app.goo.gl/kffQGV5yDBqu8p2WA?g_st=ic) - [Balaban Supermarkt](https://maps.app.goo.gl/cEb3KnCHnWV9iYDR6?g_st=ic) - [Eurogida](https://maps.app.goo.gl/PgrL5whn4ANpKmPi9?g_st=ic) (Turkish chain, multiple locations) - [Bolu](https://maps.app.goo.gl/npf6Jivdxb8aesmu8?g_st=ic) (Turkish chain, multiple locations) ## Movies in English Most cinemas show German dubbed versions. [OV Berlin](https://ov-berlin.info/) aggregates showtimes for original versions. Look for: - `OV` - Original Version - `OmeU` - Original with English Subtitles - `OmU` - Original with German Subtitles Also useful: [English cinemas in Berlin](https://allaboutberlin.com/guides/english-cinemas-berlin). ## Moving apartments **DIY** - Buy boxes and supplies from Amazon or Bauhaus. Get friends to help (🍕 mandatory). Rent a van via Sixt or Miles, or get one with a driver. **Hands-off** - [Umzug365](https://www.umzug-365.de/), [Moovick](https://moovick.com/service/berlin), or [Movinga](https://www.movinga.com/de/de/). ## Recycling Germans take this seriously. [Here's how to sort your trash](https://allaboutberlin.com/guides/sorting-trash-in-germany). ## Expat newsletters - [20 Prozent](https://substack.com/@20prozent) - [Berlin Daily] (https://berlindaily.org/) - [The Berliner](https://www.the-berliner.com/) ## Exploring the city [Curated places on Google Maps](https://maps.app.goo.gl/NPoteGfn1Zyw8hqSA) ### Staying in the Know URL: https://sarmadsaleem.com/thoughts/staying-in-the-know/ Date: 2023-02-03 Tags: meta, resources I was recently asked what are some of my favourite tech sources to stay in the know, so here comes the rolodex: ## Twitter Still the fastest signal for what's actually happening. Follow the right people, mute the noise. The algo is garbage but the network effects are unmatched. ## [Devo](https://chromewebstore.google.com/detail/devo/elkhalpmbmbaeoemecpcfdcoekmpgmdm?hl=en) Browser extension that turns your new tab into a dev news feed. Low effort, high surface area. Good for catching things you didn't know you should care about. ## Tech Radar by ThoughtWorks Quarterly. Opinionated. Useful for understanding where the industry *thinks* it's heading vs. where it actually is. I don't agree with all their takes, but that's the point. ## Newsletters - [TLDR](https://tldr.tech). Daily, skimmable, no fluff. - [The Overflow](https://stackoverflow.blog/newsletter). Stack Overflow's weekly. Good pulse on what devs are actually asking about. - [Level Up](https://levelup.patkua.com). Pat Kua's newsletter for tech leads. Short, practical, no fluff. - [The Pragmatic Engineer](https://newsletter.pragmaticengineer.com). Gergely Orosz writes the real stuff about big tech eng culture. - [Architecture Notes](https://architecturenotes.co). Deep dives that don't waste your time. - [ByteByteGo](https://blog.bytebytego.com). System design porn, great visuals. ## YouTube - [Fireship](https://www.youtube.com/@Fireship). 100 seconds of dopamine. Perfect for mass-produced context on things you'll never actually use. - [This is My Architecture](https://www.youtube.com/playlist?list=PLhr1KZpdzukdeX8mQ2qO73bg6UKQHYsHb). AWS customers explaining their real architectures. Less polished, more honest. ### Coaching Lighthouse URL: https://sarmadsaleem.com/thoughts/coaching-lighthouse/ Date: 2021-09-18 Tags: leadership, engineering-management A few years into coaching engineers, crystallizing this for myself. ## Feedback Adopt Radical Candor. Care personally, but challenge directly. Most people drift toward empathy without challenge (ruinous) or challenge without empathy (obnoxious). You need the intersection. ## Inclusion The loudest voice fills the vacuum. Introverted team members often have the best technical insights but the least desire to fight for airtime. Your job is to moderate the volume. ## Psychological Safety Blameless postmortems are non-negotiable. Assume good intent, especially in remote/async work where tone is lost. ## Emotional Intelligence The most useful signal isn't on Grafana. It's someone going quiet in standup, or crossed arms during a retro. If *you* feel triggered, pause. Replace your reaction with a clarifying question. ## Self Awareness Normalize "I don't know." Validate your own assumptions by asking "What if I'm wrong?" If you hide your blind spots, your team will hide theirs. ## Conflict Resolution Move from *what* to *why*. If two engineers are deadlocked on implementation, use a "time-boxed experiment." Prototype both for 3 days and let the metrics pick the winner. ## Maker to Multiplier The hardest mental shift. Your code is no longer the product. The team is the product. (See: Pat Kua's [Maker vs Multiplier](https://www.patkua.com/blog/maker-vs-multiplier/)). ### Kubernetes Monitoring URL: https://sarmadsaleem.com/thoughts/kubernetes-monitoring/ Date: 2021-02-23 Tags: devops, kubernetes, monitoring As increasing number of software teams opt for microservices architecture, they tend to choose containers as their preferred way of packaging and shipping their applications. Some of the benefits that containers can afford us, as opposed to bare metal or virtual machines, include superior fault isolation, better resource utilization and ability to scale workloads faster. Even though there are other container runtimes out there, people tend to use containers and Docker almost interchangeably now, that goes to show the impact Docker has had on the landscape. While containers by themselves are extremely useful, they can become quite challenging to deploy, manage, and scale across multiple hosts in different environments. That's where container orchestration comes into play. Similar to how Docker became the de-facto for containerization, the industry has found Kubernetes as the front runner in container orchestration landscape. At its core, Kubernetes eases management of containerized apps at scale by pools underlying resources, like compute, memory, disk space etc, into one unified blob and allows you to manage that pool using different objects via an API. These objects work together to deliver a robust, reliable and resilient application experience. You might have heard the analogy of pets vs cattle in cloud native context, the idea is that old way of maintaining indispensable servers was pet way of doing things, whereas moving towards low-maintenance ephemeral servers is close to cattle model as they don't require the same level of nurturing and attention. This paradigm shift allows tools like Kubernetes to really shine. While Kubernetes drastically streamlines container lifecycle management at scale, it also comes with its own set of complexities and maintenance overhead. Legacy ways of monitoring infrastructure and trouble shooting workflows which were primarily focused on static target (pet model) fall flat in todays dynamic world of Kubernetes (cattle model). The abstractions that make Kubernetes powerful, also force us to redefine the way we monitor underlying infrastructure, cluster itself and the applications running atop. In this guide, we'll unpack Kubernetes monitoring by discussing why legacy methods fail, why we need specialized monitoring, what should we monitor and some of the popular tools teams tend to reach out for. ## Kubernetes primer Before we jump into the why of monitoring, let's use this primer as a brief recap of Kubernetes. It's open-source software that has become the gold standard for orchestrating containerized workloads in private, public, and hybrid cloud environments. At its core, it orchestrates clusters of virtual machines and schedules containerized apps to run atop based on specified constraints. On a higher level, Kubernetes architecture is comprised of a control plane and group of worker nodes. Control plane is the brain responsible for accepting user instructions and figuring out the best way to execute them. Whereas worker nodes are machines responsible for obeying instructions from the control plane and running containerized workloads. Control plane and worker nodes use Kubernetes native objects like pods, deployments, stateful set, deamon set, service, ingress etc to orchestrate workloads. Containers are grouped into pods, the basic operational unit for Kubernetes, and those pods scale to your desired state typically using high level objects like deployment, stateful set or daemon set. These high level objects are then tied to services for automatic service discovery. Kubernetes also incorporates load balancing, tracks resource allocation, and scales based on compute utilization. And, it checks the health of individual resources and enables apps to self-heal by automatically restarting or replicating containers. Some of the major features include self-healing, horizontal scaling, load balancing, service discovery, automated rollouts and rollbacks, secrets and configuration management, storage orchestration and batch execution. ## Why do we need monitoring? Generally speaking, monitoring helps maximize availability and performance for your applications and services by collecting, analyzing and acting on telemetry from your environment. It not only helps you understand how your applications are performing but also proactively identify issues so that you can resolve them thus minimizing them time end-user is affected. In other words, its the feedback from your applications. Monitoring helps achieve the goal of ensuring high availability by minimizing time to detect (TTD) and time to mitigate (TTM). Additionally monitoring also enables validated learning by tracking usage. The concept of validated learning relates to development teams collecting data to support the hypotheses that led to the development of a feature and eventual deployment. If a hypothesis is proved wrong, team can fail fast or pivot, thus enforcing a strong feedback loop. Effective monitoring is essential to allow DevOps teams to deliver at speed, get feedback from production, and increase customers satisfaction, acquisition and retention. From developer experience perspective, Kubernetes abstractions may seem straight forward. For example your containers are placed inside pods, pods are grouped under deployments, deployments are exposed via services and services accept external traffic via ingresses. The learning curve seems to be manageable here, but in this example the scope is limited to your application workloads. If we change the scope to running a Kubernetes cluster itself, there's a lot more that we need to think about. For example, how to provision cluster, secure it, monitor it, make it scale, bootstrap it with necessary add-ons as per application needs, perform back-ups, upgrade to newer version and handle disaster recovery. This can be an overwhelming task and requires in-depth knowledge about Kubernetes. Now, what if someone can run Kubernetes for you? That's exactly the pain-point that managed Kubernetes services like AWS Elastic Kubernetes Service, GCP Google Kubernetes Engine & Microsoft Azure Kubernetes Service among others, attempt to solve. The idea is that cloud provider will manage the master nodes for us given some SLA, that way fair bit of additional overhead is outsourced. Regardless of whether you run your own Kubernetes cluster on bare metal or use a managed Kubernetes service, the approach towards monitoring your cluster remains relatively similar. You define metrics worth monitoring, establish rules and alert respective teams if any of those rules are breached. ## What should I monitor? Kubernetes solved some old problems but it also created some new challenges like added complexity to logging and monitoring in a dynamic environment. There are bunch of variables at play, underlying hosts, the platform itself, containers and objects that orchestrate those containers, all of which must to be monitored. You can look at this question of what to monitor from multiple lenses, for example control plane metrics, cluster state metrics, resource metrics, bandwidth, latency, application health and performance. To keep things digestable, let's try to transform afore-mentioned lenses into two supersets i.e. Kubernetes monitoring & application monitoring. ### Key metrics for Kubernetes monitoring Kubernetes exposes a bunch of metrics out of the box. Additionally you can pair some add-ons that can help aggregate data collected from the kubelet on each node in a more streamlined manner. This can then be used by a number of visualization tools. If we zoom in a little bit, the interface of the metrics registry consists of three separate APIs, namely Resource Metrics API, Custom Metrics API & External Metrics API. For now, we'll only focus on Resource Metrics API. Some key metrics to consider monitoring include: - **Cluster state metrics**, including the health and availability of pods. - **Node status**, including readiness, memory, disk or processor overload, and network availability. - **Pod availability**, since unavailable pods can indicate configuration issues or poorly designed readiness probes. - **Memory utilization** at the pod and node level. - **Disk utilization** including lack of space for file system and index nodes. - **CPU utilization** in relation to the amount of CPU resource allocated to the pod. - **API request latency** measured in milliseconds, where the lower the number the better the latency. ### What about my applications? How do I monitor them? Our application workloads are composed of Kubernetes objects like pods, deployments, services, ingresses etc., most of which are covered by the key metrics we discussed above. Let's zoom into the probe mechanism, something we brushed on above. Some of above-mentioned metrics depend on this mechanism heavily. Kubernetes uses liveness probes to determine when to restart a container, whereas it uses readiness probes to decide when a container is ready to accept traffic. Using these probes Kubernetes can understand when containers are alive and ready to accept traffic, this understanding feeds into state of different objects (pods, deployments etc) and gives us a holistic picture. But that's not where the story of monitoring our applications ends. So far we have talked about monitoring pod availability, resource utilization and using probes, all of these things are very Kubernetes specific. What about application specific metrics? One of the ways to do application specific custom metrics on Kubernetes is through the Custom Metrics API. For example you could instrument your application to expose the metric as a Prometheus metric, install and configure Prometheus to collect it from all pods and use Prometheus Adapter to expose it as custom metric. This metric can then be used for visualization, reporting and even dynamically autoscaling workloads if needed. Another way is to keep application specific metrics isolated from the infrastructure and orchestration layer altogether. What about application logging? A container running on Kubernetes writes its logs to stdout and stderr streams, which are them picked up by kubelet service running on that node and delegated to container engine for handling based on logging driver configured. Another way to handle logs is to leverage a side-car pattern, most third party logger aggregation services use similar pattern to implement their agent. Typically purpose of logging is to create an ongoing record of application events so that teams can refer to this record to troubleshoot issues. Monitoring on the other hand is more metric focused, these metrics are then used to alert teams of any analogies. You can think of monitoring as security alarm whereas logs as security camera footage. ## Kubernetes monitoring tools Now that we have gone through a primer on Kubernetes, looked into why we need monitoring and discussed what we should monitor, let's dive into some of the monitoring tools. Let's start with some rudimentary out of the box options and build upto more popular production ready solutions. In this section we'll learn by example how all these tools stack up against each other. We'll be using a cluster created on AWS EKS (Elastic Kubernetes Service) using [eksctl](https://eksctl.io/) command-line utility. As an alternative you can also try to follow along locally using Kubernetes bundled with Docker Desktop or Minikube. As long as you have access to a cluster via kubectl, following examples should apply as is. ```bash # Install eksctl brew tap weaveworks/tap brew install weaveworks/tap/eksctl eksctl version ``` The cluster manifest (`cluster.yaml`): ```yaml apiVersion: eksctl.io/v1alpha5 kind: ClusterConfig metadata: name: basic-cluster region: eu-central-1 nodeGroups: - name: ng-1 instanceType: t3a.small desiredCapacity: 3 volumeSize: 100 ``` ```bash # Provision cluster on AWS EKS based on manifest eksctl create cluster -f cluster.yaml ``` ### kubectl Kubernetes API server exposes a number of metrics out of the box that are useful for monitoring and analysis. These metrics are exposed internally through a metrics endpoint that refers to the `/metrics` HTTP API. Let's see it in action: ```bash # Get raw metrics kubectl get --raw /metrics ... # HELP rest_client_requests_total Number of HTTP requests, partitioned by status code, method, and host. # TYPE rest_client_requests_total counter rest_client_requests_total{code="200",host="127.0.0.1:21362",method="POST"} 4994 rest_client_requests_total{code="200",host="127.0.0.1:443",method="GET"} 1.326086e+06 rest_client_requests_total{code="200",host="127.0.0.1:443",method="PUT"} 862173 ``` ### Kubernetes Metric Server [Kubernetes Metrics Server](https://github.com/kubernetes-sigs/metrics-server) is a scalable, efficient source of container resource metrics for Kubernetes built-in autoscaling pipelines. This add-on collects resource metrics from Kubelets and exposes them in Kubernetes apiserver through Metrics API for use by Horizontal Pod Autoscaler and Vertical Pod Autoscaler. Metrics API can also be accessed by `kubectl top`, making it easier to debug autoscaling pipelines. While metric server is useful to get resource usage metrics for nodes or pods, by itself its very bare bones and is typically used in conjunction with pod scalers, Kubernetes Dashboard etc. ```bash # Deploy metric-server kubectl apply -f https://github.com/kubernetes-sigs/metrics-server/releases/download/v0.3.6/components.yaml # Verify deployment kubectl get deployment metrics-server -n kube-system NAME READY UP-TO-DATE AVAILABLE AGE metrics-server 1/1 1 1 42m # Get node resource usage kubectl top nodes NAME CPU(cores) CPU% MEMORY(bytes) MEMORY% ip-192-168-31-120.eu-central-1.compute.internal 53m 2% 428Mi 28% ip-192-168-38-58.eu-central-1.compute.internal 64m 3% 409Mi 26% ip-192-168-79-154.eu-central-1.compute.internal 62m 3% 438Mi 28% ``` ### Kubernetes Dashboard [Kubernetes Dashboard](https://github.com/kubernetes/dashboard) is a general purpose, web-based UI for Kubernetes clusters. It allows users to manage applications running in the cluster and troubleshoot them, as well as manage the cluster itself. Dashboard also provides information on the state of Kubernetes resources in your cluster and on any errors that may have occurred. ```bash # Deploy kubernetes dashboard kubectl apply -f https://raw.githubusercontent.com/kubernetes/dashboard/v2.0.0/aio/deploy/recommended.yaml # Verify deployment kubectl get deployment -n kubernetes-dashboard # Get access token kubectl -n kube-system describe secret $(kubectl -n kube-system get secret | grep eks-admin | awk '{print $1}') # Start proxy kubectl proxy # Launch dashboard open http://localhost:8001/api/v1/namespaces/kubernetes-dashboard/services/https:kubernetes-dashboard:/proxy/#!/login ``` ![Kubernetes Dashboard demo](/images/blog/k8s-dashboard-demo.gif) Kubernetes Dashboard incorporates raw metrics and metric server to provide a web based interface which provides a high level overview as well as resource level drill down functionality. It's a good place to start but teams soon outgrow this dashboard and reach out for more production-ready solutions like Prometheus and kubewatch. ### Lens [Lens](https://k8slens.dev/) advertises itself as the only Kubernetes IDE you'll ever need to take control of your Kubernetes clusters. It is a standalone application for MacOS, Windows and Linux operating systems. It is open source and free. It extends a lot of Kubernetes dashboard functionality while capitalizing on good usability, multi-cluster management, terminal access to nodes and pods to name a few. ```bash # Install Lens brew cask install lens # Launch Lens & select desired kubeconfig to connect to your cluster open /Applications/Lens.app ``` ![Lens IDE demo](/images/blog/k8s-lens-demo.gif) Lens is quite comparable to Kubernetes Dashboard, a good GUI alternative that can save some time doing repetitive tasks over the API. However it still falls under basic monitoring as there's no concept of rules or alerts out of the box. For that kind of setup, you'll have to side with Prometheus or third party services like Datadog, NewRelic etc. ### Prometheus, Grafana & Alert Manager So far the tools we have been looking at are following a progression from rudimentary to more production ready setups. Typically such a monitoring system consists of a time-series database that houses metric data and a visualization layer. In addition, an alerting layer creates and manages alerts, handing them off to integrations and external services as necessary. Finally, one or more components generate or expose the metric data that will be stored, visualized, and processed for alerts by the stack. One popular monitoring solution is the open-source Prometheus, Grafana, and Alertmanager stack, deployed alongside kube-state-metrics and node_exporter to expose cluster-level Kubernetes object metrics as well as machine-level metrics like CPU and memory usage. - **Prometheus** - It is time series database and monitoring tool that works by polling metrics endpoints and scraping and processing the data exposed by these endpoints. It allows you to query this data using PromQL, a time series data query language. - **Grafana** - It is data visualization and analytics tool that allows you to build dashboards and graphs for your metrics data. - **Alertmanager** - Typically installed alongside Prometheus, forms the alerting layer of the stack, handling alerts generated by Prometheus and deduplicating, grouping, and routing them to integrations like email or PagerDuty. - **kube-state-metrics** - Add-on agent that listens to the Kubernetes API server and generates metrics about the state of Kubernetes objects like Deployments and Pods. These metrics are served as plaintext on HTTP endpoints and consumed by Prometheus. - **node-exporter** - Prometheus exporter that runs on cluster nodes and provides OS and hardware metrics like CPU and memory usage to Prometheus. We'll be using Kubernetes package manager [Helm](https://helm.sh/) to install this to our cluster instead of applying multiple manifests individually. ```bash # Create namespace kubectl create namespace prometheus # Deploy prometheus helm repo add prometheus-community https://prometheus-community.github.io/helm-charts helm repo add stable https://charts.helm.sh/stable helm install prometheus prometheus-community/prometheus \ --namespace prometheus \ --set alertmanager.persistentVolume.storageClass="gp2" \ --set server.persistentVolume.storageClass="gp2" # Verify deployment kubectl get deployment -n prometheus NAME READY UP-TO-DATE AVAILABLE AGE prometheus-alertmanager 1/1 1 1 8m20s prometheus-kube-state-metrics 1/1 1 1 8m20s prometheus-pushgateway 1/1 1 1 8m20s prometheus-server 1/1 1 1 8m20s # Port forward prometheus kubectl --namespace=prometheus port-forward deploy/prometheus-server 9090 open localhost:9090 ``` ![Prometheus dashboard demo](/images/blog/k8s-prometheus-demo.gif) This gives you an insight into how Prometheus gives us an expressive query language to sift through plethora of metrics and define alerts based on that. Let's throw Grafana into the mix and visualize some of the metrics. ```bash # Create namespace kubectl create namespace grafana ``` The Grafana datasource config (`grafana-values.yaml`): ```yaml datasources: datasources.yaml: apiVersion: 1 datasources: - name: Prometheus type: prometheus url: http://prometheus-server.prometheus.svc.cluster.local access: proxy isDefault: true ``` ```bash # Deploy grafana helm repo add grafana https://grafana.github.io/helm-charts helm install grafana grafana/grafana \ --namespace grafana \ --set persistence.storageClassName="gp2" \ --set persistence.enabled=true \ --set adminPassword='super-secret-password' \ --values grafana-values.yaml \ --set service.type=LoadBalancer # Verify deployment kubectl get deployment -n grafana # Access grafana dashboard export ELB=$(kubectl get svc -n grafana grafana -o jsonpath='{.status.loadBalancer.ingress[0].hostname}') open "http://$ELB" ``` ![Grafana dashboard demo](/images/blog/k8s-grafana-demo.gif) Now that we have seen Prometheus and Grafana in action, it's easy to imagine how alerts can be setup on some of the metrics using AlertManager. Just to recap, Prometheus collects metrics and forwards them to AlertManager which takes care of deduplicating, grouping, and routing them to the correct receiver integration such as email, PagerDuty, or Slack. An example rule may look something like this: ```yaml # Sample configuration for prometheus rule and alert serverFiles: alerting_rules.yml: groups: - name: Instances rules: - alert: InstanceDown expr: up == 0 for: 5m labels: severity: page annotations: description: '{{ $labels.instance }} of job {{ $labels.job }} has been down for more than 5 minutes.' summary: 'Instance {{ $labels.instance }} down' alertmanagerFiles: alertmanager.yml: global: null route: group_by: [alertname, job] receiver: slack_alerting receivers: - name: slack_alerting slack_configs: - api_url: 'https://hooks.slack.com/foo' username: 'Alertmanager' channel: '#prometheus-alerts' send_resolved: true ``` Key take away here is that Prometheus and AlertManager allows us to declare rules based on metrics and notify teams using mediums like Slack, PagerDuty, Email etc when those rules are breached. ### Proprietary solutions Similar to above mentioned open source solutions, monitoring landscape is ripe with proprietary ones as well. They take the pain away from setting up and maintaining your own monitoring and alerting system. Some notable options include: - AWS CloudWatch with EKS - Google Cloud Operations with GKE - Microsoft Azure Monitor with AKS - Datadog - Epsagon - Dynatrace ## Wrapping up We started by revisiting some of the core concepts of Kubernetes, unpacked why, what and how of Kubernetes monitoring. We explored some open-source tools by example and briefly touched upon managed monitoring services that aim at reducing the overhead of Kubernetes monitoring. Once we start to move beyond infrastructure monitoring and focus on the most important piece of the puzzle (our applications) then application performance monitoring becomes increasingly important. Instrumenting your apps, understanding request flows, and identifying bottlenecks at the code level is the next step in building a truly observable system. ### Container Orchestration Overview URL: https://sarmadsaleem.com/thoughts/container-orchestration-overview/ Date: 2020-10-29 Tags: devops, kubernetes, containers The way we write, ship and maintain software today has evolved drastically in last few years. How we consume underlying infrastructure to run our software has matured significantly, in that we have seen a transition from bare metal to virtual machines to containers to micro-VMs. Rise in adoption of microservices has certainly paved the way for containers to be the primary way organizations are packaging and shipping their applications today. Amid this evolution, we have seen Docker become almost synonymous with containers and Kubernetes emerge as the gold standard to orchestrate those containers. Some of the primary benefits of this transition include fault isolation, resource utilization and scaling of workloads, all of which have a direct impact on the business. In this post, we'll get into the what and why of container orchestration. We'll also take a look at some of the leading tools out there and stack them up against each other with the aim to help you choose the right tool for the job. ## What are containers again? Why do we even need them? Back in the day, we ran applications on bare-metal, which is another way of saying physical servers or on-prem. It was time intensive, expensive and error-prone endeavor which was extremely slow to scale. In came virtual machines to solve these pain points, a layer of abstraction on top of physical servers allowing us to run multiple operating systems in complete isolation on the same physical servers enforcing better security, resource utilization, scalability and significant reduction in costs. Great, if we have already addressed aforementioned pain points with virtual machines, then why are we even talking about containers? Well, containers take it up a notch. You can think of them as mini virtual machines that, instead of packaging a full-fledged operating system, try to leverage the underlying host OS for most things. Container-based virtualization guarantees higher application density and maximum utilization of server resources. An important distinction between virtual machines and containers is that VM virtualizes underlying hardware whereas the container virtualizes the underlying operating system. Both have their use cases, in fact, many container deployments use VM as their host operating system rather than running directly on bare metal. Whether you are building a monolith or a microservices based architecture, if resilience and ability to scale fast are important to you, for most types of workloads, containerization is your best bet when it comes to packaging your application. ## What, exactly, is Container Orchestration? While containers by themselves are extremely useful, they can become quite challenging to deploy, manage, and scale across multiple hosts in different environments. Container orchestration is another fancy word for streamlining this process. Let's unpack it a bit further. ![Container orchestration diagram](../../assets/blog/container-orchestration-diagram.png) At its core, container orchestration is about managing lifecycle of containers. Whether you are running a monolith or a bunch of microservices, container orchestration tool can help you streamline the container lifecycle management in both scenarios. However, it's real utility really shines through at scale in complex dynamic environments. Tools in this space help teams to control and automate many tasks including: - Recover from failures when encounters, ensuring that your apps are self-healing, robust and resilient. - Provision and schedule containers by allocating required resources based on define configuration - Scale services by adding or removing containers, typically based on some metrics. - Monitor health of containers and hosts - Expose services to the outside world - Load balance traffic between multiple containers seamlessly Most container orchestration tools follow similar mechanism from a consumer point of view. You describe configuration of your application in a tool specific DSL. These configuration files (typically YAML or JSON files) tell the orchestration tool things like where to get container images, how to do networking, how to handle storage volumes and where to push logs. An application may be deployed in multiple environments like development, staging and production. In most cases, software teams version control their environment specific configuration files to make things auditable and reproducible. These configuration files are handed off to the tool using an interface (typically CLI), tool then schedules the deployment, selects the best host to place the containers based on constraints defined in the configuration. Once containers are up and running, tool continuously monitors the app by matching desired state with actual state in addition to querying health checks, if anything doesn't add up, it tries to recover from that failure automatically. Being able to have run these orchestration tools in disparate environments ranging from a desktop to bare metal to cloud based VMs is a big selling point. ## Popular tools When Docker emerged in 2013, containers exploded in popularity. A number of tools have since been developed to make container management easier, while they have been around for years, many consider 2017 to be the year that container tools came of age. As of today there are several open-source and proprietary solutions to manage containers out there. In the open source space, Kubernetes, Docker Swarm, Apache Marathon on Mesos and Hashicorp Nomad are some of the notable players. While the proprietary space is dominated by leading cloud providers, some of the notable examples include Amazon Web Services (AWS) Elastic Container Service, Google Cloud Platform (GCP) Compute Engine & Cloud Run, Microsoft Azure Container Instances & Web Apps for Containers. Let's zoom into some of the most popular ones, stack them up against each other and try to better understand how they differ from each other. ### Kubernetes - The market leader Similar to how Docker became the de-facto for containerization, the industry has found Kubernetes to rule the container orchestration landscape. That's why most major cloud providers have started to offer managed Kubernetes service as well. It's an open-source software that has become the gold standard for orchestrating containerized workloads in private, public, and hybrid cloud environments. Initially developed by engineers at Google, who distilled years of experience in running production workloads at scale into Kubernetes. It was open-sourced in 2014 and has since been maintained by CNCF (Cloud Native Computing Foundation). It's often abbreviated as k8s which is a numeronym (starting with the letter "k" and ending with "s" with 8 other characters in between). Managing containers at scale is commonly referred to as quite challenging, why is that? Running a single Docker container on your laptop may seem trivial but doing that for a large number of containers across multiple hosts in an automated fashion ensuring zero downtime isn't as trivial. Let's take an example of a Netflix-like video-on-demand platform consisting of 100+ microservices resulting in 5000+ containers running atop 100+ VMs of varying sizes. Different teams are responsible for different microservices. They follow continuous integration and continuous delivery (CI/CD) driven workflow and push to production multiple times a day. The expectation from production workloads is to be always available, scale up and down automatically if demand changes, and recover from failures when encountered. In situations like these, the utility of container orchestration tools really shine. Tools like Kubernetes allow you to abstract away the underlying cluster of virtual or physical machines into one unified blob of resources. Typically they expose an API, using which you can specify how many containers you'd like to deploy for a given app and how they should behave under increased load. API-first nature of these tools allows you to automate deployment processes inside your CI pipeline, giving teams the ability to iterate quickly. Being able to manage this kind of complexity in a streamlined manner is one of the major reasons why tools like Kubernetes have gained such popularity. #### Underlying architecture & objects To understand the Kubernetes' view of the world, we need to familiarize ourselves with cluster architecture first. Kubernetes cluster is a group of physical or virtual machines which is divided into two high-level components, control plane and worker nodes. - **Control plane** - Act as the brain for the entire cluster, responsible for accepting user instructions, health checking all servers, deciding how to best schedule workloads, and orchestrating communication between components. Constituents include components like kube-apiserver, etcd, kube-scheduler, kube-controller-manager and cloud-controller-mananger. - **Worker nodes** - These are machines responsible for accepting instructions from the control plane and running containerized workloads. Each machine runs a kubelet, kube-proxy and container runtime. Now that we have some know-how of Kubernetes architecture, the next milestone in our journey is understanding the Kubernetes object model. Kubernetes has a few abstractions that make up the building blocks of any containerized workload. We'll go over a few different types of objects available in Kubernetes that you are more likely to interact with: - **Pod** - It is the smallest deployable unit of computing in Kubernetes hierarchy. It can contain one or more tightly coupled containers sharing environment, volumes, and IP space. Generally, it is discouraged for users to manage pods directly, instead, Kubernetes offers higher-level objects (deployment, statefulset & daemonset) to encapsulate that management. - **Deployment** - High-level object designed to ease the life cycle management of replicated pods. Users describe a desired state in the deployment object and the deployment controller changes the actual state to match the desired state. Generally, this is the object users interact with the most. It is best suited for stateless applications. - **Stateful Set** - You can think of it as a specialized deployment best suited for stateful applications like a relational database. They offer ordering and uniqueness guarantees. - **Daemon Set** - You can think of it as a specialized deployment when you want your pods to be on every node (or a subset of it). Best suited for cluster support services like log aggregation, security, etc. - **Secret & Config Map** - These objects allow users to store sensitive information and configuration respectively. These can then be exposed to certain apps thus allowing for more streamlined configuration and secrets management - **Service** - This object groups a set of pods together and makes them accessible through DNS within the cluster. Different types of services include NodePort, ClusterIP, and LoadBalancer. - **Ingress** - Ingress object allows for external access to the service in a cluster using an IP address or some URL. Additionally, it can provide SSL termination and load balancing as well - **Namespace** - This object is used to logically group resources inside a cluster Note: There are other objects like Replication Controller, Replica Set, Job, Cron Job, etc. that we have deliberately skipped for simplicity's sake. ### Docker Swarm - Lightweight alternative potentially approaching end of life Let's differentiate between Docker and Docker Swarm first. Docker is a container runtime comparable with rkt. Whereas Docker Swarm is a cluster management and orchestration tool embedded in the Docker Engine comparable with Kubernetes and likes. As compared to Kubernetes, it's a slightly less extensible and complex tool best suited for people who want an easier path to deploying containers. On a higher level, you'll notice a lot of similarities when it comes to the architecture of both the tools. An important thing to note here is that after Mirantis acquired Docker Enterprise, in late 2019, they announced that primary orchestrator going forward would be Kubernetes. They'll support Swarm for atleast two years and will work on making transition easier to Kubernetes. Does this mean Docker Swarm is dead and we shouldn't even talk about it? Not really! As of now, all this means is we won't be seeing many Docker Swarm as a service options out there. However for simpler use-cases, it still is a viable option owing to it's lightweight and simple nature. #### Underlying architecture To understand Docker Swarm's view of the world, we need to familiarize ourselves with cluster architecture first. Swarm by itself is a group of physical or virtual machines which is divided into two high-level components, manager node and worker nodes. - **Manager node** - Similar to Kubernetes Control Plane, it is responsible for receiving service definition from user and dispatching instructions to worker nodes on how to run that service. Additionally it also performs the orchestration and management functions necessary to sync actual state with desired state of the cluster. Manager nodes elect a single leader to conduct orchestration tasks. - **Worker node** - Similar to Kubernetes Worker nodes, it receives and execute tasks dispatched from manager nodes. An agent runs on each worker node and reports back to manager node on the assigned tasks so that manager can maintain the desired state of each worker. Now that we have some know-how of the architecture, let's get into object-level constructs of Docker Swarm. - **Task** - A task carries a Docker container and the commands to run inside the container, it is the atomic unit of scheduling within a swarm. When we declare a desired state of a service, orchestrator realizes the desired state by scheduling tasks. If the task fails the orchestrator removes the task and its container and then creates a new task to replace it according to the desired state specified by the service. - **Service** - Service is the definition of tasks to be executed on the nodes. When creating a service, user specifies which image to user and which commands to execute inside running containers. There are two types of service, replicated and global. Similar to Kubernetes deployments, in replicated services model, manager spins up specified number of replica tasks among nodes. Similar to Kubernetes daemon set, for global services, swarm runs one task for the service on every available node. - **Load balancer** - Swarm manager uses ingress load balancing to expose services to the outside world. External components like cloud load balancers can access a given service on its port while swarm uses internal load balancing to distribute requests among services within the cluster. ### Proprietary offerings by public cloud providers Just like open-source space, cloud orchestration tools have a pretty competitive propriety space mostly dominated by public cloud providers like Amazon Web Services (AWS), Google Cloud Platform (GCP) and Microsoft Azure. These cloud providers also offer a managed version of Kubernetes. What that means is that the provider is responsible for managing and maintaining the cluster. This reduces maintenance and management overhead. For the purposes of this section, we'll focus on proprietary offerings only. - **AWS Elastic Container Service** - ECS is a fully managed container orchestration service from AWS which has deep integrations with other AWS services like Route53, Secret Manager, IAM, CloudWatch etc. If offers two ways to run workloads, one is on EC2 virtual machines and second more recent one is Fargate which brings serverless capabilities to ECS. Users define configuration of their applications as task definitions and then provide these definitions as JSON documents to AWS Console or CLI interface. Depending on configuration and which mode you selected (EC2 or Fargate), ECS schedules the tasks composing services appropriately onto the cluster and monitors them constantly to maintain the desired state. - **GCP Cloud Run** - Cloud run is a fully managed serverless container orchestration platform. It abstracts away all infrastructure management by adopting the serverless model for containers. What that means if that your application can scale down to zero and you don't pay anything when there's no traffic and on the other hand it can scale upto millions of requests almost instantaneously. As a consumer of Cloud Run, all you need to do is provide the platform your Docker containers and it'll take care of the rest, which is very convenient as all the complexity has been automated or abstracted away. Under the hood, it runs [Knative](https://knative.dev/), which is a Kubernetes-based platform to deploy and manage modern serverless workloads. - **Azure Container Instance** - Container Instances is Microsoft Azure's answer to running containers on-demand in a serverless fashion. It is comparable to AWS Fargate and GCP Cloud Run. Since this serverless model allows you to not worry about underlying infrastructure and just focus on application logic, a lot of management overhead is cut down. All you need is a Docker container and specify some based configuration and the platform handles the rest for you, including things like provision resources, scaling containers up and down, necessary networking, health monitoring to name a few. ## Which one is right for me? There's no one size fits all when it comes to container orchestration tools. Choosing the right tool for the job is very use-case dependent. If you want to run your containerized apps with straight forward needs, perhaps your best bet is to use one of the serverless containers offerings out there like AWS Fargate, Google Cloud Run or Azure Container Instances. However if your needs require fine-grain control, customization and flexibility, perhaps a managed version of Kubernetes in form of AWS EKS, GCP GKE or Azure AKS are better suited for your use-case as they reduce the overhead of provisioning and running a Kubernetes cluster while ensuring smooth intra and inter cloud integrations. Lastly if your use-case has strict data residency and sovereignty constraints and you are required to run containerized workloads in an on-prem or private cloud setting, perhaps self-managed Kubernetes is the fore runner among all container orchestration tools. ### Introductory Guide: What is Kubernetes? URL: https://sarmadsaleem.com/thoughts/what-is-kubernetes/ Date: 2020-09-15 Tags: devops, kubernetes, containers It's easy to get lost in today's continuously changing landscape of cloud native technologies. The learning curve from a beginner's perspective is quite steep, and without proper context it becomes increasingly difficult to sift through all the buzzwords. If you have been developing software, chances are you may have heard of Kubernetes by now. Before we jump into what Kubernetes is, it's essential to familiarize ourselves with containerization and how it came about. In this guide, we are going to paint a contextual picture of how deployments have evolved, what's the promise of containerization, where Kubernetes fits into the picture, and common misconceptions around it. We'll also learn the basic architecture of Kubernetes, core concepts, and some examples. Our goal from this guide is to lower the barrier to entry and equip you with a mind map to navigate this landscape more confidently. ## Evolution of the deployment model This evolution can be categorized into three rough categories, namely traditional, virtualized, and containerized deployments. Let's briefly touch upon each to better actualize this evolution. ### Bare metal Some of us are old enough to remember the archaic days when the most common way of deploying applications was on in-house physical servers. Cloud wasn't a thing yet, organizations had to plan server capacity to be able to budget for it. Ordering new servers was a slow and tedious task that took weeks of vendor management and negotiations. Once shiny new servers did arrive, they came with the overhead of setup, deployment, uptime, maintenance, security, and disaster recovery. From a deployment perspective, there was no good way to define resource boundaries in this model. Multiple applications, when deployed on the same server, would interfere with each other because of lack of isolation. That forced deployment of applications on dedicated servers, which ended up being an expensive endeavor making resource utilization the biggest inefficiency. In those days, companies had to maintain their in-house server rooms and deal with prerequisites like air conditioning, uninterrupted power, and internet connectivity. Even with such high capital and operating expenses, they were limited in their ability to scale as demand increased. Adding additional capacity to handle increased load would involve installing new physical servers. ### Virtual machines Enter virtualization. This solution is a layer of abstraction on top of physical servers, such that it allows for running multiple virtual machines on any given server. It enforces a level of isolation and security, allowing better resource utilization, scalability, and reduced costs. It allows us to run multiple apps, each in a dedicated virtual machine offering complete isolation. If one goes down, it doesn't interfere with the other. Additionally, we can specify resource budgets for each. For example, allocate 40% of physical server resources to VM1 and 60% to VM2. Okay, so this addresses isolation and resource utilization issues but what about scaling with increased load? Spinning a VM is way faster than adding a physical server. However scaling of VMs is still bound by available hardware capacity. This is where public cloud providers come into the picture. They streamline the logistics of buying, maintaining, running, and scaling servers against a rental fee. This means organizations don't have to plan for capacity beforehand. This brings down the capital expense of buying the server and operating expense of maintaining it significantly. ### Containers If we have already addressed the issue of isolation, resource utilization, and scaling with virtual machines, then why are we even talking about containers? Containers take it up a notch. You can think of them as mini virtual machines that, instead of packaging a full-fledged operating system, try to leverage the underlying host OS for most things. Container-based virtualization guarantees higher application density and maximum utilization of server resources. An important distinction between virtual machines and containers is that VM virtualizes underlying hardware whereas the container virtualizes the underlying operating system. Both have their use cases, in fact, many container deployments use VM as their host operating system rather than running directly on bare metal. The emergence of Docker engine accelerated the adoption of this technology. It has now become the defacto standard to build and share containerized apps - from desktop to the cloud. Shift towards microservices as a superior approach to application development is another important factor that has fueled the rise of containerization. ## Demystifying container orchestration While containers by themselves are extremely useful, they can become quite challenging to deploy, manage, and scale across multiple hosts in different environments. Container orchestration is another fancy word for streamlining this process. As of today, there are several open-source and proprietary solutions to manage containers out there. ### Open-source landscape If we look at the open-source landscape, some notable options include: - Kubernetes - Docker Swarm - Apache Marathon on Mesos - Hashicorp Nomad - Titus by Netflix ### Proprietary landscape On the other hand, if we look at the propriety landscape, most of it is dominated by major public cloud providers. All of them came up with their home-grown solution to manage containers. Some of the notable mentions include: - Amazon Web Services (AWS) - Elastic Beanstalk - Elastic Container Service (ECS) - Fargate - Google Cloud Platform (GCP) - Cloud Run - Compute Engine - Microsoft Azure - Container Instances - Web Apps for Containers ### Gold standard Similar to how Docker became the de-facto for containerization, the industry has found Kubernetes to rule the container orchestration landscape. That's why most major cloud providers have started to offer managed Kubernetes service as well. We'll learn more about them later in the ecosystem section. ## What exactly is Kubernetes? Kubernetes is open-source software that has become the defacto standard for orchestrating containerized workloads in private, public, and hybrid cloud environments. It was initially developed by engineers at Google, who [distilled years of experience](https://queue.acm.org/detail.cfm?id=2898444) in running production workloads at scale into Kubernetes. It was open-sourced in 2014 and has since been maintained by CNCF (Cloud Native Computing Foundation). It's often abbreviated as k8s which is a [numeronym](https://en.wikipedia.org/wiki/Numeronym) (starting with the letter "k" and ending with "s" with 8 other characters in between). Managing containers at scale is commonly referred to as quite challenging, why is that? Running a single Docker container on your laptop may seem trivial (we'll see this in the example below) but doing that for a large number of containers across multiple hosts in an automated fashion ensuring zero downtime isn't as trivial. Let's take an example of a Netflix-like video-on-demand platform consisting of 100+ microservices resulting in 5000+ containers running atop 100+ VMs of varying sizes. Different teams are responsible for different microservices. They follow [continuous integration and continuous delivery](https://m.youtube.com/watch?v=scEDHsr3APg) (CI/CD) driven workflow and push to production multiple times a day. The expectation from production workloads is to be always available, scale up and down automatically if demand changes, and recover from failures when encountered. In situations like these, the utility of container orchestration tools really shine. Tools like Kubernetes allow you to abstract away the underlying cluster of virtual or physical machines into one unified blob of resources. Typically they expose an API, using which you can specify how many containers you'd like to deploy for a given app and how they should behave under increased load. API-first nature of these tools allows you to automate deployment processes inside your CI pipeline, giving teams the ability to iterate quickly. Being able to manage this kind of complexity in a streamlined manner is one of the major reasons why tools like Kubernetes have gained such popularity. ### Kubernetes Architecture To understand the Kubernetes' view of the world, we need to familiarize ourselves with cluster architecture first. Kubernetes cluster is a group of physical or virtual machines which is divided into two high-level components, control plane and worker nodes. It's okay if some of the terminologies mentioned below don't make much sense yet. - **Control plane** - Acts as the brain for the entire cluster. In that, it is responsible for accepting instruction from the users, health checking all servers, deciding how to best schedule workloads, and orchestrating communication between components. Constituents of control plane include: - **kube-apiserver** - Responsible for exposing Kubernetes API. In other words, this is the gateway into Kubernetes - **etcd** - Distributed, reliable key-value store that is used as a backing store for all cluster data - **kube-scheduler** - Responsible for selecting a worker node for newly created pods (also known as scheduling) - **kube-controller-manager** - Responsible for running controller processes like Node, Replication, Endpoints, etc. These controllers will start to make more sense after we discuss k8s objects - **cloud-controller-manager** - Holds cloud-specific control logic - **Worker nodes** - These are machines responsible for accepting instructions from the control plane and running containerized workloads. Node has the following sub-components: - **kubelet** - An agent that makes sure all containers are running in any given pod. We'll get to what that means in a bit. - **kube-proxy** - A network proxy that is used to implement the concept of service. We'll get to what that means in a bit. - **Container runtime** - This is the software responsible for running containers. Kubernetes supports Docker, containerd, rkt to name a few. The key takeaway here is that the control plane is the brain responsible for accepting user instructions and figuring out the best way to execute them. Whereas worker nodes are machines responsible for obeying instructions from the control plane and running containerized workloads. ### Kubernetes Objects Now that we have some know-how of Kubernetes architecture, the next milestone in our journey is understanding the Kubernetes object model. Kubernetes has a few abstractions that make up the building blocks of any containerized workload. We'll go over a few different types of objects available in Kubernetes that you are more likely to interact with: - **Pod** - It is the smallest deployable unit of computing in Kubernetes hierarchy. It can contain one or more tightly coupled containers sharing environment, volumes, and IP space. Generally, it is discouraged for users to manage pods directly, instead, Kubernetes offers higher-level objects (deployment, statefulset & daemonset) to encapsulate that management. - **Deployment** - High-level object designed to ease the life cycle management of replicated pods. Users describe a desired state in the deployment object and the deployment controller changes the actual state to match the desired state. Generally, this is the object users interact with the most. It is best suited for stateless applications. - **Stateful Set** - You can think of it as a specialized deployment best suited for stateful applications like a relational database. They offer ordering and uniqueness guarantees. - **Daemon Set** - You can think of it as a specialized deployment when you want your pods to be on every node (or a subset of it). Best suited for cluster support services like log aggregation, security, etc. - **Secret & Config Map** - These objects allow users to store sensitive information and configuration respectively. These can then be exposed to certain apps thus allowing for more streamlined configuration and secrets management - **Service** - This object groups a set of pods together and makes them accessible through DNS within the cluster. Different types of services include NodePort, ClusterIP, and LoadBalancer. - **Ingress** - Ingress object allows for external access to the service in a cluster using an IP address or some URL. Additionally, it can provide SSL termination and load balancing as well - **Namespace** - This object is used to logically group resources inside a cluster Note: There are other objects like Replication Controller, Replica Set, Job, Cron Job, etc. that we have deliberately skipped for simplicity's sake. ## Show me by example Now that we have touched upon some of the most commonly used Kubernetes objects that act as the building blocks of containerized workloads, let's put it to work. For this example we'll do the following: - Setup prerequisites like Docker, Kubernetes & kubectl - Create a very simple hello world node app - Containerize it using Docker and push it to a public docker registry - Use some of the above-explained Kubernetes objects to declare our containerized workload specification ### Setup This guide assumes that you are on macOS, have Docker Desktop installed and running. It comes with a standalone Kubernetes instance which is a single-node cluster, an excellent choice these days to run Kubernetes locally. Additionally, you must also have kubectl installed. It is a command-line tool that allows us to run commands against Kubernetes clusters. The easiest way to get up and running on Mac is to use Homebrew package manager like so: ```bash # Install docker brew cask install docker # Install kubectl brew install kubectl ``` Setup instructions and source code for this exercise can be found [on GitHub](https://github.com/sarmadsaleem/scout-apm-intro-to-k8s). ### Sample hello world app Here we have a very simple hello world app written in NodeJS. It creates an HTTP server, listens on port 3000, and responds with "Hello World". ### Containerize sample app To dockerize our app, we'll need to create a Dockerfile. It describes how to assemble the image. Let's use Docker CLI to build and test the image: ```bash # Build docker image docker build -t sarmadsaleem/scout-apm:node-app . # Run container based on docker image docker run -p 3000:3000 -it --rm sarmadsaleem/scout-apm:node-app # Verify functionality curl http://localhost:3000 ``` Now that we have verified our container works fine locally, let's push this docker image to a public registry. For this example, we'll be using a public repository on Docker Hub. ```bash # Docker hub login docker login --username sarmadsaleem # Push image to Docker Hub repository docker push sarmadsaleem/scout-apm:node-app ``` At this point, we have packaged our node app in a docker container and made it public in the form of an image. Anyone can pull it from our public repository and run it anywhere. ### Define workload specification using Kubernetes objects With our dockerized sample app ready, the only thing remaining to do is declare our desired state of workload using Kubernetes objects. We'll be dealing with Namespace, Deployment & Service in this example. Typically Kubernetes manifests are declared in YAML files that describe the desired state. It is then passed to the Kubernetes API using kubectl. It's okay if some of the things in this manifest don't make sense yet. The key takeaway is that we have declared specification for our containerized workload using Kubernetes objects. A Namespace object is just a wrapper to group things together. Deployment object does the heavy lifting of creating pods (which hold containers), maintaining a specified number of replicas, and managing their lifecycle. Service object streamlines network access to all pods. Now that we have the manifest ready, let's select the correct context for kubectl and apply the manifest: ```bash # Switch kube context kubectl config use-context docker-desktop # Apply kube manifests kubectl apply -f path/to/app.yaml ``` That was fast, what just happened? Kubernetes accepted our declared manifest and tried to execute it. In doing so it created a namespace, a bunch of pods, a replica set, a deployment, and a service. We should be able to verify all of them by doing so: ```bash # Get state of kube objects kubectl get pod,replicaset,deployment,service -n scout-apm ``` Does this mean our sample app is running? Yes! We should be able to reach our node app: ```bash # Get service port kubectl get service -n scout-apm # Verify functionality curl http://localhost: # Let's clean up after ourselves kubectl delete -f path/to/app.yaml ``` In a more production-ready setup, you'll probably have to deal with considerations like optimizing your docker images, automating deployments, setting up health checks, securing the cluster, managing incoming traffic over SSL, instrumenting your apps for observability, etc. The goal of this simple exercise was to jump from theory into a practical playground where we can see Docker in action, learn how to interact with Kubernetes, and deploy a sample workload. ## Basic features Let's touch upon some of the major Kubernetes features: - **Self-healing** - One of the biggest selling points of Kubernetes is that it provides self-healing capabilities thus making your applications more robust and resilient. - **Horizontal scaling & load balancing** - Kubernetes advertises itself to be planet scale. In that, it provides multiple ways to scale applications up and down. These horizontal scaling capabilities stretch from containers to underlying nodes. - **Automated rollouts and rollbacks** - It provides the ability to progressively roll out changes while monitoring application health. If anything goes wrong, changes are rolled back automatically to the last working version. - **Service discovery** - Streamlines service discovery mechanism by giving pods their IP and consolidating all those IPs behind a service so that a single DNS name can be used to reach all pods and load can be balanced across them. - **Secret and configuration management** - Inline with [Twelve Factor App](https://12factor.net/), Kubernetes provides tools to externalize secrets and configurations. This means sensitive information and configuration can be updated without rebuilding your image and the risk of embedding secrets in images is neutralized. - **Storage orchestration** - Provides ways to mount storage systems like local storage or something from a public cloud provider. This mounted storage is then made available to workloads via a unified API. - **Batch execution** - In addition to long-lived workloads, it also provides ways to manage short-lived batch workloads ## Kubernetes ecosystem In the past few years, the Kubernetes ecosystem has grown exponentially. The popularity of Kubernetes has led to greater adoption thus inspiring innovation in different verticals like the following: - **Cluster provisioning** - Managed Kubernetes services like EKS, GKE, and AKS are rather recent additions to the landscape. Most of these platforms provide API, CLI, and GUI interfaces to provide clusters. Before these, provisioning a new cluster and maintaining your control plane used to involve non-trivial effort. Some of the popular tools to provision and bootstrap clusters include Kops, Kubeadm, Kubespray, and Rancher. - **Managed control plane** - These days most cloud providers offer a managed Kubernetes offering. What that means is that the provider is responsible for managing and maintaining the cluster. This reduces maintenance and management overhead. However, this is very use-case specific, in some cases you'd like to self-manage it to have better flexibility. Some of the managed Kubernetes offerings include EKS, GKE, and AKS. - **Cluster management** - Most common way to interact and manage clusters is still kubectl however this vertical has been evolving at a rapid pace. Helm has emerged as Kubernetes goto package manager, think of it as Homebrew for your cluster. To streamline the management of YAML templates, Kustomize is leading the charge and has already been integrated with kubectl. - **Secrets management** - Built-in Kubernetes secrets are base64 encoded, not encrypted. Companies find themselves outgrow the built-in functionality quite quickly and turn to more sophisticated solutions like Sealed Secrets, Vault and Cloud Managed Secret Stores (AWS Secrets Manager, Cloud KMS, Azure Key Vault, etc.) - **Monitoring, logging & tracing** - Kubernetes provides some basic features around monitoring, logging, and tracing but as soon as you get into running multiple microservices, these features can fall short. A common monitoring stack in the industry today is Prometheus, Alertmanager & Grafana. When it comes to log aggregation, ELK stack has gained popularity. Lastly for distributed tracing, Jaeger and Zipkin are the two open-source tools that seem to have a high adoption rate. - **Development toolkits** - Oketo, Tilt, and Garden are some of the famous development toolkits that help streamline development workflow while working with multiple microservices in a cloud-native context. - **Infrastructure as code** - Infrastructure as code (IaC) is the management of infrastructure in a descriptive model. Tools like Terraform and Pulumi have become popular choices and have support for provisioning Kubernetes related resources using code. There are several benefits to this strategy include version-controlled infrastructure, audibility, static analysis, and automation to name a few. - **Traffic management** - Typically Kubernetes uses Ingress controllers to expose services to the outside world. This is where SSL termination can also happen. Some of the popular ingress controllers include NGINX, AWS ALB Ingress Controller, Istio, Kong, and Traefik. Additionally, new constructs like API Gateway & Service Mesh have been introduced into the ecosystem lately and the boundary between how one can manage ingress traffic (north to south) and internal traffic (east to west) has become increasingly blurry. ## Common Questions ### Is Kubernetes free? Open-source version of Kubernetes itself is free to download, build, extend, and use for everyone, so there are no costs associated with the software itself. Typically organizations run their Kubernetes clusters in public, private, hybrid, or multi-cloud environments, in those cases, they have to pay for the underlying resources. ### Who created Kubernetes? Kubernetes originated at Google and distilled years of experience in running production workloads at scale. It was founded by Joe Beda, Brendan Burns, Craig McLuckie who were quickly joined by other Google engineers in their endeavor. It was later donated to Cloud Native Computing Foundation (CNCF) and is now being maintained by the foundation along with the open-source community under Apache License. ### What is the difference between Docker and Kubernetes? Docker is a container runtime meant to run on a single node whereas Kubernetes is a container orchestration tool meant to run across a cluster of nodes. They are not opposing technologies. They complement one another. ### How do you upgrade Kubernetes? Upgrading Kubernetes version is a common practice to keep up with the latest security patches, new features, and bug fixes. This process is typically dictated by the tool you used to provision the cluster. If it's a managed control plane, the cloud provider exposes an API to trigger the upgrade. If it's a self-managed control plane, bootstrapping tools like kops and kubeadm simplify this workflow. ### Is Kubernetes the best way to run containers in production today? This can be a controversial one. Kubernetes is arguably the most feature-complete container orchestration tool with a vibrant community and buzzing ecosystem. Is it the best way to run containers in production today? That depends on your use-case. Perhaps in some cases, you can get by using a PaaS solution like Heroku. Alternatively, you may be able to leverage new-age serverless container services like Google Cloud Run or AWS Fargate. In other cases where you want complete control and flexibility over your workloads, Kubernetes may be the front runner among container orchestration tools. --- ## Work Experience ### BCG X (2018 - Now) Principal & Venture CTO, Berlin, Germany Leading founding teams. Healthcare, media, fintech, and beyond. - **Clinical Trials Platform**: Medical imaging workflows for clinical trials. Multi-tenant, compliance-first, AI-ready. - **AI Endoscopy Platform**: End-to-end workflow management meets AI marketplace for real-time detection, diagnostics, and quality scoring on live video feeds. - **Travel Experience Marketplace**: Travel experiences marketplace for a major airline. ML-driven personalization. PoC, team setup, foundational architecture. - **Neobank Platform**: Digital-first banking platform for a major financial consortium. PoC, team setup, foundational architecture. - **Source2Sea**: Maritime procurement marketplace. Zero to MVP, team setup, foundational architecture. - **Hybrid Recommendation Engine**: Unified recommender blending collaborative filtering, content signals, and editorial curation for streaming. ### Bogo (2015 - 2021) Fractional CTO, Karachi, Pakistan Scaled consumer lifestyle platform from zero to exit. - **Core Platform**: API services powering web, mobile, and backoffice clients. Built to handle deal-hungry users at scale. - **Consumer Website**: Browse deals, redeem offers, repeat. The web experience. - **Merchant Backoffice**: Deal configuration, redemption tracking, analytics. Where merchants watched the magic happen. ### Al Jazeera Media Network (2016 - 2018) Senior Software Engineer, Doha, Qatar Shipped election dashboards and data-driven interactives. - **US Election 2016**: Live results. Real-time visualizations. - **Pakistan Election 2018**: Live results. Constituency-level tracking. - **India Election 2019**: Live results. 900M+ eligible voters. ### Dawn Media Group (2015 - 2016) Digital Project Manager, Karachi, Pakistan Led digital transformation - web, apps, video, ad monetization. - **Dawn Classifieds**: Online classifieds and payments platform. Built and monetized. - **Dawn Obituaries**: Digital obituary platform with integrated payments. - **Dawn TV App**: Mobile app for Dawn's Urdu news channel. Ad-monetized video delivery. ### Aga Khan University (2013 - 2015) Programmer Analyst, Karachi, Pakistan Developed ERP systems, campus workflows, and integrations. - **Online Admissions**: Replaced paper applications with a web portal on Oracle PeopleSoft Campus Solutions. Thousands of students, zero lost forms. - **Databook Integration**: BI dashboards that turned campus data into decisions leadership actually acted on. - **Legacy Digitization**: Decades of academic records, digitized. Ctrl+F beats filing cabinets. ### Xipnox (2010 - 2013) Co-Founder & CTO, Karachi, Pakistan Bootstrapped design and dev studio for SMEs. - **Law Enforcement Dispatch**: Dispatch system for local law enforcement. Web and mobile. Real emergencies, real stakes. - **Linkagoal**: Social network for setting and tracking goals. Frontend build. - **Virtual Desktop Delivery**: Enterprise virtual desktop infrastructure built on Ulteo. Thin clients, centralized management. --- ## Tools & Setup ### Hardware - **development machine**: MacBook Pro 14" (M4 Max, 36GB RAM) - **gaming rig**: Legion 5 (RTX 3060, Ryzen 5, 32GB RAM) - **monitor**: Gigabyte M27Q (27", 1440p, 170Hz, KVM) - **desk**: FlexiSpot E7 - **keyboards**: WOBKEY Zen 65, Magic Keyboard with Touch ID - **pointing devices**: Magic Trackpad, Logitech G Pro X Superlight - **audio**: AirPods Pro 2, Bose QC 35 II, Beyerdynamic DT 770 Pro, Rode PodMic USB ### Software - **editor**: Zed, Cursor, VS Code - **terminal**: Ghostty - **ai**: Claude, ChatGPT, Perplexity, Granola, Wispr Flow, Google Notebook LM - **swe agents**: Claude Code, Codex, Opencode, Cursor, Lovable - **design**: Figma - **notes & tasks**: Notion, Linear, Apple Notes - **utilities**: Maccy, Zappy, Magnet ### Coffee - **espresso machine**: La Marzocco Linea Micra - **grinder**: Timemore Sculptor 064S - **beans**: Codos, Coffee Circle, 19grams, Studio Natura