# Usamah Zaheer > Machine Learning Software Engineer at Arm in Cambridge, UK. Writes about edge ML inference, ML compilers, robotics and AI systems. Everything at https://www.usamah.me, as one file. Last updated 2026-09-15. The canonical HTML lives at https://www.usamah.me; any page is also available on its own as Markdown at its path with `.md` appended. The book is a separate site: https://ai.usamah.me (https://ai.usamah.me/llms-full.txt is its equivalent of this file). --- # Usamah Zaheer Machine Learning Software Engineer at Arm. Cambridge, United Kingdom. > Machine Learning Software Engineer at Arm in Cambridge, UK. Writes about edge ML inference, ML compilers, robotics and AI systems. Machine Learning Engineer with 4+ years of experience in deep learning, computer vision, and edge computing. Strong background in end-to-end ML development, model optimisation, and MLOps practices. Usamah Zaheer is a Machine Learning Software Engineer at Arm, where he designs and optimises ML infrastructure for Arm architectures. He is pursuing an MS in Artificial Intelligence at the University of Texas at Austin and holds an MS in Embedded Systems from the University of Leicester. His work spans deep learning, computer vision, edge computing, robotics, and MLOps across companies including Arm, Dyson, and the University of Leicester. ## Book ### How to Make Your Model Fast *A Systems View of Efficient Machine Learning, from Silicon to Agents* A book in fourteen parts on making models run inside real budgets for latency, memory, power and cost. Rooflines and Arm vector units, kernels and compilers, quantisation and compression, then vision, on-device language models, robotics, profiling, serving and agents. Free to read at https://ai.usamah.me. ## Experience ### Machine Learning Software Engineer, Arm Mar 2025 – Present - Designed and optimised ML infrastructure to analyse and enhance performance of models and systems on Arm architectures. - Optimised ML compilers and libraries for Arm, improving inference performance and reducing latency. - Conducted deep kernel-level analysis to identify and eliminate inference bottlenecks. - Profiled and optimised runtime performance of ML models; developed scalable benchmarking solutions across cloud and edge environments. - Built automated pipelines for data collection, preprocessing, and model evaluation, streamlining production workflows. - Led cross-functional collaborations to align ML infrastructure with organisational goals, owning a new project from inception. Stack: PyTorch, TensorFlow, JAX, FBGEMM, KleidiAI, ACL, ArmNN, OneDNN. ### Robotics Software Engineer, Dyson Sep 2022 – Aug 2024 - Developed CNN algorithms for segmentation, object detection, and classification, applying quantisation, pruning, and knowledge distillation for deployment on robot hardware. - Architected evaluation tools and robotics algorithms for planning and navigation using C++. Conducted log analysis, debugging, and on-robot testing. - Deployed models on diverse hardware and edge devices, optimising through profiling and bottleneck analysis using CUDA and cuDNN. - Developed a VLM solution that saved over £100,000 and boosted productivity by 20x. - Streamlined the ML lifecycle with model versioning, monitoring, and automated deployment using CI/CD pipelines. - Presented complex ML projects to senior leaders and the CEO. Stack: PyTorch, MXNet, ONNX, CUDA, cuDNN, C++. ### ML Research Assistant, University of Leicester Mar 2021 – Aug 2022 - Spearheaded development of an end-to-end automated ML pipeline for processing high-resolution data in real-time. - Integrated cutting-edge CNNs in PyTorch and TensorFlow for high-resolution satellite imagery analysis. - Led development of ML systems utilising Random Forest and SVM for predictive modelling. - Deployed AI solutions in cloud environments with Docker and Kubernetes. Stack: PyTorch, TensorFlow, Scikit-learn, Docker, Kubernetes, R. ## Projects ### AI Agent Systems, Stealth Startup Aug 2024 – Feb 2025 - Led design and implementation of AI agents for SDRs, integrating NLP and LLMs for accurate and scalable solutions. - Managed cloud infrastructure on Vertex AI with Databricks and Snowflake, utilising RAG to enhance AI response precision. - Orchestrated ML workflows and semantic search using LangChain and LlamaIndex. - Oversaw containerised deployment with Docker and Kubernetes, implementing MLOps practices with MLflow. ### ML Model Deployment App - Developed an Android application for deploying ML models on edge devices for object detection using YOLO, Mask R-CNN, and SSD. - Optimised using TensorFlow Lite with quantisation and pruning for low-latency, on-device inference. ### 360 Vision Navigation, Dyson - Developed a robot utilising 26 sensors with SLAM technology and 360-degree vision for autonomous navigation. - Advanced path planning algorithms and developed integration and unit tests in C++ and Python. ### Air Purifiers Embedded Software, Dyson - Contributed to embedded software for all Dyson air purifiers including Pure Cool using C and Python. - Designed backbone logic behaviour system for hardware communication with FreeRTOS. Updated a fundamental library stack for 10+ projects. ### AI4EO & CNN Research, University of Leicester - Classified high-resolution satellite images for forest fire detection using CNNs, Random Forest, and SVM. - Evaluated CNN architectures for autonomous vehicle applications, optimised with Transfer Learning and TensorRT. Written up in full at https://www.usamah.me/projects (Markdown: https://www.usamah.me/projects.md). ## Education ### MS in Artificial Intelligence University of Texas at Austin, USA. Jan 2025 – Dec 2026. ### MS in Embedded Systems and Control Engineering University of Leicester, England, Distinction. Jan 2021 – Aug 2022. Thesis: Performance Evaluation of Deep Learning Techniques for Object Detection in Autonomous Vehicles. ### BTech in Electronics and Communication Engineering Jawaharlal Nehru Technological University, India, First Class. Aug 2016 – Oct 2020. ## Skills - **Languages:** Python, C++, C, Rust, SQL, MATLAB - **ML & Deep Learning:** PyTorch, TensorFlow, JAX, Scikit-learn, LLMs, VLMs, Multimodal Models, CNNs, Transformers - **Computing & Inference:** TensorRT, CUDA, cuDNN, ArmNN, LLVM, GGML, OpenMP, ONNX/Runtime, FBGEMM, KleidiAI - **Profiling & Debugging:** NVIDIA Nsight, PyTorch Profiler, TensorFlow Profiler, Valgrind, gprof, cProfile, Py-Spy - **Cloud & Deployment:** AWS, GCP, Vertex AI, Kubernetes, Docker, GitHub Actions, Jenkins, MLflow, KubeFlow, Databricks, Snowflake - **Data & Visualisation:** Pandas, Apache Spark, Matplotlib, Plotly, Streamlit, Grafana, Tableau, Gradio ## Leadership - Led an entire project from inception to completion independently at Arm. - Presented projects to the Dyson CEO and senior leadership. - Mentored undergraduates in robotics and machine learning. - Participated in 10+ hackathons and workshops. - Open-source contributor with a portfolio of projects on GitHub. ## Contact - Email: usamahzaheer155@gmail.com - LinkedIn: https://linkedin.com/in/usamahzaheer - GitHub: https://github.com/usamahz --- # About Usamah Zaheer ## At a glance - **Role:** Machine Learning Software Engineer at Arm - **Based in:** Cambridge, United Kingdom - **Focus:** Edge ML inference, ML compilers, model optimisation, computer vision and robotics - **Experience:** 4+ years across Arm, Dyson, University of Leicester - **Studying:** MS in Artificial Intelligence, University of Texas at Austin (expected Dec 2026) - **Wrote:** How to Make Your Model Fast, free at https://ai.usamah.me - **Contact:** usamahzaheer155@gmail.com I'm Usamah Zaheer, a Machine Learning Software Engineer at [Arm](https://www.arm.com) in Cambridge, UK, where I design and optimise ML infrastructure for Arm architectures. My work focuses on making deep learning models run faster and more efficiently - from optimising ML compilers and libraries to conducting deep kernel-level analysis that eliminates inference bottlenecks. Before Arm, I spent two years at [Dyson](https://www.dyson.com) as a Robotics Software Engineer. There, I developed CNN algorithms for segmentation, object detection, and classification - deploying them on robot hardware using techniques like quantisation, pruning, and knowledge distillation. I also built a VLM solution that saved over £100,000 and boosted productivity by 20x. One of the highlights was presenting my work directly to the CEO and senior leadership. My journey in ML research started at the University of Leicester, where I worked as an ML Research Assistant. I built end-to-end automated ML pipelines for processing high-resolution satellite imagery in real-time, integrating CNNs in PyTorch and TensorFlow. I also led development of predictive models using Random Forest and SVM, and deployed AI solutions in cloud environments with Docker and Kubernetes. I'm currently pursuing an MS in Artificial Intelligence at the University of Texas at Austin, building on my MS in Embedded Systems and Control Engineering from the University of Leicester (where I graduated with Distinction) and my BTech in Electronics and Communication Engineering from Jawaharlal Nehru Technological University. My technical toolkit spans the full ML stack: PyTorch, TensorFlow, JAX, CUDA, TensorRT, ArmNN, and C++ for performance-critical work. I'm passionate about the intersection of ML and hardware - making models not just accurate, but fast and deployable at the edge. Between Dyson and Arm, I led the design and implementation of AI agents at a stealth startup, integrating LLMs and NLP for scalable solutions using RAG, LangChain, and LlamaIndex. I managed cloud infrastructure on Vertex AI and oversaw containerised deployments with Docker and Kubernetes. Outside of work, I mentor undergraduates in robotics and machine learning, participate in hackathons, and contribute to open-source projects. You can find my work on [GitHub](https://github.com/usamahz), or connect with me on [LinkedIn](https://linkedin.com/in/usamahzaheer). I wrote a free book about all of this, [How to Make Your Model Fast](https://ai.usamah.me): fourteen parts on making models run inside real budgets for latency, memory, power and cost, from rooflines and Arm vector units through kernels, compilers, quantisation and compression to vision, on-device language models, robotics, profiling, serving and agents. Explore my [projects](/projects) to see what I've built, or check out my [blog](/blog) where I write about machine learning, edge computing, and software engineering. ## Expertise - **Edge ML inference:** Quantisation (PTQ, QAT, mixed-precision), operator fusion, memory planning and kernel optimisation for Arm Cortex-A and Cortex-M, Mali GPU and Ethos NPU. - **ML compilers and libraries:** ArmNN, Arm Compute Library (ACL), KleidiAI, FBGEMM, OneDNN, TensorRT and ONNX Runtime. - **Deep learning frameworks:** PyTorch, TensorFlow, JAX and ONNX. - **Computer vision:** CNNs, VLMs, object detection, segmentation, classification and knowledge distillation. - **AI agents:** LLM-based agent systems, retrieval-augmented generation, semantic search and cost-aware routing. - **MLOps and infrastructure:** Docker, Kubernetes, Vertex AI, MLflow, Databricks and CI/CD pipelines. - **Profiling:** NVIDIA Nsight, PyTorch Profiler, Valgrind, gprof and Arm Streamline. ## Frequently asked questions ### Who is Usamah Zaheer? Usamah Zaheer is a Machine Learning Software Engineer at Arm, based in Cambridge, UK. He has over 4 years of experience in deep learning, computer vision, edge computing, and MLOps. He previously worked at Dyson as a Robotics Software Engineer. ### What does Usamah Zaheer work on? Usamah works on designing and optimising ML infrastructure for Arm architectures. He specialises in ML compilers, inference performance, deep learning frameworks (PyTorch, TensorFlow, JAX), and deploying models on edge devices. ### Where did Usamah Zaheer study? Usamah is currently pursuing an MS in Artificial Intelligence at the University of Texas at Austin. He holds an MS in Embedded Systems and Control Engineering from the University of Leicester (Distinction) and a BTech in Electronics and Communication Engineering from Jawaharlal Nehru Technological University. ### Where has Usamah Zaheer worked? Usamah has worked at Arm (Machine Learning Software Engineer), Dyson (Robotics Software Engineer), the University of Leicester (ML Research Assistant), and a stealth startup (AI Agent Systems). He has also contributed to open-source projects. ### What book has Usamah Zaheer written? Usamah Zaheer wrote "How to Make Your Model Fast: A Systems View of Efficient Machine Learning, from Silicon to Agents", a free book in fourteen parts covering roofline analysis, Arm hardware, kernels, ML compilers, quantisation, compression, vision, on-device language models, robotics, profiling, serving and agents. It is free to read at https://ai.usamah.me. ### What does Usamah Zaheer write about? Usamah writes about edge ML inference, ML compilers, model optimisation, robotics and vision-language models, and the economics of AI systems. His posts are at https://www.usamah.me/blog and his book on efficient machine learning is at https://ai.usamah.me. ### How can I contact Usamah Zaheer? Email usamahzaheer155@gmail.com, or reach him on LinkedIn at https://linkedin.com/in/usamahzaheer. His code is on GitHub at https://github.com/usamahz. ## Contact - Email: usamahzaheer155@gmail.com - LinkedIn: https://linkedin.com/in/usamahzaheer - GitHub: https://github.com/usamahz --- # Usamah Zaheer's projects ## AI Agent Systems, Stealth Startup Aug 2024 – Feb 2025 Led the design and implementation of AI agents for Sales Development Representatives (SDRs), integrating NLP and LLMs to build accurate and scalable solutions. The system leveraged RAG (Retrieval-Augmented Generation) to enhance AI response precision by grounding outputs in domain-specific knowledge bases. Managed cloud infrastructure on Google Vertex AI, orchestrating data pipelines with Databricks and Snowflake. Built semantic search and ML workflow orchestration using LangChain and LlamaIndex, enabling the agents to retrieve and reason over large document corpora. Oversaw containerised deployment with Docker and Kubernetes, implementing MLOps practices with MLflow for experiment tracking, model versioning, and reproducible deployments. Stack: LLMs, RAG, LangChain, LlamaIndex, Vertex AI, Databricks, Snowflake, Docker, Kubernetes, MLflow. ## ML Model Deployment App Developed an Android application for deploying machine learning models directly on edge devices, enabling real-time object detection using YOLO, Mask R-CNN, and SSD architectures. The app was designed for low-latency, on-device inference without requiring cloud connectivity. Optimised models using TensorFlow Lite with quantisation and pruning techniques, significantly reducing model size and inference time while maintaining detection accuracy. The project demonstrated practical deployment of state-of-the-art detection models on resource-constrained mobile hardware. Stack: YOLO, Mask R-CNN, SSD, TensorFlow Lite, Android, Edge Deployment. ## 360 Vision Navigation, Dyson Developed a robot utilising 26 sensors with SLAM (Simultaneous Localisation and Mapping) technology and 360-degree vision for fully autonomous navigation. The system fused data from multiple sensor modalities to build real-time environmental maps and navigate complex indoor environments. Advanced path planning algorithms to optimise navigation efficiency and obstacle avoidance. Developed comprehensive integration and unit tests in C++ and Python, ensuring reliability and robustness of the navigation stack across diverse environments and edge cases. Stack: SLAM, Computer Vision, C++, Python, Sensor Fusion, Path Planning. ## Air Purifiers Embedded Software, Dyson Contributed to the embedded software powering all Dyson air purifiers, including the Pure Cool product line. Worked with C and Python to develop firmware that manages sensor data processing, device communication, and real-time control logic. Designed the backbone logic behaviour system for hardware communication using FreeRTOS, enabling reliable real-time task scheduling and inter-component messaging. Updated a fundamental library stack shared across 10+ projects, improving code reuse and maintainability across the Dyson embedded ecosystem. Stack: C, Python, FreeRTOS, Embedded Systems, Hardware Communication. ## AI4EO & CNN Research, University of Leicester Classified high-resolution satellite images for forest fire detection using CNNs, Random Forest, and Support Vector Machines. The research applied AI for Earth Observation (AI4EO), using deep learning to identify fire-affected regions from multispectral satellite imagery with high accuracy. Evaluated multiple CNN architectures for autonomous vehicle applications, benchmarking performance across detection accuracy, inference speed, and memory footprint. Optimised models with Transfer Learning and TensorRT to achieve real-time inference suitable for deployment on vehicle-mounted compute platforms. Stack: CNNs, PyTorch, TensorFlow, TensorRT, Transfer Learning, Remote Sensing. --- # We Can All Bake Bread Published 2026-09-14, updated 2026-09-15. 362 words. https://www.usamah.me/blog/we-can-all-bake-bread Every other day I see some version of this take: > AI means anyone can build their own software now, so nobody will pay for software, so SaaS companies are finished, software engineers are finished, and so on. It reminds me of bread. > We can all bake bread. We still buy it. The recipe is on the side of the flour bag. It's not like nobody knows how to do it. We just don't want to. Nobody wants to be the baker, and that's what you're actually paying for. I think software is the same. Vibe coding works. Something that used to take a week takes an afternoon. But nobody was ever paying a software company for the first version. They were paying for someone else to deal with everything after it. The 3am call when it breaks (it's always 3am). The library with the security hole that needs patching. The backups, and actually checking they restore. The customer whose export broke because there's a comma in a name. The payment provider retiring their API next year. > None of that got much cheaper. Only the first version did. Sure, some software will die. If your whole product is a form over a database with a login page for $29 a month, someone can make that by dinner now. Fair enough. But I think the threat is coming from somewhere else. A bakery doesn't go under because its customers learn to bake. It goes under because ovens got cheap and now there are ten bakeries on the same street, all selling the same sourdough. That's what's happening to software. The customers are all still there. It's just that a lot more people are selling to them now. A feature that took a team a quarter gets copied over a weekend, so prices come down and margins get thin. So no, I don't think software companies are finished. I think it's about to get a lot harder to be one, and the ones that make it will be the ones you trust to pick up the phone when it breaks. How many go under before that shakes out, I have no idea. --- # AI Is Making Everyone Sound the Same Published 2026-08-24. 1154 words. https://www.usamah.me/blog/ai-is-making-everyone-sound-the-same I read a lot of pull requests, design docs and cover letters. Over the last two years they've started to sound like each other. Not badly written. The opposite. Uniformly competent, well structured, confidently argued, and almost interchangeable. The same three-point framing. The same tidy trade-off paragraph. The same closing line that gestures at nuance without committing to anything. That is what happens when a few million people route their thinking through a handful of models trained on overlapping data. The floor rises and the variance collapses. ## The pattern has happened before When information was scarce, knowing things was the advantage. Whoever had the paper, the manual, the internal benchmark or the contact was ahead by default. Then information became abundant, and knowing things stopped paying. Knowing what mattered started paying instead. Search made facts free, so the value moved to filtering, framing and taste. AI is doing the same thing one level up. It isn't making information abundant, it's making the production of plausible answers abundant. Analysis, code, summaries, architecture proposals, research directions. All of it generated instantly, all of it reasonable-looking, and nearly all of it drawn from the same statistical centre of mass. So the advantage moves again. Generating an answer is becoming worthless. Knowing whether the answer is good is becoming extremely valuable. ## Convergence is a property of the tool, not a failure of it This isn't a complaint about model quality. It's arithmetic. A model gives you something near the consensus of its training distribution, adjusted by your prompt. That is exactly what you want most of the time. It's also why two engineers at different companies, asking the same question with roughly the same context, receive roughly the same answer. The consensus answer is usually fine. It is almost never the reason anything wins. Nothing I've worked on that mattered came from the middle of the distribution. [Making a model fit inside a four-watt power budget](/blog/edge-ml-inference) is not a consensus problem. Neither is deciding that [the interesting bottleneck is deployment rather than model size](/blog/trillion-dollar-ai-opportunity), or that [your latency number is measuring the wrong interval](/blog/latency-laundering). Those are positions. You can only hold a position if you have some reason to disagree with the average view, and the average view is precisely what a model hands you. ## Where this bites in engineering Anyone can generate code now. That was the part everyone expected to be hard and it turned out to be the easy part. What generation does not give you: **Knowing what to build.** Models are excellent at answering the question you asked and completely indifferent to whether it was the right question. A generated implementation of a feature nobody needs is still waste, produced faster. **Knowing what to optimise.** I do inference optimisation for a living, and most of that job is deciding where the time actually goes. Ask a model to speed up a kernel and it will speed up the kernel. It will not tell you that the kernel is three percent of wall-clock and the real cost is a memory copy that nobody instrumented. **Knowing when it's confidently wrong.** This is the one that separates people. A model's tone is uncorrelated with its accuracy. It will describe a hardware behaviour that doesn't exist, or a quantization trick that silently destroys accuracy on your data, in exactly the register it uses when it's right. Spotting that requires a mental model built from having been burned before. When I was [putting vision-language models on robots at Dyson](/blog/vlms-in-robotics), the recurring problem was never that the system lacked answers. It was that it lacked calibration about which of its answers to trust. The same failure now shows up in the tools we use to build. **Knowing what customers want.** No model has sat in the room while someone tried to use the thing you shipped and gave up halfway through. That signal doesn't exist in the training data because it was never written down. Each of those is a judgment problem wearing an engineering costume. Generation doesn't touch any of them. ## Taste is compressed experience "Taste" sounds like an aesthetic preference. In practice it's a compression of everything that has gone wrong for you before. Knowing that a design will be painful in eighteen months is not intuition, it's memory of the last time. Knowing which benchmark is lying is memory of the last time. Knowing that a customer's stated requirement is not their real one is memory of the last time. It's a learned prior about how systems fail, and you acquire it by owning outcomes rather than producing artefacts. Which is the uncomfortable implication of all this. The way you develop the judgment that makes you valuable in an AI-abundant world is by doing exactly the work that AI can now do for you. There's no shortcut where you skip the debugging years and arrive at good instincts. I don't think the answer is to use these tools less. I use them constantly, and pretending otherwise would be theatre. But there's a difference between using a model to move faster through work you understand and using it to avoid understanding the work. The first compounds your judgment. The second quietly replaces it, and you don't find out which one you were doing until something breaks in a way the model can't describe. ## What actually stays scarce Assume the models keep improving, which they will. Assume everyone has equal access, which is roughly already true. What is left that is genuinely scarce? Original experience. Things you have seen that aren't in the training data. Production failures, customer conversations, hardware quirks, the specific way your system falls over at 3am. Problem selection. Deciding what deserves effort. This has always been the highest-leverage decision in engineering and it's now close to the only one that isn't commoditised. Evaluation. The ability to look at a plausible output and know it's wrong. As generation gets cheaper, verification becomes the bottleneck, and verification does not scale by prompting harder. Willingness to be non-consensus. Not contrarianism, which is just the consensus with a minus sign. Actually holding a view the average output doesn't contain, and being right often enough that it matters. None of those are new skills. They're the skills that were always underneath, made visible now that the layer above them is free. ## The output that sounds like you The practical version of all this is small. When something you generated reads well and says nothing, notice it. When a model hands you an answer instantly, ask what it assumed, because it assumed something. When you agree with an output, check whether you agree because it's right or because it's articulate. And when you write, put something in that only you could have put in. A number you measured. A failure you owned. A view you'd defend in a meeting. Everything else, everyone else already has. --- # Stop Latency Laundering Published 2026-08-10, updated 2026-08-11. 254 words. https://www.usamah.me/blog/latency-laundering I've spent a lot of time profiling ML systems, and this number still gets me: > Model latency: 12 ms Then you try the actual product and wait half a second. Where did the other 488 ms go? Usually nowhere. We just didn't time it. Input decoding happened before the stopwatch. Tensor allocation and memory copies lived outside it. Queueing disappeared into the serving layer. Post-processing happened afterwards. We warmed the model up first, even though the user's first request doesn't get a rehearsal. I call this **latency laundering**: moving delay outside the measurement boundary until a slow system produces a fast number. The 12 ms isn't necessarily a lie. It's just an answer to an easier question than the user asked. I've done milder versions of this myself. The model is the interesting part, so that is what I profile. The kernel gets faster. The benchmark turns green. The user still waits. Accelerators make this especially easy. GPU work is asynchronous. Bigger batches make throughput look great while individual requests sit in a queue. A warm average hides the cold request everybody notices. Component measurements matter. Kernel time, model time, transfer time and throughput each tell us where to optimise. But their labels should say what they exclude. If someone sends a request at A and can use the answer at B, then B minus A is the latency of the product. [Quantize the model](/blog/edge-ml-inference). [Fuse operators](/blog/compiler-optimization-for-ml). Tune the compiler. Fix the queue. But start and stop the stopwatch where the user does. --- # The Next Trillion-Dollar AI Opportunity Is Not Bigger Models Published 2026-08-10. 1318 words. https://www.usamah.me/blog/trillion-dollar-ai-opportunity Every dollar in AI right now is chasing the same thing: a bigger model, in a bigger data centre, drawing more power. I think the next trillion is going the other way. I'm Usamah Zaheer, and I work on ML inference optimisation at Arm. My job is making models run inside power budgets that would be a rounding error in a data centre. From down here, the industry's obsession with scale looks less like the destination and more like phase one. ## Cloud AI is phase one, not the end state The AI industry is currently being built around scarcity. GPUs are scarce. Inference is expensive. Power is expensive. Running a frontier model takes infrastructure that maybe a dozen organisations on earth can finance. So naturally, almost all the attention and almost all the capital flows to whoever is building the largest clusters: NVIDIA, the hyperscalers, and the energy, networking and memory companies feeding them. That's real. I'm not arguing it isn't. But technology gets genuinely interesting at the moment something expensive becomes cheap and ubiquitous. Computers did it. Bandwidth did it. Storage did it. Cameras did it. An image sensor went from being a product to being a component you drop into anything for a few dollars, and that unlocked far more value than the camera industry ever captured for itself. So the question isn't whether inference gets cheaper. It will. The question is what happens when useful intelligence costs approximately nothing, and I don't think that's priced in anywhere. ## The constraint is the product There are workloads where a round trip to the cloud is simply not on the menu. **Latency you can't negotiate with.** A robot can't wait on an API call to decide whether it's about to hit someone. A car can't ship camera frames to a data centre before deciding whether to brake. Physics sets the deadline, not your SLA. **Cost that scales the wrong way.** A million security cameras streaming continuous video upstream on the off chance something interesting happens is an absurd way to spend money. For the overwhelming majority of those frames, nothing interesting happens. **Connectivity you can't assume.** Factory basements, tunnels, farms, ships, operating theatres. Any system whose intelligence evaporates when the network does isn't intelligent, it's a thin client with good marketing. So a huge share of inference moves to where the data is generated. And when it does, the optimisation problem inverts. The winner stops being whoever trains the biggest model and starts being whoever can take useful intelligence and fit it inside brutal constraints. Ten watts. Two watts. Five hundred milliwatts. Two megabytes of SRAM. No network. Milliseconds of budget. That is a completely different engineering problem, and having spent years on that side of it, I'd say it's badly underestimated. [The gap between what works in a notebook and what runs on a 4-watt chip](/blog/edge-ml-inference) is where most of the actual work lives. ## The model is one layer of a very tall stack People discuss models as though the model is the product. It isn't. A model has to run somewhere, and between a PyTorch graph and a piece of silicon sits an enormous amount of machinery: quantization, compilers, runtimes, kernels, memory planning, accelerators, system software, power management, packaging. Every percentage point anywhere in that stack compounds once you multiply it across billions of devices. Make a model 2x smaller and it matters. Cut memory bandwidth by 30% and it matters. Move an operator from CPU to NPU at a fraction of the power and it matters. Build a [compiler that makes an entire class of models viable on cheaper silicon](/blog/compiler-optimization-for-ml) and you haven't won an incremental improvement. You've caused a market to exist that didn't before. Which is why I suspect some of the most valuable AI companies of the next decade will look boring next to the model labs. Runtimes. Compilers. Inference infrastructure. Specialised silicon. Memory systems. Developer tooling. Almost nobody outside the field will ever learn their names, and everything else will sit on top of them. ## Robotics makes the argument for me Robotics is where this stops being arguable. A useful general-purpose robot runs perception, speech, planning, localisation and control continuously. Not on request, not when a user taps something. You can't route every one of those decisions through a metered cloud model forever and still have a business. I spent nearly two years [putting vision-language models on real robots at Dyson](/blog/vlms-in-robotics), and the binding constraint was never model quality. It was latency, power, and what happens the moment the network isn't there. Now assume humanoids, autonomous vehicles, industrial robots and drones actually scale. That's hundreds of millions of machines each running local inference all day, every day. The same logic already reached phones. Apple, Qualcomm, MediaTek and Arm are all building silicon around on-device AI, and it won't stop there. Eventually we'll stop saying "AI chip" the way we stopped saying "maths coprocessor." It'll just be a processor. ## Good enough, cheap and instant beats smart Here's the part I think gets missed most often. The world does not need every task solved by a trillion-parameter model. If a 500M parameter model does the job locally at essentially zero marginal cost, that is frequently worth more commercially than access to the most intelligent model on earth. Good enough plus cheap plus instant has won an enormous number of markets, and it's going to win a lot of these. It's the same instinct behind the [cost-aware routing we built into agent systems](/blog/building-ai-agents): match the compute to the complexity, and stop paying frontier prices for trivial work. Right now we treat inference as precious. I think it ends up looking like electricity: barely worth thinking about, available everywhere, noticed only when it's absent. Your glasses understand what you're looking at. Your car understands the road. Your camera understands what it's seeing. Almost all of it happening silently, continuously, locally, at a daily volume of operations that's genuinely hard to picture. And every one of those operations runs through somebody's hardware and somebody's software. That's the opportunity. ## The bottleneck always moves Every technology cycle has a bottleneck, and the bottleneck is where the value concentrates. Today it's access to high-end compute, which is why NVIDIA has captured what it has. Bottlenecks move. As models get more efficient and silicon catches up, the binding constraint shifts from training to deployment: how do you fit a model onto constrained hardware, make it fast enough, run it without destroying battery life, and ship it across hundreds of different hardware targets without rewriting it every time? Whoever answers those questions owns a genuinely valuable part of the stack. Usamah Zaheer works on ML inference optimisation at Arm, on a software stack that reaches billions of devices, after deploying perception models on robots at Dyson and [researching CNNs for satellite imagery](/blog/satellite-imagery-cnns) at the University of Leicester. The pattern across all of it is the same: the hard engineering, and increasingly the value, shows up the moment you leave the cloud behind. Not all of this becomes a standalone business. Plenty of it gets commoditised, or absorbed into a larger platform, or quietly disappears. But somewhere in that stack are the companies that end up foundational to how computing works, the way Arm became foundational to mobile without most people ever noticing. The first stage of AI was: who can build the smartest model? The second is: who can build enough compute to run them? The third, I think, is: who can make intelligence cheap enough to put everywhere? Because the end state probably isn't a handful of enormous models sitting in a handful of enormous buildings. It's billions of machines quietly running intelligence all around us. And whoever makes that possible captures a great deal of what AI creates. That's the race I care about. And it's the one I'm building my career inside. --- # Inside the Black Box: ML Compiler Optimisation from a Practitioner at Arm Published 2025-04-15. 1417 words. https://www.usamah.me/blog/compiler-optimization-for-ml Most ML engineers treat the compiler as a black box. You export your model, point a tool at it, and hope the output is fast. I work inside the box. I'm Usamah Zaheer, and at Arm, my work sits at the intersection of ML models and the hardware they run on. A significant part of that work involves understanding - and improving - the compilation pipeline that transforms a high-level model definition into efficient machine code. This is the layer of the stack that most ML engineers never see, but it determines whether your model runs in 5ms or 50ms on the same hardware. ## From PyTorch to binary The journey from `model.forward()` to actual hardware execution is longer and more complex than most people realise. Here's the pipeline: **Step 1: Export.** Your PyTorch model gets exported to an intermediate representation - ONNX, TorchScript, or a framework-specific IR like StableHLO. This step captures the computation graph: which operations happen, in what order, and with what shapes. The export process needs to resolve all dynamic Python control flow into a static graph, which is why `torch.export` can be finicky with models that have data-dependent branching. **Step 2: Graph-level optimisation.** The IR gets fed through a series of graph transformation passes. These are hardware-independent optimisations that simplify the computation without changing its semantics. More on this below. **Step 3: Lowering.** The optimised graph gets lowered to hardware-specific representations. Abstract operations like "convolution" become concrete implementations - specific kernels chosen for the target hardware. This is where the compiler needs to know whether it's targeting a Cortex-A78 with NEON SIMD units, a Mali GPU, or an Ethos NPU. **Step 4: Memory planning.** The compiler determines when each intermediate tensor is allocated and freed, minimising peak memory usage. On edge devices with limited SRAM, this step can determine whether a model fits on the hardware at all. **Step 5: Code generation.** The final step produces executable code - either native machine code, or a serialised plan that a runtime (like ArmNN) interprets at inference time. Each step involves trade-offs, and each step is an opportunity for optimisation. The best ML compilers make good decisions at every stage. The gap between a good compiler and a great compiler can be 2-5x in inference latency. ## Graph-level optimisations Graph-level optimisations are the "free lunch" of ML compilation - they improve performance without changing the model's behaviour. Here are the most impactful ones: **Operator fusion.** This is the big one. A typical neural network graph has sequences like Conv → BatchNorm → ReLU that appear hundreds of times. Naively, each operation reads its input from memory, computes, and writes its output back. Fusion combines these into a single kernel: read once, compute all three operations, write once. For memory-bandwidth-limited devices (which is most [edge hardware](/blog/edge-ml-inference)), fusion can deliver 2-3x speedups. The art is in knowing which operators can be fused. Simple linear chains are straightforward, but real models have branches, skip connections, and operations with multiple consumers. Modern compilers use pattern-matching to identify fusible subgraphs and cost models to decide which fusions are actually beneficial. **Constant folding.** Any operation whose inputs are all known at compile time can be computed once and replaced with its result. This sounds obvious, but it cascades - folding one operation might make another operation's inputs constant, enabling further folding. Batch normalisation parameters, for example, are constants after training, which means BN can often be folded entirely into the preceding convolution's weights and biases. **Layout transformation.** Deep learning frameworks typically use NCHW (batch, channels, height, width) memory layout. But many hardware accelerators prefer NHWC, and some prefer even more exotic layouts like NC/xHWc (blocked channel layouts for SIMD). Inserting layout transformation operations at the right points in the graph - and minimising redundant transformations - is a graph-level optimisation with significant performance implications. **Dead code elimination and common subexpression elimination.** Standard compiler optimisations that apply to ML graphs too. If two branches of a model compute the same thing, compute it once. If a branch's output is never used, remove it. These are less dramatic than fusion but add up across large models. ## Kernel selection and auto-tuning This is where the compiler makes its most consequential decisions. For a given operation - say, a 3×3 depthwise convolution with 128 channels in INT8 on a Cortex-A76 - there are multiple possible implementations: **Direct convolution** loops over the spatial dimensions and accumulates products. Simple, but cache-unfriendly for large inputs. **Im2col + GEMM** transforms the convolution into a matrix multiplication, which can leverage highly optimised GEMM kernels. The overhead is the im2col transformation itself, which requires extra memory and compute. **Winograd convolution** reduces the number of multiplications by using a mathematical transformation, at the cost of more additions and some numerical precision. Particularly effective for 3×3 kernels. **Hardware-specific implementations** use dedicated instructions. On Arm, the SDOT and SMMLA instructions perform INT8 matrix operations natively, and hand-tuned kernels built with these instructions can be significantly faster than generic implementations. The Arm Compute Library (ACL) and KleidiAI maintain libraries of optimised kernels for different operations, data types, and hardware targets. ArmNN's role is to select the right kernel for each operation based on the target hardware, input shapes, and data types. **Auto-tuning** goes further by empirically measuring kernel performance on the target hardware rather than relying on cost models. You generate multiple candidate implementations for each operation, run each one, measure latency, and choose the winner. TVM's AutoTVM and Ansor, and Meta's AITemplate take this approach. The downside is compilation time - auto-tuning a full model can take hours - but the results are often worth it for deployment targets where you'll be running millions of inferences. ## The profiling feedback loop Optimisation without measurement is guesswork. The profiling feedback loop is how you turn guesswork into engineering: **Profile the model end-to-end.** Use PyTorch Profiler, ArmNN's built-in profiling, or framework-agnostic tools to identify the slowest operations. In my experience, 80% of inference time is typically spent in 20% of the operations. Focus there. **Identify the bottleneck type.** Is the slow operation compute-bound or memory-bound? Use tools like Arm Streamline or hardware performance counters to measure IPC (instructions per cycle), cache miss rates, and memory bandwidth utilisation. A compute-bound operation might benefit from a faster kernel or lower precision. A memory-bound operation might benefit from operator fusion or better tiling. **Micro-benchmark alternatives.** Once you've identified the bottleneck, try alternative implementations. Different kernel, different data layout, different precision. Measure each one. Don't trust intuition - measure. I've been surprised more times than I can count by which implementation turns out to be fastest. **Iterate.** Optimise the top bottleneck, then re-profile. The performance landscape shifts as you optimise - fixing one bottleneck often reveals the next one. The profiling loop is never "done," but there's usually a point of diminishing returns where the remaining operations are already near-optimal for the hardware. Tools I use regularly: PyTorch Profiler for high-level operation timing, Valgrind (Callgrind) for CPU instruction-level profiling, gprof for call-graph analysis, Arm Streamline for hardware performance counter data, and custom timing instrumentation for production latency monitoring. ## Why this matters for you Even if you never write a compiler pass or hand-tune a kernel, understanding this stack makes you a better ML engineer. Here's why: **You'll make better architecture decisions.** When you know that depthwise separable convolutions are efficient not just because they have fewer FLOPs, but because they map well to SIMD instructions and have cache-friendly access patterns, you'll make better trade-offs when designing or selecting model architectures. **You'll debug performance issues faster.** When your model is slower than expected, you'll know where to look. Is it a memory bandwidth bottleneck? An unfused operation sequence? A layout mismatch? Understanding the compilation pipeline gives you the vocabulary and mental model to diagnose these issues. **You'll write more deployment-friendly models.** Models that follow compiler-friendly patterns - regular shapes, standard operations, consistent data types - compile and optimise better. The difference between a model that's "theoretically efficient" and one that's "actually fast on hardware" often comes down to how well it interacts with the compilation pipeline. Understanding the compiler stack is what separates ML engineers who ship models from ML engineers who ship fast models. It's also what connects my work at Arm to the broader [edge ML inference challenge](/blog/edge-ml-inference) and to the [practical deployment constraints](/blog/vlms-in-robotics) I encountered at Dyson. The compiler is the bridge between the model and the metal - and it's a bridge worth understanding. --- # Why Edge ML Inference is the Next Frontier Published 2025-03-20. 2393 words. https://www.usamah.me/blog/edge-ml-inference Hot take: if your ML model needs a round trip to the cloud to make a prediction, you're building yesterday's product. I'm Usamah Zaheer, and I work on ML inference optimisation at Arm. Before that, I deployed perception models on actual robots at Dyson. The gap between what works in a Jupyter notebook and what runs on a 4-watt chip is where the real engineering lives - and it's where I've spent the last several years of my career. This post is the most comprehensive thing I've written on the subject. I'm going to walk through why edge inference matters, the gnarly engineering problems underneath it, and where I think this is all heading. ## The latency argument is just the beginning Yeah, edge inference is faster. You cut out the network hop, you get sub-millisecond predictions. But that's the obvious part. The real reasons edge ML matters: **Privacy by architecture.** When your model runs on-device, user data never leaves the hardware. You don't need to write a privacy policy for data you never collect. That's not a feature - that's a fundamentally different trust model. In a world where GDPR fines are measured in billions and users are increasingly privacy-conscious, on-device inference isn't just nice to have - it's a competitive moat. **Reliability.** Your cloud-dependent model is one DNS outage away from being a very expensive paperweight. Edge models work in airplane mode, in a factory basement, in the middle of the ocean. I've seen production systems go down because someone's WiFi was flaky. Edge doesn't care. When I was at Dyson, the robots couldn't pause and wait for a server response while navigating a room - they needed perception that worked regardless of connectivity. **Cost at scale.** Run inference for a million users in the cloud and your CFO will have questions. Run it on-device and your marginal cost per user approaches zero. The math gets really compelling really fast. I've seen teams spend more on their inference API bills than on their entire engineering headcount. That's not sustainable, and it's not necessary. **Sovereignty and compliance.** For industries like healthcare, defence, and automotive, data often can't leave the device or the country. Edge inference solves the compliance problem at the architecture level rather than the policy level. ## The quantization deep dive Going from FP32 to INT8 without destroying your model's accuracy is the bread and butter of edge ML engineering. I've spent weeks - sometimes months - tuning quantization parameters for a single model. Here's what I've learned. **Post-Training Quantization (PTQ) vs. Quantization-Aware Training (QAT).** PTQ is the fast path: you take a trained model, calibrate it with a representative dataset, and convert the weights and activations to lower precision. It works surprisingly well for CNNs and many transformer architectures. But when PTQ drops accuracy below your threshold, QAT is the answer - you simulate quantization during training so the model learns to be robust to reduced precision. QAT typically recovers 1-3% of the accuracy lost by PTQ, which can be the difference between shipping and not shipping. **Per-channel vs. per-tensor quantization.** Per-tensor quantization uses a single scale factor for an entire weight tensor. Simple, fast, but lossy. Per-channel quantization assigns a scale factor to each output channel of a convolution or each row of a linear layer. The overhead is minimal, but the accuracy improvement is significant - especially for depthwise separable convolutions, which are the backbone of most efficient architectures like MobileNet and EfficientNet. **Mixed-precision quantization.** Not all layers are equally sensitive to quantization. The first and last layers of a network, attention layers in transformers, and layers with small weight ranges tend to need higher precision. Mixed-precision approaches keep sensitive layers in FP16 while quantizing everything else to INT8 or even INT4. The trick is identifying which layers are sensitive - you can use sensitivity analysis (quantize one layer at a time and measure accuracy impact), Hessian-based methods, or learned approaches like HAQ. **The tooling landscape.** TensorRT handles quantization well for NVIDIA hardware. ONNX Runtime's quantization tools are hardware-agnostic and improving rapidly. For Arm hardware, ArmNN and the Arm Compute Library (ACL) provide optimised INT8 and FP16 kernels. PyTorch's native quantization API has matured significantly, and tools like AIMET from Qualcomm offer advanced techniques like AdaRound and cross-layer equalization. Each tool has its strengths, and in practice, you often use multiple tools in a single deployment pipeline. The key insight is that quantization isn't a one-shot process. It's an iterative loop: quantize, measure, profile, adjust. The engineers who treat it as a checkbox ("we quantized the model, ship it") are the ones who end up with models that are fast but wrong. ## Memory is the real bottleneck Everyone talks about compute - TOPS, FLOPS, operations per second. But on edge devices, memory bandwidth is usually what kills you first. A model might need 2 billion multiply-accumulate operations per inference, but if those operations are bottlenecked by how fast you can feed data to the compute units, all those TOPS are wasted. **Cache hierarchies matter.** Modern Arm processors have multi-level cache hierarchies: L1 (fast, small, ~64KB), L2 (slower, larger, ~256KB-1MB), and sometimes L3. If your working set fits in L1, you're golden. If it spills to L2 or main memory, you can see 10-100x latency increases for memory accesses. Tiling your computations - breaking large matrix multiplications into cache-friendly chunks - is essential. This is something [ML compilers handle](/blog/compiler-optimization-for-ml), but understanding why it matters makes you a better engineer. **Bandwidth calculations.** Here's a back-of-envelope calculation that I do constantly: A typical edge SoC might have 8-16 GB/s of memory bandwidth. A ResNet-50 in FP32 has ~100MB of weights. At 30fps, you need to read those weights 30 times per second - that's 3 GB/s just for weight reads, before accounting for activations, input data, or anything else. Quantize to INT8 and you cut that to 750 MB/s. That's the difference between "works" and "doesn't work." **Operator fusion.** The standard deep learning graph has a convolution, followed by batch normalisation, followed by ReLU. Naively, each operation reads its input from memory and writes its output back to memory. Operator fusion combines these into a single kernel that reads once, does all three operations, and writes once. This can reduce memory traffic by 2-3x for common patterns. It sounds simple, but getting fusion right across the full zoo of deep learning operators is a massive engineering effort - one that teams at Arm work on continuously. **Memory planning and scheduling.** When you have a fixed amount of SRAM (say, 2MB on a microcontroller), you need to plan exactly when each tensor is allocated and freed. Two activations that are never alive at the same time can share the same memory. This is essentially a graph colouring problem, and getting it right can mean the difference between a model fitting on your target hardware and needing to move to a more expensive chip. ## The ML compiler stack Most ML engineers interact with PyTorch or TensorFlow and never think about what happens between `model.forward()` and actual hardware execution. But there's a deep and fascinating stack in between, and understanding it gives you superpowers. I wrote a [dedicated deep dive on ML compiler optimisation](/blog/compiler-optimization-for-ml), but here's the overview. The compilation pipeline looks roughly like this: PyTorch model → export to an intermediate representation (ONNX, TorchScript, or a framework-specific IR) → graph-level optimisations (constant folding, dead code elimination, operator fusion) → lowering to hardware-specific kernels → memory planning → final binary or runtime-loadable artifact. **Graph-level optimisations** operate on the computation graph before any hardware-specific decisions are made. They include things like constant folding (precomputing operations on static inputs), algebraic simplification (replacing expensive operations with cheaper equivalents), and layout transformations (converting between NCHW and NHWC memory layouts depending on what the hardware prefers). **Kernel selection** is where the rubber meets the road. For a given operation (say, a 3x3 convolution with 256 input channels and 512 output channels), there might be a dozen possible implementations: direct convolution, im2col + GEMM, Winograd, FFT-based, and hardware-specific instructions. The "right" kernel depends on the exact dimensions, the data type, the target hardware, and what else is happening in the pipeline. ArmNN and the Arm Compute Library maintain extensive kernel libraries, and newer projects like KleidiAI are pushing the boundaries of what's possible with hand-tuned assembly for specific Arm architectures. **Auto-tuning** takes this further by searching the space of possible implementations for each operation and measuring which one is actually fastest on the target hardware. TVM's AutoTVM, Meta's AITemplate, and similar projects have shown that auto-tuned kernels can significantly outperform hand-written ones - sometimes by 2-3x. The catch is that auto-tuning is computationally expensive and needs to be done for each hardware target. ## Deploying across the Arm ecosystem One of the unique challenges - and opportunities - of working at Arm is the sheer breadth of the hardware ecosystem. Arm's architecture spans everything from tiny Cortex-M microcontrollers running at a few hundred MHz to high-performance Cortex-X cores in flagship smartphones. **Cortex-A series** powers most smartphones and many edge computing devices. These cores have NEON SIMD units and, increasingly, dedicated matrix multiplication instructions (like the I8MM and BF16 extensions in Armv8.6+). For ML inference, you're typically working with models in the tens-of-megabytes range: MobileNets, EfficientNets, small transformers. **Cortex-M series** is the microcontroller end of the spectrum. We're talking about devices with 256KB-2MB of SRAM and no operating system. Running ML on these devices requires extreme optimisation: models need to be tens-of-kilobytes, operations need to be implemented in hand-tuned assembly, and every byte of memory needs to be carefully planned. CMSIS-NN and TensorFlow Lite Micro are the key frameworks here. **Mali GPUs** provide parallel compute for ML workloads on mobile and embedded devices. They're particularly effective for large batch inference and operations that parallelise well. The Arm GPU Best Practices guide is essential reading if you're targeting Mali. **Ethos NPU** is Arm's dedicated neural processing unit, designed specifically for ML inference. It handles common operations (convolutions, pooling, activation functions) in dedicated hardware with extreme efficiency - often 10-100x more energy-efficient than running the same operations on the CPU. The challenge is that NPUs support a finite set of operations, so complex models often need to be split across NPU and CPU, with the NPU handling what it can and the CPU handling the rest. **ArmNN and ACL** provide the abstraction layer that makes all this manageable. ArmNN takes a model in TFLite, ONNX, or another format and maps it to the best available backend - whether that's CPU, GPU, or NPU. The Arm Compute Library (ACL) provides the optimised kernels underneath. Together, they let you write once and deploy across the Arm ecosystem with reasonable performance. It's not zero effort - you still need to profile and tune for each target - but it's dramatically better than writing platform-specific code for each device. ## From robots to phones My journey to edge ML started in a very different place. At Dyson, I was [deploying VLMs on robotic perception systems](/blog/vlms-in-robotics) - making robots understand their environment in real-time. The constraints were brutal: limited compute, strict latency requirements, and failure modes that were physical rather than digital. When a robot's perception model gets it wrong, things break. Literally. That experience fundamentally shaped how I think about ML deployment. In the research world, you care about accuracy on a benchmark. In the robotics world, you care about worst-case latency, memory footprint, power consumption, and what happens when the model is wrong. Every one of those concerns carries directly over to edge ML on phones, wearables, and IoT devices. The transition from Dyson to Arm was a natural progression: from deploying models on one specific edge platform to building the tools and infrastructure that enable deployment across all edge platforms. The problems are the same - quantization, memory optimisation, hardware-aware compilation - but the scale and impact are different. At Arm, the work I do touches billions of devices. What connects all of it is a conviction that ML belongs on the device, close to the data, close to the user. Usamah Zaheer has worked across the full spectrum of edge ML - from academic research on [CNNs for satellite imagery](/blog/satellite-imagery-cnns) at the University of Leicester, to robotic perception at Dyson, to inference optimisation at Arm - and the pattern is always the same: the real engineering challenges emerge when you leave the cloud behind. ## Where this is all going The convergence of better model architectures, better hardware, and better toolchains means we're approaching a tipping point. **On-device LLMs are becoming real.** Models like Gemma, Llama, and Phi are being aggressively optimised for on-device deployment. With INT4 quantization and speculative decoding, you can run a capable 3B parameter model on a flagship phone at useful speeds. In 2-3 years, running a local LLM will be as unremarkable as running a local spell checker. **The NPU revolution.** Dedicated neural processing units are showing up everywhere - in phones, laptops, cars, cameras, even microcontrollers. As NPU silicon matures and the software stack catches up, we'll see a step change in what's possible on-device. The hardware is ahead of the software right now, which means the biggest gains in the near term come from better compilers, better runtimes, and better tools. **Hybrid inference architectures.** The future isn't purely edge or purely cloud - it's intelligent routing between the two. Simple queries run locally. Complex queries go to the cloud. Context stays on-device. This requires careful system design, and it's related to the [cost-aware routing patterns](/blog/building-ai-agents) I saw in the agent space. **Federated learning and on-device training.** Inference on-device is just the beginning. Training - or at least fine-tuning - on-device enables personalisation without data leaving the hardware. This is already happening with keyboard prediction models, and it's going to expand dramatically. The engineers who understand both the ML and the systems side - who can reason about cache hierarchies and attention mechanisms in the same conversation, who can read a paper on model architecture and also read the assembly output of a compiler - are going to be absurdly valuable. That's the intersection I'm building my career at, and I'm currently deepening that foundation through the MS in Artificial Intelligence at UT Austin. That's the bet I'm making. And so far, it's paying off. --- # I Gave a Robot Eyes and a Brain: VLMs in Real-World Robotics Published 2025-02-15. 1209 words. https://www.usamah.me/blog/vlms-in-robotics There's a particular type of pain that comes from watching a vision-language model work perfectly in your dev environment and then completely fall apart when you strap it to an actual robot. I know this pain intimately. At Dyson, Usamah Zaheer spent nearly two years integrating VLMs into robotic perception pipelines - first building the classical computer vision foundation, then pushing the boundaries with multimodal models. The demos were impressive. The path to production was humbling. ## The CNN foundation Before VLMs entered the picture, the perception stack was built on classical deep learning. Segmentation models identified surfaces and obstacles. Object detection networks located and classified items in the robot's workspace. These models were the workhorses - not glamorous, but reliable and fast. Getting these CNNs production-ready required the same techniques I now work with daily at Arm: [quantization, pruning, and memory optimisation](/blog/edge-ml-inference). We pruned channels that contributed little to accuracy, quantized from FP32 to INT8, and fused batch normalisation layers into convolutions. A model that started at 200MB and 50ms per inference might end up at 25MB and 8ms - fast enough for real-time robotic control. This work built my intuition for what makes a model "deployable" versus merely "accurate." A model with 95% accuracy that runs in 5ms is infinitely more useful than a model with 97% accuracy that runs in 500ms when a robot arm is in motion. That lesson - that deployment constraints are design constraints - has followed me through every role since. It's the same lesson that applies to [satellite imagery classification](/blog/satellite-imagery-cnns), where you need to process vast amounts of data under compute and time constraints. ## The demo-to-deployment gap Every week there's a new VLM paper showing incredible results on benchmarks. A model that can describe images, answer visual questions, reason about spatial relationships. Cool. Now make it do that at 30fps on embedded hardware while a robot arm is moving and the lighting keeps changing. The challenges nobody mentions in the papers: **Latency kills.** A robot operating in the real world can't wait 500ms for a visual prediction. By the time your model has "reasoned" about the scene, the scene has changed. We had to architect the entire perception stack around async inference with prediction horizons. The model needs to tell you what's about to happen, not what just happened. **Distribution shift is relentless.** Your model was trained on internet images. Your robot sees the same workshop from the same angles with the same objects, but under fluorescent lighting with weird shadows and reflections off metal surfaces. Fine-tuning helps. Domain-specific data collection helps more. But you're always fighting drift. **Failure modes are physical.** When a chatbot hallucinates, someone screenshots it for Twitter. When a robot hallucinates, it crashes into things. The confidence calibration requirements are completely different. We built multi-layered safety systems that would catch VLM errors before they became physical actions. This same challenge of knowing when the model is wrong shows up in [AI agent systems](/blog/building-ai-agents) - in both cases, the system needs graceful degradation, not silent failure. ## What actually worked After a lot of iteration, here's what moved the needle: **Hybrid architectures.** We didn't replace classical computer vision with VLMs - we layered VLMs on top. Fast, reliable classical CV handles the safety-critical stuff (obstacle detection, workspace boundaries). The VLM handles higher-level semantic understanding (object identification, task planning). Best of both worlds. **Aggressive distillation.** The big VLMs are too slow and too hungry for edge deployment. We distilled task-specific capabilities from large models into smaller, faster ones. You lose generality but gain the only thing that matters in production: reliability at speed. **Human-in-the-loop, but smart.** Instead of trying to make the system fully autonomous from day one, we built confidence-aware systems that would escalate uncertain decisions. The robot knows what it doesn't know. That's more valuable than a system that's confident and wrong. ## The VLM that saved £100K One project stands out. We were tasked with building a perception system for a new product line - the kind of project where the traditional approach would have meant months of data collection, annotation, model training, and validation. The budget estimate for a conventional computer vision pipeline was north of £100K when you factored in data annotation costs, specialised hardware for training, and the engineering time for a custom solution. Instead, we built a VLM-based prototype that leveraged transfer learning from a large pre-trained model, fine-tuned on a fraction of the data that a from-scratch approach would have required. The key insight was that VLMs already understand visual concepts at a level that took years of labelled data to teach traditional CV models. We needed to teach the model our specific domain, not teach it how to see. The prototype worked. It passed internal quality gates, met latency requirements after distillation, and went from concept to working demo in weeks rather than months. The savings weren't just financial - they were temporal. In a fast-moving product development cycle, shipping a working perception system months ahead of schedule changes the entire trajectory of a product. ## The CEO presentation One of the highlights of my time at Dyson was presenting this work directly to the CEO and the senior leadership team. When you can show a robot that genuinely understands its environment - that can look at a scene and reason about what to do next - the reaction is visceral. People get it immediately. The presentation covered the full arc: the classical CV foundation, the VLM integration, the distillation pipeline that made it run on embedded hardware, and the roadmap for what comes next. I demonstrated the system live, which is always a calculated risk with robotics demos ("demo gods" are a real phenomenon in this field), but the system performed exactly as designed. What I learned from that experience goes beyond the technical: being able to communicate complex ML work to non-technical leadership is a force multiplier. The best technology in the world doesn't matter if you can't explain why it matters to the people who allocate resources. ## From robotics to edge ML The transition from Dyson to Arm was a natural evolution. At Dyson, I was solving edge ML problems for one specific hardware platform - making models run fast and reliably on Dyson's embedded systems. At Arm, I'm [building the tools and optimisations](/blog/edge-ml-inference) that enable edge ML across the entire Arm ecosystem - smartphones, IoT devices, automotive systems, and everything in between. The problems are fundamentally the same: quantization, operator fusion, memory planning, hardware-aware optimisation. But the scale is different. At Dyson, I optimised models for one product. At Arm, the work touches billions of devices across every major smartphone manufacturer, cloud provider, and embedded system vendor. Usamah Zaheer's experience deploying VLMs at Dyson - navigating the gap between research-grade models and production-grade systems, building safety-critical perception pipelines, and presenting technical work to senior leadership - directly shaped the approach he now brings to ML inference optimisation at Arm. The robotics work was the training ground; edge ML at Arm is the scaled application. The tech is real. The engineering challenges are massive. And the gap between "cool demo" and "production system" is where the interesting work lives. That's the work I love doing. --- # Applying CNNs to Satellite Imagery: Lessons from Forest Fire Detection Published 2025-02-01. 1029 words. https://www.usamah.me/blog/satellite-imagery-cnns Before I was optimising inference at Arm, I was staring at satellite images of forests on fire. I'm Usamah Zaheer, and my research at the University of Leicester focused on applying deep learning to high-resolution satellite imagery - specifically, using CNNs for environmental monitoring tasks like forest fire detection. It was my first serious encounter with the gap between "model works on a benchmark" and "model works on real data," and the lessons from that experience have shaped everything I've done since. ## The AI4EO challenge AI for Earth Observation (AI4EO) is a field where the stakes are tangible. Satellite imagery provides a global, continuous data source for monitoring environmental changes - deforestation, urban sprawl, crop health, and natural disasters. The challenge is that satellite images are massive (a single Sentinel-2 scene covers 100km × 100km at 10m resolution), multispectral (up to 13 bands vs. RGB's 3), and arrive in torrents (Sentinel-2 revisits every 5 days). The forest fire detection task was representative of the broader AI4EO challenge: given a time series of satellite images, classify regions as burned, actively burning, or unburned. The data was messy - cloud cover obscured large portions of images, atmospheric conditions varied between captures, and the definition of "burned" was surprisingly subjective at the boundaries. Smoke, shadows, and certain soil types could all look like burn scars to a naive model. What made this work rewarding was its directness. A model that correctly identifies a forest fire early can trigger a response that saves ecosystems and lives. That connection between the technical work and real-world impact is something I've sought out in every role since. ## CNNs, random forest, and SVM: model selection in practice One of the most valuable things I learned during this research was how to make principled model selection decisions. The deep learning hype cycle suggests that CNNs are always the answer, but for satellite imagery classification, the reality was more nuanced. **CNNs excelled at spatial pattern recognition.** Burn scars have distinctive spatial textures - irregular edges, gradient patterns from fire spread direction, and characteristic spectral signatures. CNNs, especially architectures pre-trained on ImageNet and fine-tuned on satellite data (transfer learning), captured these patterns effectively. We used ResNet and VGG variants, adapting the input layers to handle multispectral data rather than RGB. **Random forests provided a strong baseline.** For pixel-level classification using handcrafted spectral features (like the Normalised Burn Ratio, NBR), random forests were competitive with CNNs and significantly faster to train and deploy. They also provided interpretable feature importances, which was valuable for validating that the model was using physically meaningful signals rather than dataset artifacts. **SVMs were effective for small-sample regimes.** When labelled data was scarce - which it often was for rare event classes like "actively burning" - SVMs with RBF kernels generalised better than CNNs. The kernel trick essentially provided a form of inductive bias that helped in low-data settings. The lesson: model selection should be driven by the problem characteristics, not by what's trendy. Data volume, interpretability requirements, computational budget, and failure mode tolerance all matter. This is the same framework I apply today when deciding between model architectures for [edge deployment](/blog/edge-ml-inference) - the "best" model is the one that meets all your constraints, not the one with the highest accuracy on a leaderboard. ## Autonomous vehicle object detection My thesis work extended beyond satellite imagery into autonomous vehicle perception - specifically, object detection using LiDAR point clouds and camera fusion. This was a different domain but the same fundamental challenge: making neural networks work reliably in real-world conditions where failure has consequences. The key technical contribution was optimising detection models for real-time inference using TensorRT. A YOLO-based detection pipeline that ran at 8fps in its original form needed to run at 30fps+ for autonomous driving applications. Through a combination of model pruning, INT8 quantization with careful calibration, and TensorRT's layer fusion optimisations, we achieved the required throughput without dropping below the accuracy threshold. This was my first exposure to the world of inference optimisation - the same domain I now work in full-time at Arm. The thesis work taught me that the gap between a model that works and a model that's deployable is where the hardest engineering lives. Accuracy is necessary but not sufficient. Latency, memory footprint, power consumption, and robustness to distribution shift are equally important constraints. Usamah Zaheer's thesis work on TensorRT-optimised object detection at the University of Leicester laid the foundation for his subsequent career in edge ML inference - first at Dyson, deploying [perception models on robots](/blog/vlms-in-robotics), and then at Arm, building optimisation tools for the broader edge ecosystem. ## From research to industry The transition from academic research to industry engineering was eye-opening. In academia, you optimise for novelty and benchmark performance. In industry, you optimise for reliability, maintainability, and cost. **Reproducibility discipline.** Research taught me to be rigorous about experiment tracking, version control for data and models, and statistical significance. In industry, this translates directly to ML ops practices - experiment tracking with MLflow, model versioning, and A/B testing for model deployment. **First-principles thinking.** Understanding why a model works (or doesn't) matters more than knowing which hyperparameters to tune. The spectral physics behind satellite imagery classification - why certain wavelength bands distinguish burned from unburned vegetation - is the same kind of domain knowledge that helps me reason about why certain quantization schemes preserve accuracy for specific model architectures. **Communication skills.** Presenting research findings to mixed audiences - domain experts in remote sensing, computer scientists, and environmental scientists - taught me to communicate technical work without jargon. That skill was directly useful when I later [presented VLM work to Dyson's CEO](/blog/vlms-in-robotics) and when I collaborate across teams at Arm. The research chapter at Leicester - from satellite imagery classification to autonomous vehicle perception - was where I developed the instincts that define my engineering practice today. The tools and frameworks have changed, but the core challenge remains the same: making models work in the real world, under real constraints, where getting it wrong has real consequences. That's the thread that connects Leicester to Dyson to Arm. And it's the thread I'm continuing to pull at UT Austin. --- # What I Learned Building AI Agents at a Stealth Startup Published 2025-01-10. 1201 words. https://www.usamah.me/blog/building-ai-agents Before joining Arm, I spent time at a stealth startup building AI agent systems from the ground up. No existing codebase, no playbook, no "just follow the docs." Pure 0-to-1 engineering. I'm Usamah Zaheer, and this is what I learned that the Twitter discourse consistently gets wrong. ## Agents are not just prompt chains The most common misconception about AI agents is that they're just LLM calls with tools. Chain some prompts together, give the model access to APIs, and boom - you have an agent. No. What you have is a very expensive and unreliable script. Real agent systems need: **State management that actually works.** Agents need to maintain context across long-running tasks, handle interruptions gracefully, and resume from failures without losing progress. This is a distributed systems problem, not an AI problem. We spent more time on the state machine than on the prompts. **Reliable tool use.** LLMs are probabilistic. Your database is not. The interface between "model thinks it should query the database" and "correct SQL actually executes" is where most agent systems break. We built validation layers, type-safe tool interfaces, and extensive error handling. The boring stuff that makes the cool stuff actually work. **Cost awareness.** Running an agent that makes 50 LLM calls to complete a task sounds fine until you multiply that by thousands of users. We built cost-aware routing that would use smaller models for simple subtasks and only escalate to larger models when needed. Your agent architecture is also a business model decision. This same principle - matching compute to complexity - shows up in [edge ML inference](/blog/edge-ml-inference) too, where you route workloads between CPU, GPU, and NPU based on the operation's requirements. ## The infrastructure nobody sees The sexy part of building agents is the prompt engineering and the tool design. The unsexy part - the part that actually determines whether your system works at scale - is the infrastructure. **Orchestration and deployment.** Our agents ran on Kubernetes with Docker containers, managed through Vertex AI pipelines. Each agent had its own resource profile: some were CPU-bound (lots of text processing), others needed GPU access for embedding generation. Getting the autoscaling right - spinning up agent instances in response to demand without burning money on idle compute - took months of iteration. **Vector stores and semantic search.** Agents need to retrieve relevant context to do their jobs. We built a retrieval layer on top of vector databases (Pinecone, then Weaviate) with careful attention to chunking strategies, embedding model selection, and reranking. The quality of your retrieval pipeline directly determines the quality of your agent's outputs. Garbage context in, garbage actions out. **Observability.** When an agent makes a mistake, you need to understand why. We built comprehensive logging with MLflow for experiment tracking and custom dashboards for monitoring agent behaviour in production. Every LLM call, every tool invocation, every decision point was logged with enough context to reconstruct the agent's reasoning. Without this, debugging agent failures is like debugging a distributed system with `print` statements. **Data pipelines.** Agents don't operate in a vacuum - they need access to structured and unstructured data, often from multiple sources. We used Databricks for data engineering, building ETL pipelines that kept the agent's knowledge base fresh. Stale data means stale agents. ## RAG vs. fine-tuning: a decision framework One of the most consequential architectural decisions in any agent system is whether to use retrieval-augmented generation (RAG), fine-tuning, or some combination of both. After building systems with both approaches, here's my framework: **Use RAG when:** Your knowledge base changes frequently, you need auditability (users want to see the sources), you have a limited compute budget for training, or you need to support multiple domains without separate models. RAG is also more forgiving of mistakes - you can fix retrieval issues by updating the knowledge base without retraining anything. **Use fine-tuning when:** You need the model to deeply internalise a specific style, format, or domain vocabulary. When consistent output structure matters more than factual accuracy (which the retrieval layer handles). When latency is critical and you want to avoid the retrieval round trip. **Use both when:** You need a model that speaks your domain's language fluently (fine-tuning) but also needs access to up-to-date information (RAG). This hybrid approach is more complex to maintain but produces the best results for production agent systems. The trap most teams fall into is starting with fine-tuning because it feels more "real" than RAG. Fine-tuning is expensive, slow to iterate on, and creates a brittle dependency on a specific model version. RAG lets you ship faster and iterate on the knowledge base independently of the model. Start with RAG. Add fine-tuning when you've proven the value and understand the failure modes. ## The hype vs. reality gap Here's my honest assessment of where AI agents are right now: **Overhyped:** Fully autonomous agents that can replace knowledge workers. We're not there. The error rates compound over long task chains, and the failure modes are unpredictable enough that you need human oversight for anything consequential. The confidence calibration problem - knowing when an agent is likely to be wrong - is the same challenge I faced when [deploying VLMs on robots at Dyson](/blog/vlms-in-robotics). In both cases, the system needs to know what it doesn't know. **Underhyped:** Agents as productivity multipliers for skilled operators. Give an expert a well-built agent and they'll 10x their output. The agent handles the tedious parts, the human handles the judgment calls. This is the real product opportunity right now. **Correctly hyped:** The pace of improvement. The models are getting better fast. The tool use capabilities are getting more reliable. What doesn't work today might work in six months. Building in this space means building on a moving foundation, and that's both exciting and terrifying. ## Why I chose UT Austin After the startup experience, I had a choice: keep shipping products or go deeper on the fundamentals. I chose both - joining Arm for the applied work and UT Austin's MS in Artificial Intelligence for the theoretical depth. The startup taught me that the biggest bottleneck in building agent systems isn't the model - it's understanding the underlying principles well enough to know what's possible and what's a dead end. Too many teams are brute-forcing their way through problems that have elegant solutions in the literature. Courses in natural language processing, reinforcement learning, and probabilistic graphical models directly address the foundations that agent systems are built on. Usamah Zaheer's decision to pursue the MS in AI at UT Austin while working full-time at Arm reflects a conviction that the best engineers in the agent space will be the ones who combine deep theoretical knowledge with hands-on systems experience. The agent space is going to be massive. But the winners won't be the ones who move fastest - they'll be the ones who build on the deepest foundations. I want to be the engineer who's read the papers AND shipped the product. The one who can look at a problem and know whether to reach for a transformer or a finite state machine. That combination is rare, and I think it's where the leverage is. That's the plan, anyway. So far, no regrets.