Embedded AI Podcast Titelbild

Embedded AI Podcast

Embedded AI Podcast

Von: Embedded AI Podcast
Jetzt kostenlos hören, ohne Abo

Saisonale Sparangebote | 0,99 € pro Monat für die ersten 3 Monate

Danach 6,99 €/Monat. Bedingungen gelten. Monatlich kündbar.
A podcast about using AI in embedded systems -- either as part of your product, or during development.Embedded AI Podcast
  • E24 Practical Local LLMs: Hardware, Tools, and Trade-offs
    Oct 2 2026

    We explore the resurgence of self-hosted LLMs and why they're worth considering despite the convenience of frontier models. From data privacy concerns to cost management and consistent performance, there are solid reasons to run your own models. We break down the hardware requirements (spoiler: it's all about memory bandwidth), compare different runtime environments from beginner-friendly LM Studio to production-ready VLLM, and discuss practical model choices. Luca shares his experience running dual RTX 3090s in his basement, while Ryan questions whether his 2004 Athlon is still viable (it might be). We also touch on the reality that local models lag frontier models by about six months—which means they're roughly where Claude Opus 4A was in February, and that's perfectly usable for many tasks.

    Key Topics:

    • [02:30] Why self-host? Data privacy, cost control, and avoiding unpredictable changes from hosted providers
    • [08:45] Memory bandwidth is the real bottleneck—why GPUs still win and what that means for your hardware choices
    • [15:20] Hardware options: Mac Minis vs. discrete GPUs, and why second-hand RTX 3090s offer the best price-performance
    • [22:10] Available models: Quen, Gemini, DeepSeek, and Mistral—what works well locally and the six-month lag behind frontier models
    • [28:40] Runtime environments compared: llama.cpp for flexibility, LM Studio for ease of use, Olama for Docker-like experience, and VLLM for production
    • [35:15] Practical considerations: context windows, model size vs. available VRAM, and why speed matters more than you think
    • [40:30] Announcement: UddleSpec SDD framework now available, and the Agile Embedded Podcast is now Pragmatic Embedded

    Notable Quotes:

    "You need memory bandwidth. All of the memory. Every single token filters through all of those neural network layers—potentially 8 billion parameters per token." — Luca

    "The marketing behind these tools is very much geared toward 'don't worry what's happening behind the curtain.' But there are more efficient ways of solving the same problem without incurring the same cost." — Ryan

    "I really feel like the most rational choice is a second-hand 3090 still. They have 24 gigabytes of VRAM, they're nice and fast, and they have about half the bandwidth of a 5090—which is still pretty good and very much usable." — Luca

    Resources Mentioned:

    • llama.cpp - Command-line LLM runtime environment with broad hardware support and REST API
    • LM Studio - User-friendly GUI for running local LLMs with built-in model downloader and recommendations
    • Olama - Docker-like interface for managing and running local LLMs, good for experimentation
    • VLLM - Production-grade LLM runtime with excellent parallel request handling and memory efficiency
    • Hugging Face - Repository with hundreds of thousands of models, quantizations, and variants available for download
    • UddleSpec - Luca's new SDD framework addressing real-world workflow needs—now available on GitHub
    • Pragmatic Embedded Podcast - Sister podcast (formerly Agile Embedded) covering embedded systems development
    • Agile Embedded Podcast Slack - Community discussion channel where you can reach Ryan and Luca, with a dedicated sub-channel for Embedded AI topics
    Mehr anzeigen Weniger anzeigen
    55 Min.
  • E23 How LLMs Tick
    Sep 18 2026

    Ryan and Luca pull back the curtain on large language models, explaining what's really happening when you chat with an AI. They break down tokenization, context windows, and the stateless nature of LLMs—revealing why these tools aren't actually thinking or remembering, just generating the most likely next token based on massive matrices of weights. The conversation covers practical implications like why long sessions deteriorate, how caching affects costs, and why that apologetic "I won't do it again" from your LLM is meaningless. They also tackle the misleading anthropomorphization in AI marketing and share frustrations with features like Claude's opaque memory system. If you've ever wondered why your LLM seems to forget instructions or why it confidently states nonsense, this episode explains the mathematical reality behind the conversational illusion.

    Key Topics:

    • [02:30] Tokenization: How LLMs break language into mathematical units
    • [08:45] LLMs as stochastic algorithms: Predicting the next most likely token
    • [12:20] Training models: Billions of parameters and matrix multiplication
    • [15:10] Quantization: Trading precision for memory efficiency
    • [18:00] Context windows and token limits: Why size matters and costs money
    • [22:15] Stateless processing: LLMs don't remember, they reprocess everything
    • [26:40] Caching and timeouts: The hidden costs of pausing your session
    • [30:00] Hallucinations aren't bugs—they're the default mode of operation
    • [35:20] Attention mechanisms and why large context windows cause degradation
    • [40:15] Claude's memory system: Well-intentioned but problematic in practice
    • [44:30] Mixture of experts: Sparse vs. dense models and routing tokens

    Notable Quotes:

    "The LLM doesn't even know what truth is, so it can't lie to you. It's not lying to the truth either. It's bullshitting—it doesn't care which way is true or false." — Ryan

    "All LLMs know how to do is hallucinate. They just generate chains of tokens. If you're lucky, those chains have some connection to the real world and are actually helpful." — Luca

    "I wish it wasn't programmed to just lie to me. If the LLM says 'I won't do it again,' yes it will, because it has no memory of this incident. It will behave the same way tomorrow." — Luca

    Resources Mentioned:

    • Luca's Training and Consulting - Luca's website with links to AI and embedded systems training courses
    • Tulip Tree Tech Emulator - Ryan's commercial emulator for embedded systems development (Raspberry Pi, STM, ESP, Microchip)
    • Google's Attention Paper - The foundational 2013 paper introducing the attention mechanism that enabled modern LLMs
    • On Bullshit by Harry Frankfurt - Philosophical work distinguishing bullshitting from lying—relevant to understanding LLM outputs
    • Agile Embedded Podcast Slack - Community discussion channel where you can reach Ryan and Luca, with a dedicated sub-channel for Embedded AI topics
    Mehr anzeigen Weniger anzeigen
    53 Min.
  • E22 deploying models on small microcontrollers
    Sep 4 2026

    We talk with Sebastian Boblis and Marco Meder-Mendez from Bosch about the surprisingly challenging art of deploying AI models on resource-constrained embedded devices. They share insights from their work on ETAS Embedded AI Coder, a tool that generates optimized C code from neural network models for microcontrollers—sometimes with as little as 100 bytes of RAM available.

    The conversation covers practical strategies for model compression (often by factors of 100x or more), the counterintuitive benefits of float models over quantized ones on tiny devices, and why feature engineering still matters. Sebastian and Marco explain how they navigate the trade-offs between RAM, compute time, and accuracy, and why each project presents unique constraints—from confidential hardware specs to compiler quirks. They also discuss real-world applications, including Bosch's AI-powered wall scanner that uses radar and neural networks to detect cables in walls.

    Key Topics:

    • [03:30] Why Bosch started exploring embedded AI on tiny hardware in 2020
    • [06:45] Real-world application: AI-powered wall scanner using radar to detect cables
    • [09:20] How ETAS Embedded AI Coder generates C code from neural network models
    • [12:00] Deploying models with as few as 300 parameters and 100 bytes of RAM
    • [16:30] Model compression strategies: squeezing networks down by 100x or more
    • [21:15] When float models outperform quantized ones on tiny devices
    • [28:00] Trading off RAM vs. compute time through code generation techniques
    • [33:45] Benchmarking challenges with confidential hardware and compilers
    • [40:20] AutoML and architecture search for constrained embedded targets
    • [46:00] Free tool access for universities and opportunities for students

    Notable Quotes:

    "We go down to applications where we use 100 bytes of RAM and you can still do something useful with this on a Cortex-M0. It's very surprising how small you can get with neural networks." — Sebastian Boblis

    "Sometimes we have to squeeze it down not by one or two X, sometimes it's up to 100 X and more. This is really a regular task for us." — Marco Meder-Mendez

    "One thing that's super counterintuitive for many people on these smaller devices is to go from a quantized model to a float model. It can actually help you. It's exactly the opposite of what people do on these larger systems." — Sebastian Boblis

    Resources Mentioned:

    • ETAS Embedded AI Coder - Code generation tool for deploying neural networks on embedded devices; free for universities
    • ARM CMSIS-NN - Library containing functions for neural network layers on Cortex-M devices
    • MLPerf Tiny Benchmarks - Benchmarking suite for tiny ML systems that Sebastian's team participated in
    • Bosch AI-powered Wall Scanner - Radar-based power tool using neural networks to detect cables in walls
    • Agile Embedded Podcast Slack - Community discussion channel where you can reach Ryan and Luca, with a dedicated sub-channel for Embedded AI topics
    Mehr anzeigen Weniger anzeigen
    46 Min.
adbl_web_anon_alc_button_suppression_t1
Noch keine Rezensionen vorhanden