Skip to content

What edge AI is: the five-stage data lifecycle and where a model can run

Module 1 — Getting started: run the real thing, then take it apart · Slides: slides.md · Module overview · Course page

Learn how edge AI differs from cloud AI, see the five-stage data lifecycle that structures the whole course, and where one model can run, on a board with separate control and AI cores.

By the end of this lesson, you will:

  1. Explain how cloud AI and edge AI differ by where inference happens, and give at least three of the four reasons to run on the device (latency, privacy, cost, offline).
  2. Put the five lifecycle stages DAQ → Processing → Analysis → Training → Apps in order, and name the module of this course that covers each.
  3. Match the MCU, web, Cortex-A and PC targets to what each does with the int8 .tflite file and the runtime it uses, and explain why only the MCU needs Vela.
  4. State that MicroPython code runs on the Cortex-M33 while models run on the Cortex-M55 with the Ethos-U55, and pick the model from the six-model table that matches a given sensor.

Nothing much to prepare for this first lesson — it’s pure concepts. Actually running a model comes in lessons 1.2 and 1.3. If you have time, open BENTO IDE in another tab to get a look at the tool you’ll use throughout the course.

  • Hardware: a TESAIoT Dev Kit board already flashed with BENTO’s MicroPython firmware, or the BENTO Emulator inside BENTO IDE.

This course teaches backwards: start from something that already works, then take it apart to see how it works inside (the PRIMM approach: Predict → Run → Investigate → Modify → Make). Lessons 1.1–1.3 lead you to running a six-model menu on the board, which is stage 5 of the cycle, before stepping back to build things yourself from stage 1 in later modules.

Edge AI means running a model (inference) on the device where the data is generated, without sending raw data off to a server to think for it. Both cloud AI and edge AI can use the exact same model — they differ only in where inference happens. There are four reasons it’s worth moving inference onto the device: low latency, data that never leaves the device (privacy), no per-call server cost, and it keeps working even offline. The price you pay is that the model has to be small and fast enough to fit on a tiny chip, which is exactly what this course teaches directly.

The course follows the five-stage data lifecycle engineers actually use: DAQ (collecting raw data, module 2) → Processing (maths and physics, module 3) → Analysis (DSP, FFT, features, module 4) → Training (training a model in Docker, module 5) → Apps (inference and taking action, module 6). Most edge AI courses only teach the last stage, but we’ll also understand what a model “sees,” and why it has to be squeezed so small.

A model trained once can run across the whole spectrum: on the MCU + NPU (a BENTO board, needing an extra Vela compile step so the Ethos-U55 can read it), on the web (LiteRT.js in a browser), on Cortex-A (Raspberry Pi, Jetson, a mini PC, using ai-edge-litert), and on a PC inside Docker. The int8 .tflite file is the single source of truth, because int8 is the common denominator: the MCU requires it, and everything else can accept it. You can learn from three different surfaces with the same one set of MicroPython code: the BENTO Emulator in your browser, a real board, and Python with Docker for training models.

The board has “two brains”: the Cortex-M55 (400 MHz) with the Ethos-U55 NPU is where models run, while the Cortex-M33 (200 MHz) is the control core, and where our MicroPython code runs. When you call edge_ai.result(), code on the M33 pulls the result from the M55 across an on-chip communication channel (IPC). The firmware on the TESAIoT Dev Kit ships with six ready-to-use DEEPCRAFT models (Motion, Baby Cry, Push, Cough, Alarm, Siren), each tied to one sensor — the IMU, a microphone, or radar. These models are the work of Imagimob AB, an Infineon group company, built with DEEPCRAFT Studio. The BENTO Emulator has five of them (no Push Detection), so the order doesn’t match the board — always find a model by its name.

The same questions are in quiz.yaml for automated checking.

  1. Which of these are reasons to run a model on the device instead of sending it to the cloud? (select every correct answer) (multiple choice · objective 1)

    • a) It can decide within a fraction of a second, with no round trip to a server
    • b) Raw audio and gestures never have to leave the device
    • c) A model on the device is always bigger and more accurate than one on the cloud
    • d) The device can still decide even if the network drops
    Solution

    a, b, d — the four reasons in the slides are latency, privacy, cost and offline. Size, on the other hand, is the trade-off of going to the edge: the model has to be small enough to run on a tiny chip, not bigger than a cloud one.

  2. Put the data lifecycle stages in the order this course follows (ordering · objective 2)

    • a) Training: train a model, then export it as .tflite
    • b) DAQ: collect raw data from a sensor
    • c) Apps: run inference, then take action
    • d) Analysis: DSP, FFT and features
    • e) Processing: maths and physics
    Solution

    b → e → d → a → c — DAQ → Processing → Analysis → Training → Apps, matching modules 2 through 6. This first set of lessons deliberately starts at Apps, but the real cycle starts from data.

  3. One single model_int8.tflite file runs on several targets. Which target needs an extra compile step? (single choice · objective 3)

    • a) A browser using LiteRT.js
    • b) An MCU board with an Ethos-U55, which needs Vela first
    • c) A Raspberry Pi using ai-edge-litert
    • d) A PC inside Docker
    Solution

    b — Vela converts part of the graph so the Ethos-U55 NPU can read it. Every other target can use the same int8 file directly, which is why int8 is the common denominator across every target.

  4. When MicroPython code calls edge_ai.result(), what happens? (single choice · objective 4)

    • a) The Cortex-M33 runs the model itself and returns the answer
    • b) Code on the Cortex-M33 pulls the latest inference result, computed by the Cortex-M55 and the Ethos-U55, over the on-chip communication channel
    • c) The board sends data to the cloud and waits for an answer
    • d) The Ethos-U55 runs the Python code instead of the M33
    Solution

    b — the model infers on the M55 and the NPU, while MicroPython runs on the M33. Calling result() pulls the result across cores over the on-chip communication channel.

  5. You want to detect someone reaching toward the board without using a camera or sound. Which model on the TESAIoT Dev Kit should you pick? (single choice · objective 4)

    • a) Motion Detection (IMU)
    • b) Cough Detection (MIC)
    • c) Push Detection (RADAR)
    • d) Siren Detection (MIC)
    Solution

    c — Push Detection uses 60 GHz radar, so it can detect a reaching hand without a camera. This model only exists on boards with radar — you won’t see it in the list on the BENTO Emulator.

In the next lesson, we’ll meet the edge_ai module’s four commands, which take us from the model registry all the way to an answer on screen.

Next lesson: lesson 1.2 — The edge_ai module: query the model registry, select, then read the answer

  • Which jobs at home or at work “must answer instantly” or “should never let data leave the machine,” to the point they should really be edge AI?
  • If a model can run both on a cheap chip and on a Jetson, what criteria would you use to decide where it belongs?

Review questions

Answer on your own first, then open the answer.

  1. Which are reasons to run the model on the device rather than in the cloud? (select all that apply) (Objective 1)

    1. ตัดสินใจได้ในเสี้ยววินาทีโดยไม่ต้องรอ round-trip ไปเซิร์ฟเวอร์
    2. เสียงและท่าทางดิบไม่ต้องออกจากอุปกรณ์
    3. โมเดลบนอุปกรณ์ใหญ่และแม่นกว่าโมเดลบนคลาวด์เสมอ
    4. อุปกรณ์ยังตัดสินใจได้แม้เน็ตหลุด
    Show answer

    Answer: A. ตัดสินใจได้ในเสี้ยววินาทีโดยไม่ต้องรอ round-trip ไปเซิร์ฟเวอร์ · B. เสียงและท่าทางดิบไม่ต้องออกจากอุปกรณ์ · D. อุปกรณ์ยังตัดสินใจได้แม้เน็ตหลุด

    สี่เหตุผลในสไลด์คือ latency, privacy, cost และ offline ส่วนขนาดกลับเป็นข้อแลกเปลี่ยนของ Edge คือโมเดลต้องเล็กพอจะรันบนชิปเล็ก ๆ ไม่ได้ใหญ่กว่าคลาวด์

  2. Put the stages of the data lifecycle in the order this course follows. (Objective 2)

    1. Training: ฝึกโมเดลแล้วส่งออกเป็น .tflite
    2. DAQ: เก็บข้อมูลดิบจากเซนเซอร์
    3. Apps: อนุมานแล้วสั่งการ
    4. Analysis: DSP, FFT และ feature
    5. Processing: คณิตและฟิสิกส์
    Show answer

    Correct order: B. DAQ: เก็บข้อมูลดิบจากเซนเซอร์ → E. Processing: คณิตและฟิสิกส์ → D. Analysis: DSP, FFT และ feature → A. Training: ฝึกโมเดลแล้วส่งออกเป็น .tflite → C. Apps: อนุมานแล้วสั่งการ

    DAQ → Processing → Analysis → Training → Apps ตรงกับโมดูล 2 ถึง 6 ชุดบทเรียนแรกเริ่มที่ Apps ก่อนโดยตั้งใจ แต่วงจรจริงเริ่มจากข้อมูล

  3. One model_int8.tflite goes to several targets. Which target needs an extra compile step? (Objective 3)

    1. เบราว์เซอร์ที่ใช้ LiteRT.js
    2. บอร์ด MCU ที่มี Ethos-U55 ต้องผ่าน Vela ก่อน
    3. Raspberry Pi ที่ใช้ ai-edge-litert
    4. PC ใน Docker
    Show answer

    Answer: B. บอร์ด MCU ที่มี Ethos-U55 ต้องผ่าน Vela ก่อน

    Vela แปลงส่วนของกราฟให้ NPU Ethos-U55 อ่านออก เป้าหมายอื่นใช้ไฟล์ int8 เดิมได้เลย int8 จึงเป็นตัวหารร่วมของทุกเป้าหมาย

  4. What happens when MicroPython code calls edge_ai.result()? (Objective 4)

    1. Cortex-M33 รันโมเดลเองแล้วคืนคำตอบ
    2. โค้ดบน Cortex-M33 ดึงผลอนุมานล่าสุดที่ Cortex-M55 กับ Ethos-U55 คำนวณไว้ผ่าน IPC
    3. บอร์ดส่งข้อมูลขึ้นคลาวด์แล้วรอคำตอบ
    4. Ethos-U55 รันโค้ด Python แทน M33
    Show answer

    Answer: B. โค้ดบน Cortex-M33 ดึงผลอนุมานล่าสุดที่ Cortex-M55 กับ Ethos-U55 คำนวณไว้ผ่าน IPC

    โมเดลอนุมานบน M55 และ NPU ส่วน MicroPython อยู่บน M33 การเรียก result() คือการดึง (pull) ผลข้ามคอร์ผ่านช่องสื่อสารในชิป

  5. To detect a hand pushed toward the board without a camera or sound, which model on the TESAIoT Dev Kit fits? (Objective 4)

    1. Motion Detection (IMU)
    2. Cough Detection (MIC)
    3. Push Detection (RADAR)
    4. Siren Detection (MIC)
    Show answer

    Answer: C. Push Detection (RADAR)

    Push Detection ใช้เรดาร์ 60 GHz จึงตรวจการยื่นมือได้โดยไม่ใช้กล้อง โมเดลนี้มีเฉพาะบนบอร์ดที่มีเรดาร์ บน BENTO Emulator จะไม่เห็นในรายชื่อ

Cite this lesson

If you teach from this lesson or reuse it in slides or documents, credit it with the text below. If you changed it, add (adapted) after the title.

"What edge AI is: the five-stage data lifecycle and where a model can run" from TESA Open Knowledge by the Thai Embedded Systems Association (TESA), https://github.com/tesaiot/tesa-qualification-program, licensed under CC BY-NC 4.0

Thai attribution: "Edge AI คืออะไร: วงจรชีวิตของข้อมูลห้าขั้นและเป้าหมายที่โมเดลไปรันได้" จาก TESA Open Knowledge โดยสมาคมสมองกลฝังตัวไทย (Thai Embedded Systems Association: TESA) https://github.com/tesaiot/tesa-qualification-program สัญญาอนุญาต CC BY-NC 4.0

Lesson link: https://tesaiot.github.io/tesa-qualification-program/en/courses/edge-ai-developer/m01-onboarding/l01-edge-ai-lifecycle/

Full guide: how to cite TESA

TESA Open Knowledge · © 2026 สมาคมสมองกลฝังตัวไทย (TESA) · CC BY-NC 4.0

Content is licensed CC BY-NC 4.0. Reuse it non-commercially and credit the Thai Embedded Systems Association (TESA) every time. · How to cite TESA