Hands-on: comparing three targets, MCU, web and PC
Module 5 — Training and deploying to several targets · Slides: slides.md · Module overview · Course page
Fill five points in s14_tflite_board.py to bench the int8 model on the PC as ground truth for both accuracy and latency, call Vela through quantize_vela.sh, and print an MCU, web and PC comparison table. Then cross from Python to MicroPython, read a real latency from the board to fill in the MCU row, and decide where the model should live.
Objectives
Section titled “Objectives”By the end of this lesson, you will:
- Fill the five points in practice/s14_tflite_board.py until it prints int8’s accuracy and a non-zero per-window latency on the PC, runs Vela when available, and prints the three-target table.
- Read latency_ms from the board with edge_ai, and fill the MCU row using –mcu-ms, stating clearly whether it’s your own model or the built-in Motion model used as a reference.
- Use the completed table to explain why accuracy should agree while latency differs, and choose a target for a given scenario, with reasons.
Before you start
Section titled “Before you start”You’ve been through lesson 5.8, and understand what Vela does and which file goes to which target. Copy practice/s14_tflite_board.py into shared/training, which has dataset_tools.py, model_int8.tflite, .norm.npz and quantize_vela.sh.
- Hardware: a TESAIoT Dev Kit board already flashed with BENTO’s MicroPython firmware (this lesson needs a real board) — the bench and Vela can run on the PC, but the MCU’s latency number must come from the real board, since the emulator has no NPU.
- Prior lesson: lesson 5.8 — Quantize and Vela: putting our model on the Ethos-U55
Concepts
Section titled “Concepts”The whole file reads as one sentence: bench int8 on the PC as ground truth → run Vela to get the MCU file → lay out numbers comparing three targets → point the way to reading real latency on the board. Four of the five points sit in bench_int8(): (1) normalize with mean/std (2) quantize with the input’s scale/zero (3) time only the it.invoke() span with time.perf_counter(), accumulating it in ms (4) correct += int(o.argmax() == y[i]). And one more point sits in run_vela(): (5) subprocess.run(["./quantize_vela.sh", int8_path], check=True), wrapped in try/except — a machine without vela gets a message and moves on, instead of killing the whole script.
We measure “pure inference time”, excluding normalize or quantize, so it can be compared meaningfully against the board’s numbers. On the PC, use time.perf_counter() around invoke(); in a browser, use performance.now() around model.run(); the MCU must come from the board only — select a model with edge_ai.select(n), then read edge_ai.result()["latency_ms"] or edge_ai.latency() (the latest value, in ms). Collect several readings to see both the average and the maximum, then send it back with python s14_tflite_board.py --mcu-ms <value>. The table will fill in the MCU row for you. The MCU’s accuracy cell uses pc_acc as a placeholder first, then gets confirmed with a real verdict on the board.
Getting your own model onto the NPU is a researcher’s path requiring a firmware build: wrap the Vela file with the AIM_* contract, register the model, build, then hard power-cycle before trusting the result. In the public SDK, use ai_engine_register() instead of editing ai_engine.c. If you haven’t reached that step yet, you can use the latency of the board’s built-in Motion model as a reference point, but you must clearly note it’s a different model. Success is being able to say why accuracy across three targets should agree while latency lands at completely different levels, and which kind of work belongs where.
Worked example
Section titled “Worked example”s14_tflite_board_full.py adds benching the web path (the float-I/O file built by convert_web.py), a qualitative (not actually measured) power column, a --no-vela option, and --show-mpy, which prints MicroPython code for reading latency on the board. Open it for comparison once your practice file is done.
| File | What this file teaches |
|---|---|
| examples/s14_tflite_board_full.py | A full target-comparison lab: PC vs web vs MCU, from one single model file (full version) |
This lesson’s slides also reference files under shared/:
- shared/training
- shared/training/quantize_vela.sh — Compile an int8 .tflite for the Ethos-U55 NPU on the PSoC Edge board.
Practice
Section titled “Practice”The # TODO comments are at lines 47 (normalize), 52 (quantize), 58 (timing invoke), 68 (counting correct classes), and 83 (calling quantize_vela.sh). If latency stays at 0.00 ms forever, point 58 is still empty. If accuracy looks odd, check points 47 and 52 first, since the front-end is where things break most often.
| Practice file | Topic |
|---|---|
| practice/s14_tflite_board.py | quantize -> Vela -> run on the NPU, then compare three targets (the fill-in-the-code version) |
Solution
Section titled “Solution”Open the solution after trying on your own at least once, and read how to use the solutions first.
| Solution | Pairs with |
|---|---|
| solution/s14_tflite_board.py | practice/s14_tflite_board.py |
Check your understanding
Section titled “Check your understanding”The same questions are in quiz.yaml for automated checking.
-
The table shows the PC’s latency as 0.00 ms every time. Which point is still empty? (single choice · objective 1)
- a) Point 1, normalize
- b) Point 3, timing and calling invoke()
- c) Point 4, counting correct classes
- d) Point 5, calling Vela
Solution
b — the placeholder is pass, so there’s neither an invoke call nor accumulation into total_ms. The result is 0 latency and a meaningless accuracy.
-
The script prints “vela not found on this machine.” What should you do? (single choice · objective 1)
- a) Delete point 5
- b) Run the script inside the lesson 5.3–5.5 Docker image, which has ethos-u-vela installed
- c) Retrain the model
- d) Switch to using model_web.tflite instead
Solution
b — try/except keeps the script from dying, but there will be no
_velafile until it’s run somewhere that has vela. -
Why can’t the MCU’s latency number be measured from the PC or the BENTO Emulator? (single choice · objective 2)
- a) Because the PC and the emulator have no Ethos-U55; whatever time is measured belongs to the CPU or the browser, not the NPU
- b) Because time.perf_counter() doesn’t work
- c) Because the MCU’s latency always equals the PC’s
- d) Because Vela must run on the board
Solution
a — it must be read from edge_ai on a real board only. The number on the emulator is the browser runtime’s time.
-
An animal tracking tag on a battery that must last several months, classifying motion continuously. Where should the model run? (single choice · objective 3)
- a) Cortex-A
- b) MCU + NPU, with the
_vela.tflitefile - c) A browser
- d) PC in Docker
Solution
b — this job needs the lowest possible energy per inference, and has no OS to run on. The extra cost of the Vela step at build time is worth the longer battery life every time it runs.
The MVP for lessons 5.8–5.9: a comparison table of three targets (MCU, web, PC) with real numbers, explaining why accuracy matches but latency differs, and why the MCU needs Vela.
- All five points in the practice file are filled in. Running it in the same Docker image gives int8’s accuracy and latency on the PC, and the
_vela.tflitefile. - On the board, select a model and read
edge_ai.latency()at least 20 times, noting the average and maximum, and state which model it was. - Run
python s14_tflite_board.py --mcu-ms <average>and save the table into your learning log. - Choose a target for three scenarios (a battery-powered tag, a customer demo, a gateway on the factory floor), with reasons from the table.
Going further
Section titled “Going further”In the next module (Edge AI apps), we’ll take a model and build a real app around it — one verdict, one job — starting with an app focused on a single model.
Next lesson: lesson 6.1 — Six models and the edge_ai API: an app focused on a single model
Reflect
Section titled “Reflect”- How many times faster is your PC’s latency than the NPU’s, and how fair is that number when the PC isn’t power-constrained?
- If you had to choose between 2% higher accuracy and half the latency, which would your work need?
Review questions
Answer on your own first, then open the answer.
-
The table shows the PC latency as 0.00 ms every time. Which point is still empty? (Objective 1)
- เติม 1 normalize
- เติม 3 จับเวลาและเรียก invoke()
- เติม 4 นับคลาสถูก
- เติม 5 เรียก Vela
Show answer
Answer: B. เติม 3 จับเวลาและเรียก invoke()
placeholder คือ pass จึงไม่มีทั้งการ invoke และการบวก total_ms ผลคือ latency 0 และ accuracy ที่ไม่มีความหมาย
-
The script prints that vela is not on this machine. What should you do? (Objective 1)
- ลบจุดที่ 5 ทิ้ง
- รันสคริปต์ใน Docker image ของบทเรียน 5.3–5.5 ที่ติดตั้ง ethos-u-vela ไว้
- ฝึกโมเดลใหม่
- เปลี่ยนไปใช้ไฟล์ model_web.tflite
Show answer
Answer: B. รันสคริปต์ใน Docker image ของบทเรียน 5.3–5.5 ที่ติดตั้ง ethos-u-vela ไว้
try/except ทำให้สคริปต์ไม่ตาย แต่จะยังไม่มีไฟล์ _vela จนกว่าจะรันในที่ที่มี vela
-
Why can the MCU latency not be measured on the PC or in the BENTO Emulator? (Objective 2)
- เพราะ PC กับ Emulator ไม่มี Ethos-U55 เวลาที่วัดได้จึงเป็นของ CPU หรือเบราว์เซอร์ ไม่ใช่ของ NPU
- เพราะ time.perf_counter() ใช้ไม่ได้
- เพราะ latency ของ MCU เท่ากับ PC เสมอ
- เพราะ Vela ต้องรันบนบอร์ด
Show answer
Answer: A. เพราะ PC กับ Emulator ไม่มี Ethos-U55 เวลาที่วัดได้จึงเป็นของ CPU หรือเบราว์เซอร์ ไม่ใช่ของ NPU
ต้องอ่านจาก edge_ai บนบอร์ดจริงเท่านั้น เลขบน Emulator คือเวลาของ runtime ในเบราว์เซอร์
-
A battery animal tag that must last months while classifying motion continuously. Where should the model run? (Objective 3)
- Cortex-A
- MCU + NPU ด้วยไฟล์ _vela.tflite
- เบราว์เซอร์
- PC ใน Docker
Show answer
Answer: B. MCU + NPU ด้วยไฟล์ _vela.tflite
งานนี้ต้องการพลังงานต่อการอนุมานต่ำที่สุดและไม่มี OS ให้ใช้ ขั้น Vela ที่จ่ายเพิ่มตอน build คุ้มกับแบตที่อยู่นานขึ้นทุกครั้งที่รัน
Cite this lesson
If you teach from this lesson or reuse it in slides or documents, credit it with the text below. If you changed it, add (adapted) after the title.
"Hands-on: comparing three targets, MCU, web and PC" from TESA Open Knowledge by the Thai Embedded Systems Association (TESA), https://github.com/tesaiot/tesa-qualification-program, licensed under CC BY-NC 4.0
Thai attribution: "ลงมือทำ: เทียบสามเป้าหมาย MCU, Web และ PC" จาก TESA Open Knowledge โดยสมาคมสมองกลฝังตัวไทย (Thai Embedded Systems Association: TESA) https://github.com/tesaiot/tesa-qualification-program สัญญาอนุญาต CC BY-NC 4.0
Lesson link: https://tesaiot.github.io/tesa-qualification-program/en/courses/edge-ai-developer/m05-training/l09-three-targets-lab/
TESA Open Knowledge · © 2026 สมาคมสมองกลฝังตัวไทย (TESA) · CC BY-NC 4.0
Content is licensed CC BY-NC 4.0. Reuse it non-commercially and credit the Thai Embedded Systems Association (TESA) every time. · How to cite TESA