Training and deploying to several targets
Training and deploying to several targets · Course page
Prepare a balanced dataset, train our own model in Docker, squeeze it into int8, then take one single file and run it on the web, Cortex-A and an MCU through Vela, while measuring parity.
Module goal
Section titled “Module goal”Build a model of our own, from data all the way to the chip, and be able to measure exactly what trade-off each target makes.
Lessons
Section titled “Lessons”| Lesson | Topic | Time (min) | Slides |
|---|---|---|---|
| 5.1 | Dataset engineering: class balance, windows and the train/val/test split | 65 | slides.md |
| 5.2 | Hands-on: recording a balanced dataset on the board, then splitting it on the PC | 75 | slides.md |
| 5.3 | Training a model in Docker: one artifact, four targets | 65 | slides.md |
| 5.4 | Inside training: Keras, Conv1D, gradient descent, int8 and the confusion matrix | 70 | slides.md |
| 5.5 | Hands-on: fill in a training script and run it in Docker | 75 | slides.md |
| 5.6 | Running a model on the web: LiteRT.js, int8 I/O and parity | 70 | slides.md |
| 5.7 | Hands-on: matching the web’s verdict to the PC’s, and the Cortex-A story | 75 | slides.md |
| 5.8 | Quantizing and Vela: getting our model onto the Ethos-U55 | 70 | slides.md |
| 5.9 | Hands-on: comparing three targets — MCU, web and PC | 75 | slides.md |
Lessons come in pairs: a concept lesson followed by a hands-on lesson with a practice file, a solution, and a lab.
Module checkpoint
Section titled “Module checkpoint”You pass this module once you can do all of the following (details are in the Lab section of each hands-on lesson):
- A clean, balanced, split dataset from the board — every set (train/val/test) has all three classes in close-to-equal proportion (lesson 5.2).
- Successfully train a Keras model in Docker, getting a report of float32 accuracy, int8 accuracy and a confusion matrix on a test set the model has never seen, along with the
model_int8.tfliteand.norm.npzfiles (lesson 5.5). - The web file’s verdict matches the PC side within tolerance (max|score_pc − score_web| ≤ TOL, and the winning class matches), and can explain why they don’t need to match bit for bit (lesson 5.7).
- A comparison table of three targets (MCU, web, PC) with real numbers, explaining why accuracy matches but latency differs, and why the MCU needs Vela (lesson 5.9).
TESA Open Knowledge · © 2026 สมาคมสมองกลฝังตัวไทย (TESA) · CC BY-NC 4.0
Content is licensed CC BY-NC 4.0. Reuse it non-commercially and credit the Thai Embedded Systems Association (TESA) every time. · How to cite TESA