Skip to content

Training in Docker: one artifact, four targets

Module 5 — Training and deploying to several targets · Slides: slides.md · Module overview · Course page

Stop borrowing ready-made models and train our own. Start with why training runs in Docker, the “one artifact, four targets” picture of model_int8.tflite, then run the real train.py and eval_pc.py (or on Colab) before understanding them, and learn to read what every line of the log tells you.

By the end of this lesson, you will:

  1. Give at least two reasons to train inside Docker, and explain what -v “$PWD”:/work and –rm do in the docker run command.
  2. Draw the one-artifact, four-target map of model_int8.tflite (MCU, PC, web, Cortex-A), and say why only the MCU needs Vela.
  3. Run train.py and eval_pc.py (in Docker or on Colab), and record the window count of each set, the float32 and int8 accuracy, the file size, and the file that must travel alongside the model.

You’ve been through lessons 5.1–5.2, and have data/gestures.csv from the board, or can create a synthetic set with python dataset_tools.py --synthesize --out data/gestures.csv. Have Docker installed (the first run downloads several hundred MB of TensorFlow), or a Google Colab notebook ready.

In the shared/training folder, run docker build -t edgeai-train . once, then docker run --rm -v "$PWD":/work edgeai-train python train.py --data data/gestures.csv --out model_int8.tflite. Watch accuracy climb every epoch until a model_int8.tflite file, roughly 11 KB, appears in the folder on your own machine — before knowing what’s happening inside.

Training needs TensorFlow plus a dozen other libraries, and every machine’s versions mismatch just enough to produce “it runs on my machine.” Docker fixes this with a single image: the Dockerfile starts from python:3.11-slim, installs tensorflow, ai-edge-litert, numpy, scikit-learn, and installs ethos-u-vela as a separate layer. Everyone gets the same versions on macOS, Windows and Linux, without cluttering their machine — and it can be deleted. -v "$PWD":/work places the current folder inside the container, so code, data and results all live in the same place. --rm deletes the container when it’s done. If you don’t have Docker yet, use the train_edge_ai.ipynb notebook on Google Colab instead (the notebook synthesizes its own data in units of g, so it’s good for practising the pipeline, but a model meant for the board must be trained from the board’s own CSV).

The whole module revolves around one file, model_int8.tflite, following train once, run everywhere: the PC runs it with eval_pc.py via ai-edge-litert; Cortex-A (Raspberry Pi, Jetson) runs the same file with the same kind of script; the browser uses the version convert_web.py prepares; and the MCU needs an extra compile step with quantize_vela.sh (Vela), because the Ethos-U55 needs a graph converted into NPU ops. The resulting _vela.tflite file therefore only works on the MCU.

train.py’s log, with the synthetic set, tells you line by line: train/val/test windows: 101 21 21 is the window count for the three sets; every epoch has accuracy paired with val_accuracy, which should climb together; then float32 test accuracy on the set the model has never seen; the size of the file written (the course’s reference file is 11,504 bytes); and model_int8.tflite.norm.npz, which stores the mean/std. This file must always travel with the model. eval_pc.py runs the real int8 file and prints int8 test accuracy along with a confusion matrix — this is the ground truth before taking it anywhere else. The synthetic set separates classes easily, so the numbers look great; real data having some confusion is normal.

This lesson’s slides also reference files in another lesson and under shared/:

The same questions are in quiz.yaml for automated checking.

  1. In docker run –rm -v “$PWD”:/work edgeai-train python train.py, what does -v “$PWD”:/work do? (single choice · objective 1)

    • a) Deletes the container when it’s done
    • b) Places the current folder on your machine at /work inside the container, so code, data and results all share the same place
    • c) Names the image
    • d) Installs TensorFlow on your machine
    Solution

    b — the bind mount means the model_int8.tflite file written inside /work appears on your machine right away. –rm deletes the container, and -t during build names the image.

  2. Which target needs to compile model_int8.tflite an extra step before use? (single choice · objective 2)

    • a) A PC, via ai-edge-litert
    • b) Cortex-A, such as a Raspberry Pi
    • c) An MCU using the Ethos-U55, which needs Vela
    • d) No target needs anything extra
    Solution

    c — Vela converts the parts the NPU can run into Ethos-U ops. The resulting file therefore only runs on the MCU; PC and Cortex-A use the same int8 file directly.

  3. The log says float32 test accuracy 0.980, but eval_pc.py says int8 test accuracy 0.600. What should you suspect first? (single choice · objective 3)

    • a) The data is unbalanced
    • b) Something’s wrong with the int8 compression (calibration, or a mismatched normalize)
    • c) Docker is too slow
    • d) Nothing’s wrong — int8 should always be much less accurate
    Solution

    b — the same model should get int8 accuracy close to float32. If it drops sharply, check the representative dataset and the .norm.npz file used to normalize at test time.

  4. Which file must travel alongside model_int8.tflite to every target? (single choice · objective 3)

    • a) The Dockerfile
    • b) model_int8.tflite.norm.npz, which holds the training set’s mean/std
    • c) The whole gestures.csv file
    • d) train.py
    Solution

    b — whoever uses the model must normalize data with the same mean/std used during training, or the model won’t error out — it’ll just quietly predict wrong.

  • Build the image, then run train.py and eval_pc.py (or run the Colab notebook to the end), and note every number the log explains into your learning log.
  • Draw a map from model_int8.tflite to the four targets, writing the name of the script for each path.
  • Delete model_int8.tflite.norm.npz, then run eval_pc.py again. Note what happens, and why.

In lesson 5.4, we’ll open train.py and take apart its four beats — build, fit, convert, eval — along with the maths behind each one.

Next lesson: lesson 5.4 — Inside training: Keras, Conv1D, gradient descent, int8 and the confusion matrix

  • What other work have you run into “it runs on my machine” problems with? Could Docker have helped?
  • If you had to send a model to a web team and a firmware team at the same time, which files would you send to whom?

Review questions

Answer on your own first, then open the answer.

  1. In docker run --rm -v "$PWD":/work edgeai-train python train.py, what does -v "$PWD":/work do? (Objective 1)

    1. ลบ container ทิ้งเมื่อจบ
    2. เอาโฟลเดอร์ปัจจุบันบนเครื่องไปวางที่ /work ใน container โค้ด ข้อมูล และผลลัพธ์จึงอยู่ที่เดียวกัน
    3. ตั้งชื่ออิมเมจ
    4. ติดตั้ง TensorFlow ลงเครื่อง
    Show answer

    Answer: B. เอาโฟลเดอร์ปัจจุบันบนเครื่องไปวางที่ /work ใน container โค้ด ข้อมูล และผลลัพธ์จึงอยู่ที่เดียวกัน

    bind mount ทำให้ไฟล์ model_int8.tflite ที่เขียนใน /work โผล่บนเครื่องเราทันที ส่วน --rm คือลบ container และ -t ตอน build คือตั้งชื่อ

  2. Which target needs an extra compile step on model_int8.tflite before use? (Objective 2)

    1. PC ผ่าน ai-edge-litert
    2. Cortex-A เช่น Raspberry Pi
    3. MCU ที่ใช้ Ethos-U55 ต้องผ่าน Vela
    4. ไม่มีเป้าหมายใดต้องทำเพิ่ม
    Show answer

    Answer: C. MCU ที่ใช้ Ethos-U55 ต้องผ่าน Vela

    Vela แปลงส่วนที่ NPU รันได้เป็น op ของ Ethos-U ไฟล์ที่ได้จึงรันได้แค่บน MCU ส่วน PC และ Cortex-A ใช้ไฟล์ int8 เดิม

  3. The log shows float32 test accuracy 0.980 but eval_pc.py reports int8 test accuracy 0.600. What do you suspect first? (Objective 3)

    1. ข้อมูลไม่สมดุล
    2. การบีบเป็น int8 (calibration หรือ normalize ที่ไม่ตรงกัน) มีปัญหา
    3. Docker ช้าเกินไป
    4. ไม่มีอะไรผิด int8 ควรแม่นน้อยกว่ามากเสมอ
    Show answer

    Answer: B. การบีบเป็น int8 (calibration หรือ normalize ที่ไม่ตรงกัน) มีปัญหา

    โมเดลเดียวกันควรได้ int8 ใกล้ float32 ถ้าตกมากให้ตรวจ representative dataset และไฟล์ .norm.npz ที่ใช้ normalize ตอนทดสอบ

  4. Which file must travel with model_int8.tflite to every target? (Objective 3)

    1. Dockerfile
    2. model_int8.tflite.norm.npz ที่เก็บ mean/std ของชุดฝึก
    3. gestures.csv ทั้งไฟล์
    4. train.py
    Show answer

    Answer: B. model_int8.tflite.norm.npz ที่เก็บ mean/std ของชุดฝึก

    ฝั่งที่ใช้โมเดลต้อง normalize ข้อมูลด้วย mean/std ชุดเดียวกับตอนฝึก ไม่งั้นโมเดลไม่ error แต่ทายผิดเงียบ ๆ

Cite this lesson

If you teach from this lesson or reuse it in slides or documents, credit it with the text below. If you changed it, add (adapted) after the title.

"Training in Docker: one artifact, four targets" from TESA Open Knowledge by the Thai Embedded Systems Association (TESA), https://github.com/tesaiot/tesa-qualification-program, licensed under CC BY-NC 4.0

Thai attribution: "ฝึกโมเดลใน Docker: หนึ่งชิ้นงาน สี่เป้าหมาย" จาก TESA Open Knowledge โดยสมาคมสมองกลฝังตัวไทย (Thai Embedded Systems Association: TESA) https://github.com/tesaiot/tesa-qualification-program สัญญาอนุญาต CC BY-NC 4.0

Lesson link: https://tesaiot.github.io/tesa-qualification-program/en/courses/edge-ai-developer/m05-training/l03-training-pipeline/

Full guide: how to cite TESA

TESA Open Knowledge · © 2026 สมาคมสมองกลฝังตัวไทย (TESA) · CC BY-NC 4.0

Content is licensed CC BY-NC 4.0. Reuse it non-commercially and credit the Thai Embedded Systems Association (TESA) every time. · How to cite TESA