Skip to content

Quantize and Vela: putting our model on the Ethos-U55

Module 5 — Training and deploying to several targets · Slides: slides.md · Module overview · Course page

The last and hardest target of “train once, run everywhere” is the MCU with its Ethos-U55. Understand why the NPU needs an extra Vela compile, which of three .tflite files goes where, the four-function AIM_* contract that lets edge_ai call a model, the two ways to wrap a model, and why the NPU is faster and cheaper.

By the end of this lesson, you will:

  1. Explain why only the MCU needs a Vela compile, and match model_int8, model_int8_vela and model_web to the right targets.
  2. Run quantize_vela.sh in the same Docker image to get output/model_int8_vela.tflite, and explain the options –accelerator-config ethos-u55-128 and –optimise Performance.
  3. Describe the four-function AIM_* contract (init, enqueue, dequeue, finalize) and its return codes, and name feature parity, not the graph, as the main trap when wrapping a model.
  4. Explain why the NPU is faster and uses less energy per inference than the CPU, and read a comparison table looking at worst-case latency and per-class errors.

You’ve been through lessons 5.3–5.7, and have model_int8.tflite that passed eval_pc.py, and understand quantize and dequantize. Open quantize_vela.sh to read alongside the slides.

In the same Docker image, run ./quantize_vela.sh model_int8.tflite, and look at the new file in output/. This _vela.tflite file can no longer be opened by a browser or a PC. Ask yourself: what does Vela change inside the file, and why shouldn’t accuracy change after it?

Web, Cortex-A and PC can all run an ordinary TFLite graph directly, on CPU or GPU. But an NPU is not a CPU — it’s a specialized int8 matrix-multiply circuit. Vela (Arm’s ethos-u-vela) therefore takes a full-integer int8 .tflite, and replaces the subgraph the NPU can run with a single custom op named ethos-u. Whatever the NPU can’t do is left for the Cortex-M55. quantize_vela.sh wraps the command vela --accelerator-config ethos-u55-128 --optimise Performance (a U55 at 128 MAC per cycle, tuned for speed), producing output/model_int8_vela.tflite, usable only on the MCU. At this point, there is one set of weights, three packages: model_int8.tflite (PC, Cortex-A, and Vela’s source), _vela.tflite (MCU), and model_web.tflite (browser). Loading the wrong file in the wrong place fails to load immediately.

Vela doesn’t touch the model’s maths — $q = \mathrm{round}(x/s) + z$ and $x = (q - z)\cdot s$ still use the $s, z$ that PTQ embedded into every tensor, so accuracy on the NPU should match int8 on the PC (differing only at the level of kernel rounding). If the board diverges dramatically, suspect the front-end first — a Vela file still isn’t a model the board recognizes. It must be wrapped with a four-function contract: <SLOT>_init, <SLOT>_enqueue(const float *in), <SLOT>_dequeue(float *out), and <SLOT>_finalize — Imagimob’s IPWIN streaming ABI, which DEEPCRAFT Studio generates. dequeue returning −1 (NODATA) while the window isn’t full yet is normal. There are two ways to wrap it: DEEPCRAFT’s own converter, which generates C along with the front-end, or wrapping TFLite-Micro yourself and writing a front-end that matches training exactly. The number-one trap is mismatched features, not the graph. The BENTO firmware source this course uses isn’t public yet — in the public SDK, the engine ships as a prebuilt library, so a new model registers with ai_engine_register(), or replaces an existing slot, following the Filling a model slot guide.

The NPU isn’t smarter than the CPU — it does one job, multiplying int8 matrices, with 128 parallel MAC circuits per cycle, so the same work finishes in far fewer cycles. Finishing quickly means waking the circuit briefly then going back to sleep, so energy per inference is correspondingly low. When reading a comparison table, watch for three things: overall accuracy can hide a missed class (check the confusion matrix), average latency hides the worst case, and the web can differ slightly because its activations are float.

This lesson’s slides also reference files under shared/:

The same questions are in quiz.yaml for automated checking.

  1. Which file should be sent to a web page? (single choice · objective 1)

    • a) model_int8_vela.tflite
    • b) model_web.tflite, whose I/O is float and has no NPU custom op
    • c) model.keras
    • d) Any file works
    Solution

    b — the Vela file has the ethos-u op, which only runs on the NPU. model_web.tflite is built for the browser, and model_int8.tflite is used with the PC and Cortex-A.

  2. After running ./quantize_vela.sh model_int8.tflite, where is the MCU’s file? (single choice · objective 2)

    • a) ./model_int8.tflite, overwriting the original
    • b) ./output/model_int8_vela.tflite
    • c) ./vela/model.bin
    • d) Directly on the board
    Solution

    b — Vela writes its result into the output folder, appending _vela to the name. The original int8 file remains, used for the PC and Cortex-A.

  3. AIM_GESTURE_dequeue returns −1 right after start. What does that mean? (single choice · objective 3)

    • a) The model is broken and needs reflashing
    • b) NODATA: the window hasn’t finished collecting data yet — normal
    • c) The NPU is overheating
    • d) The winning class is −1
    Solution

    b — per the IPWIN ABI, −1 means no result yet. Keep enqueueing until the window is full, and dequeue then returns 0 with a score.

  4. int8 is 0.95 accurate on the PC, but the board gets nearly everything wrong. What should you suspect first? (single choice · objective 3)

    • a) Vela computed something wrong
    • b) The board’s front-end (normalize, quantize, or the feature) doesn’t match training
    • c) The NPU is too slow
    • d) It needs retraining with more epochs
    Solution

    b — Vela never touches the maths. A gap this large usually comes from data fed into the model on the board being at a different scale or a different feature than during training.

  5. Why does the Ethos-U55 use less energy per inference than the Cortex-M55 doing it itself? (single choice · objective 4)

    • a) Because it uses float32
    • b) Because it multiplies int8 matrices with 128 parallel MACs per cycle, finishing the work in fewer cycles, so the board can go back to sleep sooner
    • c) Because it skips normalize
    • d) Because it uses no power at all
    Solution

    b — speed and energy come from the same reason: a specialized circuit doing one job in parallel uses far fewer clock cycles.

  • Run ./quantize_vela.sh model_int8.tflite in the same Docker image, and note the _vela.tflite file’s size against model_int8.tflite in your learning log.
  • Read the report Vela prints, and find whether any op fell back to running on the CPU.
  • Draw a “one set of weights, three packages” map, writing which script each file comes from and which target it goes to.

In lesson 5.9, we’ll fill in s14_tflite_board.py to benchmark int8 on the PC, run Vela, and fill in a three-target comparison table, including reading real latency from the board.

Next lesson: lesson 5.9 — Hands-on: comparing three targets — MCU, web and PC

  • If your model had several ops the Ethos-U55 can’t run, what would that do to latency, and where would you fix it?
  • What kind of work would you accept an extra compile step for, in exchange for a longer battery life?

Review questions

Answer on your own first, then open the answer.

  1. Which file should go to the web page? (Objective 1)

    1. model_int8_vela.tflite
    2. model_web.tflite ที่ I/O เป็น float และไม่มี custom op ของ NPU
    3. model.keras
    4. ไฟล์ใดก็ได้
    Show answer

    Answer: B. model_web.tflite ที่ I/O เป็น float และไม่มี custom op ของ NPU

    ไฟล์ Vela มี op ethos-u ที่รันได้เฉพาะบน NPU ส่วน model_web.tflite ทำมาเพื่อเบราว์เซอร์ และ model_int8.tflite ใช้กับ PC และ Cortex-A

  2. After ./quantize_vela.sh model_int8.tflite, where is the MCU file? (Objective 2)

    1. ./model_int8.tflite ทับไฟล์เดิม
    2. ./output/model_int8_vela.tflite
    3. ./vela/model.bin
    4. ในบอร์ดโดยตรง
    Show answer

    Answer: B. ./output/model_int8_vela.tflite

    Vela เขียนผลลงโฟลเดอร์ output ต่อท้ายชื่อด้วย _vela ไฟล์ int8 เดิมยังอยู่ไว้ใช้กับ PC และ Cortex-A

  3. AIM_GESTURE_dequeue returns −1 just after start. What does it mean? (Objective 3)

    1. โมเดลพัง ต้อง flash ใหม่
    2. NODATA: หน้าต่างยังเก็บข้อมูลไม่เต็ม เป็นเรื่องปกติ
    3. NPU ร้อนเกินไป
    4. คลาสที่ชนะคือ −1
    Show answer

    Answer: B. NODATA: หน้าต่างยังเก็บข้อมูลไม่เต็ม เป็นเรื่องปกติ

    ตาม IPWIN ABI ค่า −1 คือยังไม่มีผล enqueue ต่อไปจนหน้าต่างเต็ม dequeue จึงคืน 0 พร้อมคะแนน

  4. int8 on the PC scores 0.95 but the board is wrong almost every time. What do you suspect first? (Objective 3)

    1. Vela คำนวณผิด
    2. front-end บนบอร์ด (normalize, quantize หรือ feature) ไม่ตรงกับตอนฝึก
    3. NPU ช้าเกินไป
    4. ต้องฝึกใหม่ด้วย epoch มากขึ้น
    Show answer

    Answer: B. front-end บนบอร์ด (normalize, quantize หรือ feature) ไม่ตรงกับตอนฝึก

    Vela ไม่แก้คณิต ความต่างใหญ่ขนาดนี้มักมาจากข้อมูลที่ป้อนเข้าโมเดลบนบอร์ดคนละสเกลหรือคนละ feature กับตอนฝึก

  5. Why does the Ethos-U55 use less energy per inference than the Cortex-M55 doing the work itself? (Objective 4)

    1. เพราะใช้ float32
    2. เพราะคูณเมทริกซ์ int8 ได้ขนาน 128 MAC ต่อรอบ งานเสร็จในรอบน้อยกว่า แล้วบอร์ดกลับไปหลับได้เร็ว
    3. เพราะข้ามการ normalize
    4. เพราะไม่ต้องใช้ไฟ
    Show answer

    Answer: B. เพราะคูณเมทริกซ์ int8 ได้ขนาน 128 MAC ต่อรอบ งานเสร็จในรอบน้อยกว่า แล้วบอร์ดกลับไปหลับได้เร็ว

    ความเร็วกับพลังงานมาจากเหตุผลเดียวกัน วงจรเฉพาะทางทำงานเดียวได้ขนาน ใช้รอบสัญญาณนาฬิกาน้อยกว่ามาก

Cite this lesson

If you teach from this lesson or reuse it in slides or documents, credit it with the text below. If you changed it, add (adapted) after the title.

"Quantize and Vela: putting our model on the Ethos-U55" from TESA Open Knowledge by the Thai Embedded Systems Association (TESA), https://github.com/tesaiot/tesa-qualification-program, licensed under CC BY-NC 4.0

Thai attribution: "quantize และ Vela: เอาโมเดลของเราขึ้น Ethos-U55" จาก TESA Open Knowledge โดยสมาคมสมองกลฝังตัวไทย (Thai Embedded Systems Association: TESA) https://github.com/tesaiot/tesa-qualification-program สัญญาอนุญาต CC BY-NC 4.0

Lesson link: https://tesaiot.github.io/tesa-qualification-program/en/courses/edge-ai-developer/m05-training/l08-quantize-and-vela/

Full guide: how to cite TESA

TESA Open Knowledge · © 2026 สมาคมสมองกลฝังตัวไทย (TESA) · CC BY-NC 4.0

Content is licensed CC BY-NC 4.0. Reuse it non-commercially and credit the Thai Embedded Systems Association (TESA) every time. · How to cite TESA