Skip to content

Running the model on the web: LiteRT.js, int8 I/O and parity

Module 5 — Training and deploying to several targets · Slides: slides.md · Module overview · Course page

Take the same model into the browser. Choose a runtime with reasons (LiteRT.js, ONNX Runtime Web, tfjs-tflite, WebNN), understand the int8 I/O trap that calls for a second web file, the front-end that lives outside the graph, the maths of quantize and dequantize, and parity defined by max-abs-diff against a TOL.

By the end of this lesson, you will:

  1. Compare four browser runtimes and justify a choice, including why the BENTO Emulator converts the model to ONNX and runs it with ONNX Runtime Web.
  2. Explain why the MCU’s full-integer int8 model may need a second web file, and how convert_web.py builds a weight-only int8 file with float I/O from the same Keras weights.
  3. Compute quantize q = clip(round(x/s + z), −128, 127) and dequantize x̂ = (q − z)·s from a given scale and zero-point.
  4. Judge parity by d = max|s_pc − s_web| ≤ TOL with the same winning class, and explain why d need not be zero, and why the front-end is where parity breaks most often.

You’ve been through lessons 5.3–5.5, and have model_int8.tflite and model_int8.tflite.norm.npz. If you’ll make a web file, run train.py --save-keras to get model.keras. Open the BENTO Emulator, select the Motion model, and flip the REAL switch on to watch alongside.

In the BENTO Emulator’s Edge AI panel, select the Motion model and flip the REAL switch on (the label changes from MOCK to REAL). The three-class confidence bars will move with the simulated IMU. This is a hand-gesture model trained with the same pipeline as ours, running in a browser tab with no server. Ask yourself: how does it know the model’s mean/std and scale?

A browser model needs no installation — send a link, and a verdict is right there to try before flashing anything, and data never leaves the user’s machine. The author surveyed the runtimes as follows: LiteRT.js (@litertjs/core) loads a .tflite file — the same format as the MCU — via WASM or WebGPU, with loadLiteRt(), loadAndCompile() and run(), making it the primary choice. @tensorflow/tfjs-tflite is no longer developed. ONNX Runtime Web is fully mature and handles int8 well, but needs an extra step converting TFLite to ONNX. WebNN isn’t yet production-ready. The BENTO Emulator chose ONNX Runtime Web’s path: convert model_int8.tflite with tf2onnx once (int8 in/out), then build the front-end itself in JS.

The MCU’s model is full-integer int8 (int8 in/out), as the Ethos-U55 requires. The author found that LiteRT.js accepts I/O as float32/int32, so this file might fail to load. convert_web.py fixes this by converting the same Keras model into model_web.tflite, dynamic-range style (Optimize.DEFAULT with no representative dataset) — weights are int8, but I/O is float, roughly four times smaller than pure float, and built from the exact same weights as the MCU’s. What makes a file “browser-clean” is having no NPU custom ops; a file that’s already been through Vela is therefore kept for the MCU only.

The graph only ever sees an already-prepared feature — the front-end lives outside the graph. Ours is normalizing with mean/std from .norm.npz (audio models carry a much heavier front-end: FFT, Mel, log). An int8 file needs its input quantized, $q = \mathrm{clip}(\mathrm{round}(x/s + z), -128, 127)$, and its output dequantized, $\hat{x} = (q - z)\cdot s$, always reading $s, z$ from the model itself. Parity is a measurement, not an assumption: the browser uses XNNPACK’s kernels, while the board uses CMSIS-NN, and the two aren’t bit-exact with each other. We judge it with $d = \max_k |s^{pc}_k - s^{web}_k| \le \mathrm{TOL}$ (for example, 0.02), with the winning class also required to match. If it fails, nine times out of ten the problem is in the front-end, especially normalization.

This lesson’s slides also reference files in another lesson and under shared/:

The same questions are in quiz.yaml for automated checking.

  1. What’s LiteRT.js’s main advantage over ONNX Runtime Web in this course? (single choice · objective 1)

    • a) Always faster on every machine
    • b) Loads the .tflite file, the same format as the MCU, directly, with no cross-format conversion
    • c) Only supports int8
    • d) Needs no front-end
    Solution

    b — this removes a conversion step that could introduce distortion. ONNX Runtime Web must convert TFLite to ONNX first, which is the path the BENTO Emulator chose to run an int8 in/out file.

  2. convert_web.py sets Optimize.DEFAULT without attaching a representative dataset. What kind of file does it produce? (single choice · objective 2)

    • a) Full-integer int8 (int8 in/out)
    • b) Dynamic-range: int8 weights, but float activations and I/O
    • c) Pure float16
    • d) A file that’s already been through Vela
    Solution

    b — with no samples to calibrate against, the converter can only compress the weights. The file shrinks about four times, while I/O stays float, which the browser can accept.

  3. An input has scale s = 0.05 and zero-point z = −10. What does x = 1.2 quantize to? (single choice · objective 3)

    • a) 14
    • b) 24
    • c) 34
    • d) −10
    Solution

    a — round(1.2 / 0.05 + (−10)) = round(24 − 10) = 14, and dequantizing gets back (14 + 10) × 0.05 = 1.2.

  4. s_pc = [0.10, 0.85, 0.05] and s_web = [0.07, 0.88, 0.05], with TOL = 0.02. What’s the result? (single choice · objective 4)

    • a) Passes, because the winning class matches
    • b) Fails, because d = 0.03 exceeds TOL, even though the winning class matches
    • c) Passes, because d = 0.00
    • d) Fails, because the winning class differs
    Solution

    b — both conditions must pass. d = max(0.03, 0.03, 0.00) = 0.03 > 0.02, so it fails. Check the front-end first.

  5. The browser predicts a different class from the PC on nearly every window. What should you check first? (single choice · objective 4)

    • a) The XNNPACK versus CMSIS-NN kernels
    • b) Normalization — whether both sides use the same mean/std set from .norm.npz
    • c) Internet speed
    • d) The number of epochs
    Solution

    b — different kernels cause small, consistent score differences. But if classes are wrong across the board, the front-end — especially normalize — is usually the cause.

  • Write a comparison table of the four runtimes in your learning log, with your reasoning for which one you’d choose for your own demo web page.
  • If you have Docker installed, run train.py --save-keras, then convert_web.py, and compare the sizes of model_int8.tflite and model_web.tflite.
  • Compute the quantize and dequantize of x = 0.8 with s = 0.04, z = −5, and see how far the value that comes back differs from the original.

In lesson 5.7, we’ll fill in s13_web.py to complete all four steps: build PC-side ground truth, measure parity, and hear the Cortex-A story.

Next lesson: lesson 5.7 — Hands-on: matching the web’s verdict to the PC’s, and the Cortex-A story

  • If you had to send a demo to a customer with no board, which runtime would you choose, and which files would you need to send?
  • Why isn’t “the winning class matches” alone enough evidence of parity?

Review questions

Answer on your own first, then open the answer.

  1. What is the main advantage of LiteRT.js over ONNX Runtime Web in this course? (Objective 1)

    1. เร็วกว่าเสมอทุกเครื่อง
    2. โหลด .tflite ฟอร์แมตเดียวกับ MCU ได้ตรง ๆ ไม่ต้องแปลงข้ามฟอร์แมต
    3. รองรับเฉพาะ int8
    4. ไม่ต้องทำ front-end
    Show answer

    Answer: B. โหลด .tflite ฟอร์แมตเดียวกับ MCU ได้ตรง ๆ ไม่ต้องแปลงข้ามฟอร์แมต

    ลดขั้นแปลงที่อาจทำให้เพี้ยน ส่วน ONNX Runtime Web ต้องแปลง TFLite เป็น ONNX ก่อน ซึ่งเป็นทางที่ BENTO Emulator เลือกเพื่อรันไฟล์ int8 in/out

  2. convert_web.py sets Optimize.DEFAULT without a representative dataset. What kind of file results? (Objective 2)

    1. full-integer int8 (int8 in/out)
    2. dynamic-range: น้ำหนัก int8 แต่ activation และ I/O เป็น float
    3. float16 ล้วน
    4. ไฟล์ที่ผ่าน Vela แล้ว
    Show answer

    Answer: B. dynamic-range: น้ำหนัก int8 แต่ activation และ I/O เป็น float

    ไม่มีตัวอย่างไว้ calibrate converter จึงบีบได้แค่น้ำหนัก ไฟล์เล็กลงราวสี่เท่าและ I/O ยังเป็น float ที่เบราว์เซอร์รับได้

  3. The input has scale s = 0.05 and zero-point z = −10. What does x = 1.2 quantize to? (Objective 3)

    1. 14
    2. 24
    3. 34
    4. −10
    Show answer

    Answer: A. 14

    round(1.2 / 0.05 + (−10)) = round(24 − 10) = 14 และ dequantize กลับได้ (14 + 10) × 0.05 = 1.2

  4. s_pc = [0.10, 0.85, 0.05] and s_web = [0.07, 0.88, 0.05] with TOL = 0.02. What is the result? (Objective 4)

    1. ผ่าน เพราะคลาสที่ชนะตรงกัน
    2. ไม่ผ่าน เพราะ d = 0.03 เกิน TOL แม้คลาสที่ชนะจะตรงกัน
    3. ผ่าน เพราะ d = 0.00
    4. ไม่ผ่าน เพราะคลาสที่ชนะต่างกัน
    Show answer

    Answer: B. ไม่ผ่าน เพราะ d = 0.03 เกิน TOL แม้คลาสที่ชนะจะตรงกัน

    ต้องผ่านทั้งสองเงื่อนไข d = max(0.03, 0.03, 0.00) = 0.03 > 0.02 จึงไม่ผ่าน ให้ไล่ตรวจ front-end ก่อน

  5. The browser picks a different class from the PC on almost every window. What do you check first? (Objective 4)

    1. kernel XNNPACK กับ CMSIS-NN
    2. normalization ว่าทั้งสองฝั่งใช้ mean/std ชุดเดียวกันจาก .norm.npz หรือไม่
    3. ความเร็วอินเทอร์เน็ต
    4. จำนวน epoch
    Show answer

    Answer: B. normalization ว่าทั้งสองฝั่งใช้ mean/std ชุดเดียวกันจาก .norm.npz หรือไม่

    kernel ต่างกันทำให้คะแนนต่างเล็กน้อยสม่ำเสมอ แต่ถ้าคลาสเพี้ยนทั้งกระดาน front-end โดยเฉพาะ normalize มักเป็นต้นเหตุ

Cite this lesson

If you teach from this lesson or reuse it in slides or documents, credit it with the text below. If you changed it, add (adapted) after the title.

"Running the model on the web: LiteRT.js, int8 I/O and parity" from TESA Open Knowledge by the Thai Embedded Systems Association (TESA), https://github.com/tesaiot/tesa-qualification-program, licensed under CC BY-NC 4.0

Thai attribution: "รันโมเดลบนเว็บ: LiteRT.js, int8 I/O และ parity" จาก TESA Open Knowledge โดยสมาคมสมองกลฝังตัวไทย (Thai Embedded Systems Association: TESA) https://github.com/tesaiot/tesa-qualification-program สัญญาอนุญาต CC BY-NC 4.0

Lesson link: https://tesaiot.github.io/tesa-qualification-program/en/courses/edge-ai-developer/m05-training/l06-web-runtime/

Full guide: how to cite TESA

TESA Open Knowledge · © 2026 สมาคมสมองกลฝังตัวไทย (TESA) · CC BY-NC 4.0

Content is licensed CC BY-NC 4.0. Reuse it non-commercially and credit the Thai Embedded Systems Association (TESA) every time. · How to cite TESA