Skip to content

Diagnosing faults from evidence

By the end of this lesson, you will be able to

  1. Form at least three hypotheses for a “no result” symptom, and choose evidence that separates each one from the others.
  2. Read the SDK’s diagnostic counters and identify which stage is failing.
  3. Explain why a function should return an honest result — such as unavailable or no data — instead of pretending success.

Takes about 70 minutes (concepts 15 · practice 25 · lab 25 · check 5).

Two review questions from lesson 3.1.

  1. When CM33 is halted at a breakpoint, what keeps running, and why does halting change the system’s behaviour?
  2. In this template’s toolset, can we attach a debugger to CM55? If not, where must CM55’s evidence come from instead?

Open examples/06_pipeline_counters.c. This program simulates a three-stage pipeline like the SDK’s Edge AI one, and produces three kinds of fault that look identical from the screen. Predict before you run it: in the no verdict case, which counter’s difference will be zero, and what will the result column say?

Terminal window
gcc -std=c11 -Wall -Wextra -o pipeline examples/06_pipeline_counters.c
./pipeline

Every case shows result=OK, because the most recent result from when things were still fine is still sitting there, and the cumulative totals are all large too. Only the difference in counters over the measurement window tells you which stage the pipeline stopped at. The SDK writes about this exact issue in the 01_first_inference example: “A SNAPSHOT OUTLIVES ITS SESSION,” and in 07_engine_health: “A big number is not health; a big number that is not growing is a stall.”

1. One symptom, many causes: form hypotheses before touching the code

Section titled “1. One symptom, many causes: form hypotheses before touching the code”

“The screen shows 0% and nothing happens.” The header of 07_engine_health.c says this symptom “has four different causes and they need four different fixes.” A fix that actually works starts by writing down every hypothesis, then choosing evidence that gives a different answer for each one. Evidence every hypothesis predicts the same way separates nothing at all.

Hypothesis If true, the one-second difference will be Fix where
No data reaches the model (the sensor or CM33 isn’t sending) feeds +0 The data source side
The processing task isn’t running, or is stuck inside feeds moves, dq_calls +0 The task and model selection
Processing runs but produces no result (the data window isn’t full yet, or the NPU is stuck) dq_calls moves, dq_ok +0 Wait for it to fill, or check the NPU
The model never loaded successfully in the first place Use a different set of counters (next section) Model loading

The order you read in matters: start from the beginning of the pipeline, because if no data is going in, no processing cycles happening isn’t news. And check first that the measurement itself is valid — if the active model changed between two readings, the counters were cleared, and the difference means nothing. The SDK’s example reports “the active model changed under the measurement” instead of interpreting the numbers in that case.

A real story from the SDK showing why picking the right evidence matters: the header of bento_bgt60trxx_platform.c recounts the radar getting stuck repeatedly, while the sensor’s own stall watchdog reported zero stalls — because it “sits at the BOTTOM of the same loop that was stuck.” A measuring tool sitting underneath the point that’s stuck will never see the stall. Usable evidence has to sit outside whatever it’s measuring.

2. Reading the SDK’s counters: totals, differences, and ordering

Section titled “2. Reading the SDK’s counters: totals, differences, and ordering”

The SDK provides two sets of counters that answer different questions.

  • “Is it running right now?” ai_engine_feeds(), ai_engine_dq_calls(), ai_engine_dq_ok() in 07_engine_health are read as a difference between two readings a second apart, using a saturating delta32() (“Saturating, because a counter is cleared on a model switch”).
  • “Did it ever load successfully?” ai_engine_init_calls(), ai_engine_init_returns(), ai_engine_inits(), ai_engine_last_init_rc() in 10_model_load_diagnosis.c are read once, because loading either happened or didn’t, and a sentinel value 0x7FFFFFFF separates “init was never called” from a 0 that means “init succeeded”.

Another piece of evidence that’s always usable on the mtb-only variant is the [HB] line every ten seconds (the SDK documentation’s chapter G2). If it’s still coming, CM33 is still scheduling tasks — the problem lies in one particular task, the screen, or CM55, not a dead core. And if CM55 crashes with a serious fault, proj_cm55/main.c blinks an LED code — 1 blink for a stack overflow, 2 for a failed malloc, 3 for a HardFault — and writes a value of 0xDEAD0001 through 0xDEAD0003 to a fixed address in SRAM. Even a core with no console can leave evidence behind, if you design for it in advance.

A function that has no data and returns 0 while claiming success lets whoever’s downstream make a confident, wrong decision. The SDK’s example catalogue makes it a rule that “Every file returns an honest result code” — SDK_EX_OK, SDK_EX_UNAVAILABLE, SDK_EX_BUSY, SDK_EX_REFUSED, SDK_EX_NO_DATA, SDK_EX_STARTED — and “If the hardware is absent the example says so rather than pretending to succeed” (the catalogue’s README, section 5).

Examples of how an ambiguous result does real harm, recorded by the SDK’s own documentation:

  • Appendix X #25: MQTT’s 0x08060009 code gets overwritten once retries run out, “overwriting whatever result already held.” A TLS failure and being denied authorization end up printing the same code — the documentation notes that three separate bugs across three layers all printed this same code, and it never changed while each was fixed in turn.
  • The sensor task’s mask counter in 05_auto_push_task.c: a missing bit could mean “you disabled it” or “it isn’t working,” and “the API cannot tell you which.”
  • radar_dsp_snapshot() returns false when there is no first frame yet, which is different from target == 0, meaning no target. The SDK deliberately keeps these two apart.

The other side of honesty is that a reassuring-looking log is not evidence. Chapter A1 warns that on mtb-only, a successful boot prints almost nothing, and some lines exist in the source but are never actually printed — “If a document tells you to wait for one of them, the document is stale.” Good evidence is what the system actually emits, not what’s written somewhere in the code.

examples/06_pipeline_counters.c runs in three parts.

  • Part 1: pipeline_tick() simulates one beat of a pipeline; each kind of fault stops it at a different stage.
  • Part 2: get_confidence() returns RESULT_NO_DATA if there has never been a result — but once there has been one, it returns the latest result, which might be stale.
  • Part 3: run_case() reads the counters twice and prints both the cumulative value and the difference.

Try changing things and predicting the result before you run it.

  1. Shrink the initial normally-running period from 50 rounds to 0. What does the result column of the failing case change to, and why is this more honest?
  2. Add the time of the most recent successful result into get_confidence(), and have it return RESULT_NO_DATA if that result is older than some limit. Decide for yourself what “too old” means.
  3. Write a hypothesis table like the one in concept 1 for the symptom “the screen never updates the humidity value,” with at least three hypotheses and the evidence that separates each one.

Open practice/06_diagnose.c. There are 5 gaps to fill in, following the same logic as the SDK’s 07_engine_health and 10_model_load_diagnosis.

  1. A saturating delta32().
  2. diagnose() checks first whether the measurement itself is valid (the model didn’t change).
  3. diagnose() decides based on the differences, checking the earlier stage before the later one.
  4. load_diagnosis() covers five causes, in order.
  5. read_latest() returns an honest result, and does not touch the caller’s value when there’s no data.
Terminal window
gcc -std=c11 -Wall -Wextra -o diagnose practice/06_diagnose.c && ./diagnose

Notice that some tests pass even before you fill anything in — for example delta32(10u, 15u) == 0u, and the STAGE_HEALTHY case, because the starting code returns 0, and 0 already matches HEALTHY. A test that passes against code that does nothing at all proves nothing. This is the heart of lesson 6.1.

Try it yourself for at least 15 minutes first, then open solution/06_diagnose.c. The comments in the solution say which part of the system each result points to, because a good diagnosis doesn’t end at naming a cause — it ends at “go look here next.” Notice LOAD_NO_RC, the case where the counters’ own bookkeeping doesn’t add up: the solution reports it as its own state, rather than letting it fall through to LOAD_OK.

Answer the 5 questions in quiz.yaml (shown at the bottom of this page on the website), covering all three objectives. Getting 4 or more right counts as finishing the lesson.

Task: diagnose the Edge AI pipeline on the board from its counters, then audit the SDK’s own diagnostic example with evidence.

  1. Build the template with make build -j ENABLE_PAGE_EXAMPLES=1, flash it, unplug and replug the cable, and open the serial console.
  2. Run cm55/edge_ai/07_engine_health from SDK Examples with no model started yet. Record the message you get (it should say no model is running, and return NO_DATA).
  3. Start a model from the firmware’s Edge AI page, then run 07_engine_health again. Record every difference value and the VERDICT line.
  4. Predict first, then run cm55/edge_ai/10_model_load_diagnosis. Record what you see on screen from this example. Then open its file and compare it against three things: the rule in section 5 of the catalogue’s README (on the CM55 side, “Never printf … use sdk_example_logf()”), Appendix X #1 (printf on CM55 is a no-op once libbento_edge_ai.a is linked), and the line declaring this function in sdk_examples_table.c line 41 versus the definition in the example file — do the return type and parameters match? Write down what you find, with the lines as evidence, and say which conclusions you saw on the board with your own eyes and which ones you inferred from reading the code.
  5. On the serial console, check that [HB] keeps arriving every ten seconds throughout the experiment. If it stops for a while, record the time and what you were doing at that moment.

Evidence to keep in your portfolio: screenshots of 07’s output both times, your table of differences and diagnosis, a short report for question 4 that separates “actually observed” from “inferred,” and the [HB] log.

  • Read Appendix X items #25 and #27 in the SDK documentation. Both are examples of symptoms that point to the wrong place — write a hypothesis table for each.
  • Challenge: design a “black box” struct for your own project. See a design example in diag_blackbox.h (at this commit, in the mtb-only template, there’s only the header — no .c file uses it yet, which is itself another example of code existing not being evidence that it works).

Next lesson, moving into module 4: lesson 4.1, GPIO and interrupts

  • The last time you fixed a bug that came back — did you fix it based on the one hypothesis you happened to think of, or did you separate hypotheses with evidence first?
  • Which function in your own code returns 0 or true when there’s no data, and how could a caller be misled by that?

Review questions

Answer on your own first, then open the answer.

  1. Symptom: 'the Edge AI screen shows 0% and never changes'. Which are hypotheses that different evidence can separate? (choose all that apply) (Objective 1)

    1. ไม่มีข้อมูลเซนเซอร์ไปถึงโมเดล (feeds ไม่ขยับ)
    2. task ประมวลผลไม่ทำงานหรือค้าง (feeds ขยับ แต่ dq_calls ไม่ขยับ)
    3. ประมวลผลได้แต่ไม่มีผล (dq_calls ขยับ แต่ dq_ok ไม่ขยับ)
    4. เฟิร์มแวร์มีบั๊ก (ไม่ระบุว่าที่ไหน)
    Show answer

    Answer: A. ไม่มีข้อมูลเซนเซอร์ไปถึงโมเดล (feeds ไม่ขยับ) · B. task ประมวลผลไม่ทำงานหรือค้าง (feeds ขยับ แต่ dq_calls ไม่ขยับ) · C. ประมวลผลได้แต่ไม่มีผล (dq_calls ขยับ แต่ dq_ok ไม่ขยับ)

    สามข้อแรกทำนายส่วนต่างของตัวนับต่างกัน จึงแยกกันได้ด้วยการวัดครั้งเดียว ส่วน 'เฟิร์มแวร์มีบั๊ก' เป็นจริงกับทุกกรณีและไม่บอกว่าควรไปดูที่ไหน จึงไม่ใช่สมมติฐานที่ทดสอบได้

  2. The radar stall watchdog sits at the bottom of the radar task loop, and the task hangs mid-loop. What does the watchdog report, and what is the lesson? (Objective 1)

    1. รายงานว่าค้าง เพราะมันถูกออกแบบมาเพื่อจับการค้าง
    2. รายงาน 0 ครั้ง เพราะมันอยู่ใต้จุดที่ค้างและไม่เคยได้ทำงาน หลักฐานต้องมาจากนอกสิ่งที่มันวัด
    3. รีเซ็ตบอร์ดเอง
    4. รายงานค่าสุ่ม
    Show answer

    Answer: B. รายงาน 0 ครั้ง เพราะมันอยู่ใต้จุดที่ค้างและไม่เคยได้ทำงาน หลักฐานต้องมาจากนอกสิ่งที่มันวัด

    หัวไฟล์ bento_bgt60trxx_platform.c ของ SDK บันทึกเหตุการณ์นี้ไว้จริง: sensor's own stall watchdog reporting 0 attempts because it sits at the BOTTOM of the same loop that was stuck

  3. Over one second 07_engine_health reads feeds +50, dq_calls +50, dq_ok +0. Where is the fault? (Objective 2)

    1. แหล่งข้อมูล: เซนเซอร์ไม่ส่ง
    2. task ประมวลผลไม่ทำงาน
    3. ประมวลผลครบรอบแต่ไม่มีผล: หน้าต่างข้อมูลยังไม่เต็ม หรือ NPU ค้าง
    4. ปกติดี
    Show answer

    Answer: C. ประมวลผลครบรอบแต่ไม่มีผล: หน้าต่างข้อมูลยังไม่เต็ม หรือ NPU ค้าง

    ข้อมูลมาถึงและรอบประมวลผลจบ แต่ไม่มีผลออกมา ตัวอย่างของ SDK เรียกรูปแบบนี้ว่า the signature of a stalled NPU ถ้ามันไม่หายไปหลังหน้าต่างข้อมูลเต็ม

  4. Order the checks of a 10_model_load_diagnosis-style load diagnosis (first to rule out comes first). (Objective 2)

    1. last_rc != 0 → โมเดลปฏิเสธ
    2. init_calls == 0 → ไม่เคยถูกเรียก
    3. inits == 0 → โหลดไม่จบ
    4. init_returns < init_calls → ค้างอยู่ใน init
    Show answer

    Correct order: B. init_calls == 0 → ไม่เคยถูกเรียก → D. init_returns < init_calls → ค้างอยู่ใน init → A. last_rc != 0 → โมเดลปฏิเสธ → C. inits == 0 → โหลดไม่จบ

    แต่ละข้อมีความหมายเมื่อข้อก่อนหน้าถูกตัดทิ้งแล้วเท่านั้น: ถ้าไม่เคยถูกเรียก การดูรหัสผลลัพธ์ไม่มีความหมาย ถ้ายังค้างอยู่ รหัสล่าสุดก็ยังไม่ได้เขียน

  5. A humidity read-from-cache function has never had a successful sensor read. What should it do? (Objective 3)

    1. คืน 0 พร้อมรหัสสำเร็จ หน้าจอจะได้ไม่ว่าง
    2. คืนค่าที่อ่านได้ครั้งล่าสุดของเซนเซอร์ตัวอื่นแทน
    3. คืนรหัสแบบ NO_DATA (หรือ UNAVAILABLE ถ้าไม่มีเซนเซอร์) และไม่แตะค่าที่ผู้เรียกถืออยู่
    4. หยุดโปรแกรมด้วย assert
    Show answer

    Answer: C. คืนรหัสแบบ NO_DATA (หรือ UNAVAILABLE ถ้าไม่มีเซนเซอร์) และไม่แตะค่าที่ผู้เรียกถืออยู่

    ผลที่บอกความจริงให้ผู้เรียกตัดสินใจเองได้ เช่นแสดง '--' แทน 0% ที่ดูน่าเชื่อ แคตตาล็อกของ SDK ตั้งกติกาว่า if the hardware is absent the example says so rather than pretending to succeed

Cite this lesson

If you teach from this lesson or reuse it in slides or documents, credit it with the text below. If you changed it, add (adapted) after the title.

"Diagnosing faults from evidence" from TESA Open Knowledge by the Thai Embedded Systems Association (TESA), https://github.com/tesaiot/tesa-qualification-program, licensed under CC BY-NC 4.0

Thai attribution: "วินิจฉัยความผิดพลาดจากหลักฐาน" จาก TESA Open Knowledge โดยสมาคมสมองกลฝังตัวไทย (Thai Embedded Systems Association: TESA) https://github.com/tesaiot/tesa-qualification-program สัญญาอนุญาต CC BY-NC 4.0

Lesson link: https://tesaiot.github.io/tesa-qualification-program/en/courses/embedded-c-foundations/m03-debugging/l02-diagnosing-faults/

This lesson adapts the source below; keep its credit too.
https://github.com/tesaiot/tesaiot-pse84-devkit-sdk/tree/ef72c1b658178eee8c38b1e47d28b006f80a59b5 · SDK examples and docs are linked at this commit, not copied into this course. Lessons quote short excerpts (at most 25 lines) with a link to the file at this commit and the credit (Apache-2.0, tesaiot-pse84-devkit-sdk).

Full guide: how to cite TESA

TESA Open Knowledge · © 2026 สมาคมสมองกลฝังตัวไทย (TESA) · CC BY-NC 4.0

Content is licensed CC BY-NC 4.0. Reuse it non-commercially and credit the Thai Embedded Systems Association (TESA) every time. · How to cite TESA