Sampling to match the model: rate, Nyquist, windows and the CSV schema
Module 2 — Collecting sensor data (DAQ) · Slides: slides.md · Module overview · Course page
Start the cycle upstream: understand what DAQ is, why the sampling rate must be steady and match what the model was trained on (50 Hz), use the Nyquist rule and the window-length formula T = N/fs, and design the CSV schema written to the board’s flash.
Objectives
Section titled “Objectives”By the end of this lesson, you will:
- Explain why DAQ is stage 1 of the data cycle, and name at least three things that spoil a dataset (garbage in, garbage out).
- Use the Nyquist rule fs ≥ 2·fmax to judge whether a given sampling rate is enough for a signal, and compute the length of one burst with T = N/fs.
- Explain why the logger samples at 50 Hz to match the Motion model on the board, and what happens to the waveform if you record at another rate.
- Write the schema label,ax,ay,az,gx,gy,gz and a one-sample CSV line correctly, and choose the “a” or “w” file mode for the situation.
Before you start
Section titled “Before you start”You’ve been through module 1, and know the four-beat skeleton and sensors.bmi270.motion(). Keep BENTO IDE’s REPL open, and try calling sensors.bmi270.motion() while still, compared with while shaking.
- Hardware: a TESAIoT Dev Kit board already flashed with BENTO’s MicroPython firmware, or the BENTO Emulator inside BENTO IDE
- Prior lesson: lesson 1.7 — Hands-on: from verdict to action on the board
Concepts
Section titled “Concepts”DAQ (Data Acquisition) means collecting raw sensor data into a set you can store, search, and use later — the starting raw material for every model. The Motion model we played with in module 1 came into being because someone shook a sensor and collected thousands of samples first. A dataset’s quality depends on correct labels, a steady sampling rate, balanced classes, and variety in real-world conditions. Garbage in, garbage out.
The sampling rate is how many times per second the sensor is read. 50 Hz means reading every 20 ms, so the file sets RATE_MS = 20 — not a random number, but exactly the rate the Motion model on the board was trained on. Recording at a different rate stretches or shrinks the waveform, so it no longer looks like what the model has seen. The Nyquist rule says you must sample at least twice the highest frequency you want to capture, $f_s \ge 2 f_{max}$ — otherwise a fast signal disguises itself as a fake slow wave (aliasing). One burst’s length is $T = N / f_s$ — for example, BURST = 200 at 50 Hz gives 4 seconds, long enough to complete a full motion, since a motion is a pattern over time — the model has to see many samples in a row to tell circle apart from shaking.
The CSV schema is the first line, label,ax,ay,az,gx,gy,gz — a contract between whoever collects the data and whoever trains on it. Every following line is one snapshot of the IMU’s six axes with the motion’s name attached. The file is written to the board’s flash (LittleFS) with open()/write(), the same way as Python on a computer. Mode "a" appends; "w" overwrites. Only use "w" when you’re certain the file doesn’t exist yet. Before writing code, it’s worth querying the sensor in the REPL first, to see what “still” and “shaking” values actually look like, so you can tell whether what lands in the file “makes sense.”
Check your understanding
Section titled “Check your understanding”The same questions are in quiz.yaml for automated checking.
-
Which of these spoil a dataset even though the program never throws an error? (select every correct answer) (multiple choice · objective 1)
- a) Pressing the shaking button while the board sits still
- b) Collecting 1000 idle samples but only 50 shaking samples
- c) A sampling rate that drifts, sometimes fast, sometimes slow
- d) Collecting many people doing many motions, for variety
Solution
a, b, c — a wrong label, unbalanced classes, and an unsteady rate all quietly teach the model the wrong thing. Variety in real-world conditions, on the other hand, makes a model more robust.
-
A hand shake tops out around 10 Hz. What’s the minimum sampling rate the Nyquist rule requires? (single choice · objective 2)
- a) 5 Hz
- b) 10 Hz
- c) 20 Hz
- d) 100 Hz
Solution
c — fs ≥ 2·fmax = 2 × 10 = 20 Hz. Below this, aliasing occurs. Recording at 50 Hz leaves comfortable headroom.
-
Collecting N = 150 samples at fs = 50 Hz, how long is one burst? (single choice · objective 2)
- a) 0.33 seconds
- b) 3 seconds
- c) 7.5 seconds
- d) 200 seconds
Solution
b — T = N / fs = 150 / 50 = 3 seconds.
-
If you collect a dataset at 25 Hz, then train or compare against a model trained at 50 Hz, what’s the problem? (single choice · objective 3)
- a) No problem, since each sample’s value is still the same
- b) The waveform over time is spaced out and stretched differently from what the model has seen, so it may predict wrong
- c) The file becomes twice as large
- d) The board can no longer write the file
Solution
b — the sampling rate used while collecting data must match the rate used in actual deployment. This is a core rule of edge AI: no matter how good a dataset looks, it’s useless if the rate doesn’t match.
-
A logger is used to record three bursts, but the file ends up with only the last burst. What’s the likely cause? (single choice · objective 4)
- a) The file is opened inside the recording loop with mode “w”, which overwrites it every time, instead of using “a”
- b) Forgot to write \n at the end of each line
- c) Called motion() instead of acceleration()
- d) Flash is full
Solution
a — mode “a” appends without erasing old content. Use “w” only when creating a brand new file with its header row.
- In the REPL, call
sensors.bmi270.motion()while still and while shaking, and note the values in your learning log. - Compute T = N/fs for BURST = 200 at 50 Hz, and at 25 Hz.
- Write the schema and one example line, by hand, from a sample you actually read.
Going further
Section titled “Going further”In lesson 2.2, we’ll write a real logger following the four beats — schema → sample → record → rate — and collect our first dataset file.
Next lesson: lesson 2.2 — Hands-on: a DAQ logger that saves a dataset to CSV
Reflect
Section titled “Reflect”- If your sensor captures a signal that oscillates as fast as 12 Hz, what sampling rate would you set, and why not set it higher than necessary?
- What effect does one wrongly labelled burst have on a model, among thousands of lines in a dataset?
Review questions
Answer on your own first, then open the answer.
-
Which of these spoil a dataset even though the program raises no error? (select all that apply) (Objective 1)
- กดปุ่ม shaking แต่วางบอร์ดนิ่ง
- เก็บ idle 1000 sample แต่ shaking แค่ 50 sample
- อัตราสุ่มแกว่งไปมาเดี๋ยวเร็วเดี๋ยวช้า
- เก็บหลายคนหลายท่าทางให้หลากหลาย
Show answer
Answer: A. กดปุ่ม shaking แต่วางบอร์ดนิ่ง · B. เก็บ idle 1000 sample แต่ shaking แค่ 50 sample · C. อัตราสุ่มแกว่งไปมาเดี๋ยวเร็วเดี๋ยวช้า
label ผิด คลาสไม่สมดุล และอัตราไม่คงที่ ล้วนสอนโมเดลผิดแบบเงียบ ๆ ส่วนความหลากหลายของสภาพจริงทำให้โมเดลทนขึ้น
-
Hand shaking reaches about 10 Hz. What is the minimum sampling rate by the Nyquist rule? (Objective 2)
- 5 Hz
- 10 Hz
- 20 Hz
- 100 Hz
Show answer
Answer: C. 20 Hz
fs ≥ 2·fmax = 2 × 10 = 20 Hz ถ้าต่ำกว่านี้จะเกิด aliasing การเก็บที่ 50 Hz จึงเผื่อไว้พอ
-
Collecting N = 150 samples at fs = 50 Hz, how long is one burst? (Objective 2)
- 0.33 วินาที
- 3 วินาที
- 7.5 วินาที
- 200 วินาที
Show answer
Answer: B. 3 วินาที
T = N / fs = 150 / 50 = 3 วินาที
-
If you record the dataset at 25 Hz to train or compare with a model trained at 50 Hz, what is the problem? (Objective 3)
- ไม่มีปัญหา เพราะค่าแต่ละ sample เหมือนเดิม
- รูปคลื่นตามเวลาจะห่างและยืดต่างจากที่โมเดลเคยเห็น โมเดลจึงอาจทายผิด
- ไฟล์จะใหญ่ขึ้นสองเท่า
- บอร์ดจะเขียนไฟล์ไม่ได้
Show answer
Answer: B. รูปคลื่นตามเวลาจะห่างและยืดต่างจากที่โมเดลเคยเห็น โมเดลจึงอาจทายผิด
อัตราสุ่มตอนเก็บข้อมูลต้องเท่ากับตอนใช้งานจริง นี่คือหลักสำคัญของ Edge AI dataset สวยแค่ไหนก็ใช้ไม่ได้ถ้าอัตราไม่ตรง
-
The logger records three bursts but the file only keeps the last one. What is the likely cause? (Objective 4)
- เปิดไฟล์ในลูปบันทึกด้วยโหมด "w" ซึ่งทับไฟล์ทุกครั้ง แทนที่จะใช้ "a"
- ลืมเขียน \n ท้ายบรรทัด
- เรียก motion() แทน acceleration()
- flash เต็ม
Show answer
Answer: A. เปิดไฟล์ในลูปบันทึกด้วยโหมด "w" ซึ่งทับไฟล์ทุกครั้ง แทนที่จะใช้ "a"
โหมด "a" เขียนต่อท้ายโดยไม่ลบของเก่า ใช้ "w" ได้เฉพาะตอนสร้างไฟล์ใหม่พร้อมหัวตาราง
Cite this lesson
If you teach from this lesson or reuse it in slides or documents, credit it with the text below. If you changed it, add (adapted) after the title.
"Sampling to match the model: rate, Nyquist, windows and the CSV schema" from TESA Open Knowledge by the Thai Embedded Systems Association (TESA), https://github.com/tesaiot/tesa-qualification-program, licensed under CC BY-NC 4.0
Thai attribution: "สุ่มสัญญาณให้ตรงกับโมเดล: อัตราสุ่ม Nyquist หน้าต่าง และ schema ของ CSV" จาก TESA Open Knowledge โดยสมาคมสมองกลฝังตัวไทย (Thai Embedded Systems Association: TESA) https://github.com/tesaiot/tesa-qualification-program สัญญาอนุญาต CC BY-NC 4.0
Lesson link: https://tesaiot.github.io/tesa-qualification-program/en/courses/edge-ai-developer/m02-daq/l01-sampling-and-schema/
TESA Open Knowledge · © 2026 สมาคมสมองกลฝังตัวไทย (TESA) · CC BY-NC 4.0
Content is licensed CC BY-NC 4.0. Reuse it non-commercially and credit the Thai Embedded Systems Association (TESA) every time. · How to cite TESA