Skip to content

Sampling to match the model: rate, Nyquist, windows and the CSV schema

Module 2 — Collecting sensor data (DAQ) · Slides: slides.md · Module overview · Course page

Start the cycle upstream: understand what DAQ is, why the sampling rate must be steady and match what the model was trained on (50 Hz), use the Nyquist rule and the window-length formula T = N/fs, and design the CSV schema written to the board’s flash.

By the end of this lesson, you will:

  1. Explain why DAQ is stage 1 of the data cycle, and name at least three things that spoil a dataset (garbage in, garbage out).
  2. Use the Nyquist rule fs ≥ 2·fmax to judge whether a given sampling rate is enough for a signal, and compute the length of one burst with T = N/fs.
  3. Explain why the logger samples at 50 Hz to match the Motion model on the board, and what happens to the waveform if you record at another rate.
  4. Write the schema label,ax,ay,az,gx,gy,gz and a one-sample CSV line correctly, and choose the “a” or “w” file mode for the situation.

You’ve been through module 1, and know the four-beat skeleton and sensors.bmi270.motion(). Keep BENTO IDE’s REPL open, and try calling sensors.bmi270.motion() while still, compared with while shaking.

DAQ (Data Acquisition) means collecting raw sensor data into a set you can store, search, and use later — the starting raw material for every model. The Motion model we played with in module 1 came into being because someone shook a sensor and collected thousands of samples first. A dataset’s quality depends on correct labels, a steady sampling rate, balanced classes, and variety in real-world conditions. Garbage in, garbage out.

The sampling rate is how many times per second the sensor is read. 50 Hz means reading every 20 ms, so the file sets RATE_MS = 20 — not a random number, but exactly the rate the Motion model on the board was trained on. Recording at a different rate stretches or shrinks the waveform, so it no longer looks like what the model has seen. The Nyquist rule says you must sample at least twice the highest frequency you want to capture, $f_s \ge 2 f_{max}$ — otherwise a fast signal disguises itself as a fake slow wave (aliasing). One burst’s length is $T = N / f_s$ — for example, BURST = 200 at 50 Hz gives 4 seconds, long enough to complete a full motion, since a motion is a pattern over time — the model has to see many samples in a row to tell circle apart from shaking.

The CSV schema is the first line, label,ax,ay,az,gx,gy,gz — a contract between whoever collects the data and whoever trains on it. Every following line is one snapshot of the IMU’s six axes with the motion’s name attached. The file is written to the board’s flash (LittleFS) with open()/write(), the same way as Python on a computer. Mode "a" appends; "w" overwrites. Only use "w" when you’re certain the file doesn’t exist yet. Before writing code, it’s worth querying the sensor in the REPL first, to see what “still” and “shaking” values actually look like, so you can tell whether what lands in the file “makes sense.”

The same questions are in quiz.yaml for automated checking.

  1. Which of these spoil a dataset even though the program never throws an error? (select every correct answer) (multiple choice · objective 1)

    • a) Pressing the shaking button while the board sits still
    • b) Collecting 1000 idle samples but only 50 shaking samples
    • c) A sampling rate that drifts, sometimes fast, sometimes slow
    • d) Collecting many people doing many motions, for variety
    Solution

    a, b, c — a wrong label, unbalanced classes, and an unsteady rate all quietly teach the model the wrong thing. Variety in real-world conditions, on the other hand, makes a model more robust.

  2. A hand shake tops out around 10 Hz. What’s the minimum sampling rate the Nyquist rule requires? (single choice · objective 2)

    • a) 5 Hz
    • b) 10 Hz
    • c) 20 Hz
    • d) 100 Hz
    Solution

    c — fs ≥ 2·fmax = 2 × 10 = 20 Hz. Below this, aliasing occurs. Recording at 50 Hz leaves comfortable headroom.

  3. Collecting N = 150 samples at fs = 50 Hz, how long is one burst? (single choice · objective 2)

    • a) 0.33 seconds
    • b) 3 seconds
    • c) 7.5 seconds
    • d) 200 seconds
    Solution

    b — T = N / fs = 150 / 50 = 3 seconds.

  4. If you collect a dataset at 25 Hz, then train or compare against a model trained at 50 Hz, what’s the problem? (single choice · objective 3)

    • a) No problem, since each sample’s value is still the same
    • b) The waveform over time is spaced out and stretched differently from what the model has seen, so it may predict wrong
    • c) The file becomes twice as large
    • d) The board can no longer write the file
    Solution

    b — the sampling rate used while collecting data must match the rate used in actual deployment. This is a core rule of edge AI: no matter how good a dataset looks, it’s useless if the rate doesn’t match.

  5. A logger is used to record three bursts, but the file ends up with only the last burst. What’s the likely cause? (single choice · objective 4)

    • a) The file is opened inside the recording loop with mode “w”, which overwrites it every time, instead of using “a”
    • b) Forgot to write \n at the end of each line
    • c) Called motion() instead of acceleration()
    • d) Flash is full
    Solution

    a — mode “a” appends without erasing old content. Use “w” only when creating a brand new file with its header row.

  • In the REPL, call sensors.bmi270.motion() while still and while shaking, and note the values in your learning log.
  • Compute T = N/fs for BURST = 200 at 50 Hz, and at 25 Hz.
  • Write the schema and one example line, by hand, from a sample you actually read.

In lesson 2.2, we’ll write a real logger following the four beats — schema → sample → record → rate — and collect our first dataset file.

Next lesson: lesson 2.2 — Hands-on: a DAQ logger that saves a dataset to CSV

  • If your sensor captures a signal that oscillates as fast as 12 Hz, what sampling rate would you set, and why not set it higher than necessary?
  • What effect does one wrongly labelled burst have on a model, among thousands of lines in a dataset?

Review questions

Answer on your own first, then open the answer.

  1. Which of these spoil a dataset even though the program raises no error? (select all that apply) (Objective 1)

    1. กดปุ่ม shaking แต่วางบอร์ดนิ่ง
    2. เก็บ idle 1000 sample แต่ shaking แค่ 50 sample
    3. อัตราสุ่มแกว่งไปมาเดี๋ยวเร็วเดี๋ยวช้า
    4. เก็บหลายคนหลายท่าทางให้หลากหลาย
    Show answer

    Answer: A. กดปุ่ม shaking แต่วางบอร์ดนิ่ง · B. เก็บ idle 1000 sample แต่ shaking แค่ 50 sample · C. อัตราสุ่มแกว่งไปมาเดี๋ยวเร็วเดี๋ยวช้า

    label ผิด คลาสไม่สมดุล และอัตราไม่คงที่ ล้วนสอนโมเดลผิดแบบเงียบ ๆ ส่วนความหลากหลายของสภาพจริงทำให้โมเดลทนขึ้น

  2. Hand shaking reaches about 10 Hz. What is the minimum sampling rate by the Nyquist rule? (Objective 2)

    1. 5 Hz
    2. 10 Hz
    3. 20 Hz
    4. 100 Hz
    Show answer

    Answer: C. 20 Hz

    fs ≥ 2·fmax = 2 × 10 = 20 Hz ถ้าต่ำกว่านี้จะเกิด aliasing การเก็บที่ 50 Hz จึงเผื่อไว้พอ

  3. Collecting N = 150 samples at fs = 50 Hz, how long is one burst? (Objective 2)

    1. 0.33 วินาที
    2. 3 วินาที
    3. 7.5 วินาที
    4. 200 วินาที
    Show answer

    Answer: B. 3 วินาที

    T = N / fs = 150 / 50 = 3 วินาที

  4. If you record the dataset at 25 Hz to train or compare with a model trained at 50 Hz, what is the problem? (Objective 3)

    1. ไม่มีปัญหา เพราะค่าแต่ละ sample เหมือนเดิม
    2. รูปคลื่นตามเวลาจะห่างและยืดต่างจากที่โมเดลเคยเห็น โมเดลจึงอาจทายผิด
    3. ไฟล์จะใหญ่ขึ้นสองเท่า
    4. บอร์ดจะเขียนไฟล์ไม่ได้
    Show answer

    Answer: B. รูปคลื่นตามเวลาจะห่างและยืดต่างจากที่โมเดลเคยเห็น โมเดลจึงอาจทายผิด

    อัตราสุ่มตอนเก็บข้อมูลต้องเท่ากับตอนใช้งานจริง นี่คือหลักสำคัญของ Edge AI dataset สวยแค่ไหนก็ใช้ไม่ได้ถ้าอัตราไม่ตรง

  5. The logger records three bursts but the file only keeps the last one. What is the likely cause? (Objective 4)

    1. เปิดไฟล์ในลูปบันทึกด้วยโหมด "w" ซึ่งทับไฟล์ทุกครั้ง แทนที่จะใช้ "a"
    2. ลืมเขียน \n ท้ายบรรทัด
    3. เรียก motion() แทน acceleration()
    4. flash เต็ม
    Show answer

    Answer: A. เปิดไฟล์ในลูปบันทึกด้วยโหมด "w" ซึ่งทับไฟล์ทุกครั้ง แทนที่จะใช้ "a"

    โหมด "a" เขียนต่อท้ายโดยไม่ลบของเก่า ใช้ "w" ได้เฉพาะตอนสร้างไฟล์ใหม่พร้อมหัวตาราง

Cite this lesson

If you teach from this lesson or reuse it in slides or documents, credit it with the text below. If you changed it, add (adapted) after the title.

"Sampling to match the model: rate, Nyquist, windows and the CSV schema" from TESA Open Knowledge by the Thai Embedded Systems Association (TESA), https://github.com/tesaiot/tesa-qualification-program, licensed under CC BY-NC 4.0

Thai attribution: "สุ่มสัญญาณให้ตรงกับโมเดล: อัตราสุ่ม Nyquist หน้าต่าง และ schema ของ CSV" จาก TESA Open Knowledge โดยสมาคมสมองกลฝังตัวไทย (Thai Embedded Systems Association: TESA) https://github.com/tesaiot/tesa-qualification-program สัญญาอนุญาต CC BY-NC 4.0

Lesson link: https://tesaiot.github.io/tesa-qualification-program/en/courses/edge-ai-developer/m02-daq/l01-sampling-and-schema/

Full guide: how to cite TESA

TESA Open Knowledge · © 2026 สมาคมสมองกลฝังตัวไทย (TESA) · CC BY-NC 4.0

Content is licensed CC BY-NC 4.0. Reuse it non-commercially and credit the Thai Embedded Systems Association (TESA) every time. · How to cite TESA