Features and windowing: what the model actually sees
Module 4 — Signal analysis · Slides: slides.md · Module overview · Course page
Understand what a model actually sees. A model doesn’t eat raw samples one point at a time — it eats windows squeezed into a feature vector. Learn about sliding windows and hop, why they must overlap, statistical features (mean, std, per-band energy), and the link to the log-mel spectrogram used by audio models.
Objectives
Section titled “Objectives”By the end of this lesson, you will:
- Explain why a model must look at a window of signal instead of a single sample, and compute windows per second from WIN, HOP and the sample rate.
- State the trade-off of window overlap (small versus large hop), and why overlapping windows don’t miss short events.
- Compute the mean and std of a small window by hand, and explain how std separates still from moving, while mean shows a steady posture.
- Explain the log-mel pipeline log(Mel(|FFT(x·w)|)), and why the front-end must match between training and deployment.
Before you start
Section titled “Before you start”You’ve been through lessons 4.3–4.4, and understand the FFT, the Hann window, and the magnitude. Open the s10_windowing.py example in BENTO IDE.
- Hardware: a TESAIoT Dev Kit board already flashed with BENTO’s MicroPython firmware, or the BENTO Emulator inside BENTO IDE
- Prior lesson: lesson 4.4 — Hands-on: a live spectrum from the IMU
See it work first
Section titled “See it work first”Run s10_windowing.py right away. Alternate lying the board still and shaking it gently, and watch the mean, std and band0..3 bars move. That’s the feature vector flowing into the model — one chunk of raw signal becoming just a handful of numbers.
Concepts
Section titled “Concepts”A single cough or a single shake isn’t a value at one point — it’s the shape of a signal over a stretch of time. A model therefore eats a window — for example, WIN = 50 points, or one second at 50 Hz. The window slides forward by HOP = 25 points at a time, so it overlaps 50%, giving a new feature set every 25 points (twice a second). Overlapping makes the response quicker and stops short events falling into the gap between windows, at the cost of computing more often. The window size isn’t set arbitrarily — whatever size the model was trained with, deployment must feed it the exact same size.
A 50-point window is squeezed into a feature vector of six values: [mean, std, band0, band1, band2, band3]. mean shows the DC level, or the direction of gravity on that axis (a steady posture). std = $\sqrt{\frac{1}{n}\sum (x_i - \bar{x})^2}$ shows how strong the shaking is — low at rest, high when shaken. band0..3 splits the window into four segments in time, then measures the variance of each segment, catching whether the shaking is clustered near the start, middle, or end of the window. A model is only as small and accurate as its features are good.
The board’s audio models use the same skeleton, but split into bands by frequency instead: slice a window → multiply by Hann → FFT → take the magnitude → combine into mel bands, matching how the ear hears (fine resolution at low frequency, coarse at high) → log to compress the dynamic range. It can be written as one formula: $\text{logmel} = \log(\mathrm{Mel}(|\mathrm{FFT}(x \cdot w)|))$. The most important takeaway is: a model never sees the raw signal. It only ever sees the feature vector a front-end squeezes for it. During training, the dataset is a set of feature vectors with labels; in deployment, the board squeezes a live signal the exact same way. If the two front-ends don’t match, the model breaks instantly — a classic edge AI bug.
Worked example
Section titled “Worked example”s10_windowing.py is the reference version, using WIN = 50 and HOP = 25. Predict before running it: which bar will rise the most when shaken? Then try rotating the board slowly and compare — you’ll see mean change while std stays low.
| File | What this file teaches |
|---|---|
| examples/s10_windowing.py | What a model “sees”: windowing + a feature vector |
Check your understanding
Section titled “Check your understanding”The same questions are in quiz.yaml for automated checking.
-
WIN = 50, HOP = 25, at a sample rate of 50 Hz. How many new feature vectors per second? (single choice · objective 1)
- a) 1
- b) 2
- c) 25
- d) 50
Solution
b — a new set comes every HOP = 25 points. At 50 points per second, that’s 2 sets per second, each covering 1 second and overlapping 50%.
-
You set HOP = WIN (no overlap). What’s the main risk? (single choice · objective 2)
- a) Memory fills up
- b) A short event straddling the boundary between two windows might get split in half and missed by the model, and results come more slowly
- c) std becomes negative
- d) The sampling rate changes
Solution
b — overlap covers every point with more than one window, and produces results more often. No overlap saves effort but responds more slowly and misses more easily.
-
A window [7.8, 11.8, 7.8, 11.8] — what are its mean and std? (single choice · objective 3)
- a) mean 9.8, std 0
- b) mean 9.8, std 2.0
- c) mean 4.0, std 9.8
- d) mean 11.8, std 4.0
Solution
b — the average is 9.8; every point is 2.0 away from the average, so std is 2.0. Compare with a still window [9.8, 9.8, 9.8, 9.8], where std = 0 even though the mean is the same.
-
You rotate the board slowly from flat to upright. Which feature changes the most? (single choice · objective 3)
- a) mean, because gravity’s direction on the Z axis changes; std stays low since there’s no shaking
- b) std, because the board is moving
- c) band3 only
- d) No feature changes
Solution
a — mean shows steady posture, std shows shaking. A slow rotation changes az’s average level but barely makes the value oscillate.
-
Put the pipeline that builds a log-mel spectrogram from one audio window in order (ordering · objective 4)
- a) FFT
- b) Multiply by the Hann window (x · w)
- c) log
- d) Take the magnitude |·|
- e) Combine into Mel bands
Solution
b → a → d → e → c — x·w → FFT → |·| → Mel → log. If training and on-board deployment don’t do these steps identically, the model’s results go wrong.
- Compute the number of windows per second when WIN = 50, HOP = 25 at 50 Hz, and when HOP = 50.
- Compute the mean and std of the windows [9.8, 9.8, 9.8, 9.8] and [7.8, 11.8, 7.8, 11.8] by hand, and note them in your learning log.
- Write out the log-mel pipeline step by step in your own words, saying which step you’ve already done in lesson 4.4.
Going further
Section titled “Going further”In lesson 4.6, we’ll fill in the s10_windowing.py file to collect a buffer, squeeze features, and slide the window ourselves.
Next lesson: lesson 4.6 — Hands-on: a feature vector from a sliding window
Reflect
Section titled “Reflect”- If you had to tell “walking” apart from “running,” which of these six features do you think would help most, and what’s still missing?
- Why is it dangerous to change the window size after a model has already been trained?
Review questions
Answer on your own first, then open the answer.
-
With WIN = 50 and HOP = 25 at 50 Hz, how many new feature vectors per second? (Objective 1)
- 1 ชุด
- 2 ชุด
- 25 ชุด
- 50 ชุด
Show answer
Answer: B. 2 ชุด
ได้ชุดใหม่ทุก HOP = 25 จุด ที่ 50 จุดต่อวินาทีจึงเป็น 2 ชุดต่อวินาที แต่ละชุดครอบคลุม 1 วินาทีและซ้อนกัน 50%
-
Setting HOP = WIN (no overlap) mainly risks what? (Objective 2)
- หน่วยความจำเต็ม
- เหตุการณ์สั้นที่คร่อมรอยต่อระหว่างสองหน้าต่างอาจถูกแบ่งครึ่งจนโมเดลพลาด และได้ผลช้าลง
- std กลายเป็นลบ
- อัตราสุ่มเปลี่ยน
Show answer
Answer: B. เหตุการณ์สั้นที่คร่อมรอยต่อระหว่างสองหน้าต่างอาจถูกแบ่งครึ่งจนโมเดลพลาด และได้ผลช้าลง
การซ้อนทำให้ทุกจุดถูกครอบด้วยหน้าต่างมากกว่าหนึ่งอัน และได้ผลบ่อยขึ้น ไม่ซ้อนประหยัดแรงแต่ตอบช้าและพลาดง่าย
-
What are the mean and std of the window [7.8, 11.8, 7.8, 11.8]? (Objective 3)
- mean 9.8, std 0
- mean 9.8, std 2.0
- mean 4.0, std 9.8
- mean 11.8, std 4.0
Show answer
Answer: B. mean 9.8, std 2.0
ค่าเฉลี่ย 9.8 ทุกจุดห่างจากค่าเฉลี่ย 2.0 std จึงเป็น 2.0 เทียบกับหน้าต่างนิ่ง [9.8, 9.8, 9.8, 9.8] ที่ std = 0 แม้ mean เท่ากัน
-
You slowly rotate the board from flat to upright. Which feature changes most? (Objective 3)
- mean เพราะทิศของแรงโน้มถ่วงบนแกน Z เปลี่ยน ส่วน std ยังต่ำเพราะไม่มีการสั่น
- std เพราะบอร์ดเคลื่อนที่
- band3 เท่านั้น
- ไม่มี feature ใดเปลี่ยน
Show answer
Answer: A. mean เพราะทิศของแรงโน้มถ่วงบนแกน Z เปลี่ยน ส่วน std ยังต่ำเพราะไม่มีการสั่น
mean บอกท่าทางคงที่ std บอกการสั่น การหมุนช้า ๆ เปลี่ยนระดับเฉลี่ยของ az แต่แทบไม่ทำให้ค่าแกว่ง
-
Order the pipeline that builds a log-mel spectrogram from one audio window. (Objective 4)
- FFT
- คูณ Hann window (x · w)
- log
- เอาขนาด |·|
- รวมเป็นย่าน Mel
Show answer
Correct order: B. คูณ Hann window (x · w) → A. FFT → D. เอาขนาด |·| → E. รวมเป็นย่าน Mel → C. log
x·w → FFT → |·| → Mel → log ถ้าตอนฝึกกับตอนใช้บนบอร์ดทำขั้นเหล่านี้ไม่ตรงกัน ผลของโมเดลจะเพี้ยน
Cite this lesson
If you teach from this lesson or reuse it in slides or documents, credit it with the text below. If you changed it, add (adapted) after the title.
"Features and windowing: what the model actually sees" from TESA Open Knowledge by the Thai Embedded Systems Association (TESA), https://github.com/tesaiot/tesa-qualification-program, licensed under CC BY-NC 4.0
Thai attribution: "feature และหน้าต่าง: สิ่งที่โมเดลเห็นจริง" จาก TESA Open Knowledge โดยสมาคมสมองกลฝังตัวไทย (Thai Embedded Systems Association: TESA) https://github.com/tesaiot/tesa-qualification-program สัญญาอนุญาต CC BY-NC 4.0
TESA Open Knowledge · © 2026 สมาคมสมองกลฝังตัวไทย (TESA) · CC BY-NC 4.0
Content is licensed CC BY-NC 4.0. Reuse it non-commercially and credit the Thai Embedded Systems Association (TESA) every time. · How to cite TESA