Methodology · every number explained
How each number is made.
Aloud measures the take, never the person, and every figure is a plain signal you could compute yourself. Same take, same numbers, always. Here is exactly how each one is built, from your audio, on your device.
Pace, as an honest band
A syllable detector finds the beats of your speech (onsets of vowel energy). We count those over your speaking time to get syllables per minute, then divide by 1.45, the average syllables per word in conversational English, to estimate words per minute. We show it as a range, not a single number, because estimating words from syllables is inexact and false precision would be dishonest. With the Language lens on, your exact word count gives an exact WPM instead.
Pauses, by how long you held them
Every silence between speech runs is measured and sorted by length:
- Micro — under 0.3s. The natural gaps between words.
- Breath — 0.3 to 1.0s. A breath, a beat.
- Deliberate — 1.0 to 3.0s. A pause you chose, for weight.
- Long — over 3.0s. Held on purpose it lands; caught off guard, a breath covers it.
The longest single pause is reported on its own. A speech run longer than about 25 seconds with no breath is flagged as a rush, not because fast is bad, but because you deserve to know where you did not stop for air.
Pitch variance, against your own baseline
MicCheck records your natural median pitch. During a take we track your voiced pitch, convert it to semitones relative to that baseline (so it is your range, not a comparison to anyone), and take the standard deviation. That number is your pitch variance:
- Flat — under 1.5 semitones of movement.
- Natural — 1.5 to 3.5 semitones.
- Animated — 3.5 semitones or more.
A rolling version of the same measure marks the stretches where you held one pitch too long, the amber ticks on your timeline. None of this is a score. Flat is not a failure; it is information, and sometimes it is exactly right.
Energy, and whether you landed the ending
We take the loudness envelope (RMS) across the whole take and compare the first tenth to the last tenth. The ending drop is (open − close) / open × 100, as a percent. A take that closes within 15% of how it opened has landed its ending; a bigger drop means you trailed off, which is the single most common way a strong take loses its last line.
The words (Language lens, opt-in)
From an on-device transcript, never networked, we compute plain text signals:
- Fillers: counted with the exact words (so, um, like, basically…), and per hundred words.
- Hedges: softeners (I think, sort of, maybe…) that blur a claim, per hundred words.
- Vocabulary: unique words, and a length-robust richness measure (moving-average type-token ratio).
- Sentence rhythm: average length and its spread, so we can tell you when every sentence is the same length.
- Readability: the standard Flesch Reading Ease,
206.835 − 1.015 × (words/sentence) − 84.6 × (syllables/word), shown as plain, clear, or dense.
Why it is deterministic
None of these use a model, a random seed, or a network call in Tier 1. They are arithmetic on your audio. A test suite runs fixed recordings through the engine and fails the build if any number changes. That is what lets us promise: the same take always gives the same reading, and every reading is one you could check.
The privacy side of this is on the privacy page. If any number here is ever computed differently than described, that is a bug, and the feedback page is where to say so.