Every phrase opens in a room — the screen where you listen to it
and record yourself. At the top, the phrase, with four switches under it
(outlined below):
Translation and Furigana start hidden — read the Japanese
first, and reach for them when you need them. Phrase hides the
text itself, for practicing from the ear alone. Pitch shows or
hides the ribbon.
Listen plays the model voice; tap it again to pause, and again to
carry on from where you stopped. Replay starts the phrase over
from the top. Record opens the microphone — tap it once to start,
once to stop. Reset throws your take away and puts the room back
at the start.
The arrows step through the set's phrases in order — the counter shows
where you are — and Return to set goes back to the set's page.
The loudness band
The teal bars behind the ribbon are the model voice's
loudness, slice by slice: taller is louder, a gap is a
pause. They are the phrase's rhythm made visible — where the beats
fall and how long each part takes.
While you record, your loudness is drawn over it in red,
live, with the voice's bars ghosted underneath as a pacing guide.
If your red shape runs ahead of the ghost or lags behind it, you
see it as it happens.
With Pitch switched on, the band also draws the tune of the
phrase. Five marks, one at a time:
The step — the accent the dictionary teaches. High over high
morae, low over low, with one drop at most in each chunk of the
phrase. This is the tune you are trying to say. It has no scale: high
and low are the only two places it can be.
The curve — the pitch the voice actually produced, measured
from the recording. A real speaker following the step still rises
gradually, peaks a beat late, and drifts downward. It is never
expected to trace the step exactly.
The kana — one per beat, each placed at the moment it was
actually said, at the ribbon's real spacing. This is the start of
練習すれば…: see how the
ウ is squeezed between
シュ and ス — it lasted
a third of a beat, so it gets a third of the room.
Crowded kana are not a display fault — they are the voice
going fast there, and a place worth slowing down. As the phrase
plays, the one being spoken lights orange.
A faint kana was whispered. Japanese routinely drops the
voice on certain vowels: 好きです, here, loses
both the す in 好き and
the final す of です.
There is no pitch to measure on a whisper, so the curve skips them.
Correct Japanese, not missing data.
The orange line is where playback is. It moves as the phrase
plays and stays where you paused. While something is playing the
ribbon is locked; when it is paused or finished you can drag it
sideways to look at the parts that do not fit on screen.
Comparing your take with the voice
After you record, Hear comparison plays the model voice and then
your take, back to back. Today the comparison is for your ears:
while it plays, keep your eyes on the step and ask two different
questions.
Do I sound like the voice? That is about matching a particular
person — Naomi, or Hiro. Does my pitch drop where the step
drops? That is about saying the word with the right pitch accent.
Coming soon
Your pitch drawn in red over the voice's teal, with the step behind
both — so the second question will have a picture, not just your ear.
Recording well
The very first time you tap Record, your browser asks — each browser
draws this a little differently. Choose "allow" (or "always allow",
and it will not ask again). Listening works without it; only
recording needs it.
A headset with a microphone helps. It picks up your voice more
cleanly than a laptop or phone microphone across the room.
Leave a short beat after you tap Record before you start speaking.
The room lines your take up with the model voice by finding your first
sound, so you do not need to match the voice's start — but if the
microphone clips your first beat, the first sound it finds is your
second, and the whole comparison slides one beat out while looking
perfectly precise.
Speeds and voices
0.7× and 0.85× are real slower recordings, not a stretched
1×: every beat keeps its proportion, so the tune and the rhythm are
unchanged — just with more room to hear them. Slow it down for a phrase
that will not stick, then bring it back up.
There are four voices — two female, two male — and the pitch
accent is the same in all of them (Standard/Tokyo pitch), but the way
each one rises and falls is its own. Switching voice changes the target you are copying, which is why the
band's label changes with it. Your choice of voice is remembered on this
device. Speed stays as you set it while you move between phrases, and
starts again at 1× next time you open the room.
Your progress and your takes
On a set's page, the ✓ beside a phrase means you have
practiced it — played it and recorded a take. The highlighted row is
where the set will resume, so you can put it down mid-set and come
back.
The same ticks follow you into the room, on the Nearby list.
That progress, along with your place in each set, is saved
on this device only. There is no account and nothing is sent
anywhere.
Your recordings never leave the browser. A take exists only while
you are in the room; it is gone when you move to the next phrase or
close the page. Nothing you say is stored or uploaded.