Every Japanese word has a tune. Some beats are said a little higher, some a little lower, and the pattern is generally fixed — it belongs to the word the way its spelling does. Get the sounds right but the tune wrong and you are still understood — a native ear just hears that something is off, even if they cannot say what.
The three words below are all pronounced hashi. Only the tune tells them apart — press play and listen for where the voice falls. (Each is played with が after it — a little grammar word that follows nouns. It is there because some tunes only show themselves on the next beat, as you will see below.)
If the three sound almost the same to you, that is normal — nobody taught your ear to listen for this. Being able to hear it is the skill, and it comes from exactly the kind of practice this app is for. Seeing where the fall belongs is how you start.
And a reassurance before any of it: in real conversation, context does most of the work. Ask for hashi at dinner with the wrong tune and you will still get chopsticks. Getting the pitch wrong rarely causes a misunderstanding — it just sounds a little off. This is polish, not survival.
Textbooks skipped pitch accent for decades — it is hard to teach from paper — and native speakers cannot explain it, because they never learned it consciously: ask a native which beat of a word is high and you will usually get a puzzled look, followed by a perfect pronunciation of the word. It has only recently become something learners talk about. If you have studied Japanese for years and never heard of it, nothing went wrong; you are the normal case.
English marks a syllable by saying it louder, longer — and usually higher too, all bundled together: RE-cord the noun, re-CORD the verb. Japanese does not do that. Every beat gets about the same length and weight; what changes is only the pitch, and only between two levels, high and low.
The habit to watch: you will want to say the high beat louder and longer, because that is what English does. Don't — same loudness, same length, only higher.
The unit carrying each high or low is the mora — one beat of the kana: き is one, きょ is one, the long vowel in とう is one more, and ん is one on its own. The ribbon in the practice room writes one kana per beat for exactly this reason.
Every word falls into one of four patterns. The rule underneath all four: the pitch may go up once, and may come down once — and once it has come down, it stays down for the rest of the word. The place it comes down is called the accent. A word either has one or it does not.
Notice 桜 and 男. Said on their own they sound the same — low, then high to the end. The difference only appears on the next word: after 桜 the particle stays high; after 男 it drops. That is why the diagrams include が — and why the ribbon sometimes draws a flat high line to the end of a chunk: the drop, if there is one, belongs to whatever comes next.
A sentence is spoken in chunks — usually a word plus its particles — and each chunk gets its own little rise. "Once down, stays down" holds within a chunk, then the next chunk starts fresh. Here is the phrase from the app's home page, as the model voice says it:
Over a whole sentence the voice also drifts gradually downward, so a rise near the end is smaller than the same rise at the start. The step ignores that drift on purpose; the curve in the practice room shows it.
Japanese dictionaries — the big monolingual ones made for native speakers, and NHK's accent dictionary — write a word's accent as a number: the beat after which the pitch falls, and [0] for no fall at all. So 雨 is [1], ありがとう is [2], 男 is [3], and 桜 is [0]. The learner dictionaries have mostly left it out — though that is changing: Jisho.org is adding it. Wiktionary has it (see below). The number is worth learning to read, because it is the one thing about a word's sound that a dictionary can tell you and a textbook usually does not.
Everything on this page — and everything the model voices say — is the Tokyo (standard) accent, as dictionaries record it. Other regions genuinely differ: Kansai often puts the fall somewhere else entirely. That is dialect, not error — but the standard accent is the one every dictionary documents and the one this app models.
OJAD, the online accent dictionary from the University of Tokyo, shows the pattern for whole phrases. On Wiktionary there is no Japanese section to browse — you simply look a word up. Most Japanese entries list the pitch accent with its dictionary number under Pronunciation: here is 父 as an example, citing the NHK accent dictionary — the standard reference, the one NHK's own announcers use (a print and app dictionary, in Japanese). There are also browser plugins that add the pitch accent to the words you look up in learner dictionaries such as Jisho.org. Several English-speaking YouTubers also teach pitch accent well — searching "Japanese pitch accent" is a good afternoon.
The practice room draws all of this for every phrase — the accent as a step, the voice's pitch as a curve, each kana at its moment in time. See what the ribbon draws →