The 1000 Most Common Japanese Words: The Core, and the Words Only You Keep Meeting
A frequency list is not a syllabus. It is the diary a corpus kept: someone assembled millions of words of real Japanese, counted everything that appeared in the pile, and published the tally in descending order. That is all any page titled "the 1000 most common Japanese words" is. No teacher curated it, no exam board blessed it. It fell out of whatever pile of text happened to get counted, and the pile decides the answer far more than the counting does. So the honest version of this page starts with the piles. Usual disclosure: I build Piccard, a flashcard tool whose deepest language support is Japanese, and my argument below is that the list is real, worth learning, and still only the smaller half of the job.
The piles the numbers are counted from
The heavyweight source is the Balanced Corpus of Contemporary Written Japanese, BCCWJ, built by the National Institute for Japanese Language and Linguistics (NINJAL): 104.3 million words of written Japanese, sampled at random across registers that include general books, magazines, newspapers, business reports, blogs, and net forums, with texts spanning roughly 1976 to 2006. It is a research instrument, not a study aid, and you can query its raw counts yourself, free, through NINJAL's Shonagon tool. When a vocabulary site claims its list reflects "actual Japanese usage," this corpus, or some much smaller count like it, is usually what is sitting underneath the claim.
Where the lists on the open web actually inherit their numbers from, and what each inheritance costs:
| Where the list comes from | What it actually counts | The catch |
|---|---|---|
| BCCWJ, NINJAL's balanced corpus | 104.3 million words of written Japanese across books, magazines, newspapers, reports, blogs, and forums, 1976 to 2006 | Written only; not one word of it was ever spoken. Raw counts, not a study order |
| Newspaper frequency counts | A run of one or more papers' news pages | Register skews formal and political: 首相 ranks, casual conversation does not exist here |
| CSJ, NINJAL's spoken corpus | Roughly 7.5 million words of spontaneous Japanese speech | Under a tenth of BCCWJ's size, so its ranks wobble more; speech is its own register |
| Frequency-ordered starter decks (Kaishi 1.5k, the Core series) | A frequency list that editors already cut, ordered, and attached sentences to | Curation is judgment: the order is editorial, no longer raw counts |
| Textbook and JLPT vocabulary lists | What a syllabus chapter or an exam level expects you to know | Not frequency lists at all: ordered by curriculum, blind to usage |
That last row deserves saying out loud, because the costume gets worn a lot. A Genki chapter orders its vocabulary by grammar point; the JLPT orders by level; a "top 1000 words" page that quietly copies a textbook index is handing you a syllabus while calling it a corpus. The distinction is not academic snobbery. Frequency data tells you what Japanese runs on. A syllabus tells you what a course plans to teach. Those are different claims, and only one of them is what you searched for.
The spoken row is the one people skip and shouldn't. NINJAL and its partners (NICT and the Tokyo Institute of Technology) built the CSJ because spontaneous speech behaves differently from prose: polite forms climb, fillers and casual verbs surface, newspaper nouns evaporate. A list counted from television transcripts and a list counted from a balanced written corpus agree on the broad middle and disagree at the edges, which is exactly where your personal interests live.
The words at the top of every table
Ask several frequency lists for their exact top ten and you will get several different orders. You will also get nearly the same members. The safest claim in Japanese vocabulary is that these words live at the very top of every serious count, so here they are, deliberately unranked, because arguing about whether する edges out ある this quarter is a waste of your afternoon:
| Word | Reading | Meaning |
|---|---|---|
| する | する | to do |
| ある | ある | to exist (things) |
| いる | いる | to exist (people, animals) |
| です | です | to be (polite) |
| だ | だ | to be (plain) |
| こと | こと | thing, matter (abstract) |
| 言う | いう | to say |
| なる | なる | to become |
| 私 | わたし | I |
| これ | これ | this |
| それ | それ | that |
Exact ranks shuffle by corpus, by era, and by whether the counter included particles. If a list does include them, は and を sweep every ranking while teaching almost nothing in isolation, which is why learning-oriented lists usually drop particles and start with words like the ones above. Look at the company they keep: verbs of being, doing, saying, becoming, plus pronouns. The glue of the language arrives before anything the glue would hold together.
And notice the second property, the one that becomes this post's whole point: the most frequent words are the least self-explanatory. する does. こと is. Neither has a translation you could write on a flashcard and be usefully done with, because their job is grammatical, and grammar only means something inside a sentence. Hold that thought.
What a rank tells you, and what it can't
The case for frequency-first learning is genuine and I don't want to talk you out of it. A thin slice of the dictionary does a disproportionate share of the work in running text, and the returns flatten as you descend: your first thousand words earn their keep far faster than your fifth thousand ever will. Every corpus shows this shape. The exact coverage percentages shift with the corpus and the counting method, so treat any page quoting coverage to a decimal place as decoration, but the curve itself is not in dispute.
What the rank cannot tell you is what you will actually be reading and hearing. A thousand words counted from newspapers equips you for 首相 and 国会 before it equips you for the casual verbs of a slice-of-life monologue. A spoken corpus does the reverse. Neither list is wrong; each is a faithful diary of a different life. The question no list can answer is which life is yours, and a learner who watches cooking channels and a learner who reads financial news need different second thousands while needing almost the same first one.
There is also the inertness problem, and it is the big one. Knowing that こと ranks near the top does precisely nothing to the part of you that must recognize it at conversational speed. A list is a map of where the weight sits. Walking the territory is a separate activity, and the territory that matters is whatever you personally keep watching, reading, and re-meeting. Every learner's deck eventually splits into two zones: a shared core that any good frequency order covers, and a personal tail of words that are common in your material and nowhere else on earth. The core is a list problem. The tail is a capture problem.
Build the core, then capture the tail
Two halves, two different shapes of work. Here is how I would run both through Piccard, Japanese direction, described as it ships.
Rank the core first. Take whichever frequency-derived list or deck you fancy, in any reputable form, and feed it to the app as a document: pasted text, a PDF, a photo of a printed page. Every word comes back as a candidate card tagged high, medium, or low for how much it is worth learning, sorted with the worthwhile ones on top and the high tags already checked off, so the few hundred you already half-know from anime, games, or a semester of class sit at the bottom and get skipped in one pass instead of ground through for a month. One confirmation turns at most 20 candidates into cards. The 20 is a named constant in our codebase, not a growth mechanic: the model spends roughly 300-350 output tokens per card, and 20 keeps a confirmation well under the output ceiling. Japanese cards arrive with the reading attached, so 私 lands as 私(わたし), never as naked kanji you can't pronounce.
Capture the tail where you meet it. On a YouTube video that has Japanese captions, the extension shows the transcript line by line. Click one caption line, and candidate cards come back from that line: a word, its meaning, an example sentence. Keep what is worth keeping, discard the rest, keep watching. The clicked line only; there is no convert-the-whole-video button, and a video without captions has no lines to click, so the tool says so instead of improvising. This is the mechanism a frequency list can never be. The words you keep meeting in what you actually watch arrive with the sentence you actually met them in, not an invented one, and when such a card comes up for review it replays the video from that line's timestamp: the voice, the scene, the moment the word meant something to you.
The edges, stated plainly. A double-click on any Japanese word on any web page opens a popup dictionary running on JMdict data, which is the Japanese direction's reading bonus. Uploading audio or video files for transcription is built for English audio only, so Japanese listening material reaches your deck through the caption route above, not through file upload. And candidates cap at 20 per action on every path, same constant, same token arithmetic.
Why sentences matter most for exactly these words. Back to the property of the top of the list: its most frequent members are its most grammatical. する and こと barely have identities in isolation. In a line from a video you chose to watch, they have whatever meaning the scene gave them, and that binding of word to moment is the thing your memory actually stores. A frequency list tells you these words deserve the effort. The sentence tells you what the effort is for.
The schedule that keeps a thousand words
A thousand words learned once is a thousand words half-gone by spring. The decay was measured in the 1880s and named the forgetting curve, and the countermeasure, reviewing each item just before it drops with the gap lengthening every time you survive, is spaced repetition. Recall is the event that strengthens memory; passing your eyes over a list row is barely studying at all. What is newer than the science is not having to configure it: scheduling in Piccard is FSRS-6 with zero configuration, a 0.9 target retention, four ratings (again, hard, good, easy), and those four buttons are the entire interface between you and the algorithm.
The plan I would run:
- Load your frequency list, rank it, and start at 10 new cards a day from the top. At that pace a thousand-word core is roughly 100 days of intake, with review volume riding on top.
- Empty the review queue every day, no exceptions. The schedule only works if it is allowed to run.
- From the first captioned video you genuinely care about, let capture overtake the list as your card source. The core is finite; the tail never is.
- When a core card comes up, say the reading aloud before you flip. Frequency words you can recognize but not pronounce are half-learned, and half of them are coming to you as audio soon enough.
If your interest in lists runs exam-shaped, the sibling problem, why no two N5 lists agree and what the old spec actually said, is covered in JLPT N5 Vocabulary: The List That Isn't Official. If you would rather someone else already did the frequency-ordering for your core, Kaishi 1.5k: What the Default Japanese Starter Deck Is is precisely that, sentences attached. And the wider landscape of shared decks, Core and Tango and the anime decks included, is compared in The Best Anki Decks for Japanese.
The short version
The 1000 most common Japanese words are worth learning, and they are a corpus's diary, not a syllabus. What counts as "most common" is decided by the pile of text that got counted: the BCCWJ's 104.3 million words of written Japanese is the big credible pile, spoken corpora like the CSJ tell a different but overlapping story, and textbook or JLPT lists are not frequency data no matter what the download page says. Take one reputable frequency order, rank out what you already know, learn the core once, and expect its most frequent members to be its most grammatical, the ones that only mean things inside sentences. Then stop asking lists for the part they cannot know: the words your own material keeps using, saved with their sentence and their scene, on a schedule you never have to tune.
The free tier is 150 energy (about 150 cards, roughly a seventh of a thousand-word core), then prepaid top-ups at $5 / $10 / $20, pricing here; no subscription, and topped-up energy never expires. To start with the caption route: Piccard on the Chrome Web Store.
All guides: Piccard Blog.
Try Piccard for free
Bring your word list — paste the text or drop the PDF, and it becomes FSRS-scheduled cards. 150 free energy to start, no credit card.