A lesson in the Japanese Patent Office making Japanese spelling even crazier
This article is for that huge crowd of weebs who also work with trademarks.
Hi to all three of you!
I'm currently improving support for the Japan Patent Office in my trademark monitoring tool, BrandCat. Until recently, I was getting some Japanese trademark data through a third-party provider. Then I decided to go straight to the source and started importing data directly from JPO.
That's where the fun* began :-)
* Not really that fun 😁
If you think the Japanese writing system isn't complicated enough already, the Japanese Patent Office has your back. Consider it the level-boss fight at the end.
TL;DR for people who already know Japanese
JPO publishes a pronunciation field for trademarks called 称呼. Literally, the word means something like "what something is called" or "appellation": 称 relates to naming/designation and 呼 to calling. In trademark practice, it basically means the pronunciation by which the mark is recognized or called.
The pronunciation is represented in Katakana, regardless of whether the underlying word is Japanese, foreign, written in Kanji, Hiragana, Latin letters, or something else.
So far, perfectly sensible**
** Again, not really 😁
But JPO's pronunciation representation does something that looks very strange if you're used to ordinary Katakana: long vowels can be written out as vowel kana instead of using the usual Katakana long-vowel mark ー.
For example, ordinary Japanese writes:
And this isn't something I invented after breaking my parser. JPO's own current documentation gives its pronunciation of "JPO" as:
ジェイピイオオ
Rather than something involving long-vowel marks. JPO describes the field as 称呼, pronunciation, and publishes it in full-width Katakana.
If that already makes sense to you, congratulations. You can skip most of the next section.
For everybody else, welcome to Japanese writing.
Why is this weird in the first place?
Japan doesn't really use three "alphabets". It uses three major writing systems together: Kanji, Hiragana and Katakana.
A perfectly normal Japanese sentence*** can contain all three.
*** A perfectly normal sentence in a language that needs three writing systems to write one perfectly normal sentence 😁
Sometimes Latin letters too, just in case things were becoming too easy.1. Kanji
Kanji are characters originally imported from Chinese.
They're logographic characters, meaning that the characters are associated primarily with meanings rather than simply representing individual sounds like letters in an alphabet.
For example:
means "water".
In Japanese, one common pronunciation is みず, mizu.
In Mandarin Chinese, the same character is pronounced something like shuǐ.
In my own language, Serbian, I'd just call the thing voda.
It's translingual, the meaning is water, not sounds that make the word "water" in any language.
The real fun begins because a single Kanji can have several Japanese readings, including native Japanese readings and readings historically borrowed from varieties of Chinese.
Take Tokyo:
東 means east.
京 means capital.
So 東京 literally means something like "Eastern Capital".
Simple!
Except 東 on its own is very often pronounced ひがし, higashi.
So where the hell did the Tō in Tōkyō come from?
It comes from a Sino-Japanese reading, historically borrowed from Chinese and then adapted to Japanese pronunciation over many centuries.
Basically, medieval international telephone, except everybody is using Chinese characters.
Any questions?
No?
Excellent, because it gets worse.
2. Hiragana
Chinese writing was developed for Chinese.
**** You can claim otherwise if you're in the mood to start a fight 😁
This created certain inconveniences.
Japanese has grammatical particles, verb and adjective inflections, politeness forms and lots of other grammatical machinery that doesn't map neatly onto writing everything with Chinese characters.
Enter Hiragana.
Hiragana is a phonetic script. Each character represents a mora, roughly a rhythmic sound unit.
Take:
This means "delicious".
The first part, 美味, is written in Kanji. The ending しい is Hiragana.
The word as a whole is pronounced:
Oishii.
This is extremely common in Japanese writing. Kanji carry much of the lexical information, while Hiragana handles readings, grammatical endings, particles and words that are normally written without Kanji.
For another simple example:
means "to eat".
食 is Kanji.
べる is Hiragana.
Change the grammar and the Hiragana part changes:
So now we've got Kanji plus Hiragana.
Surely that's enough writing systems for one language?
3. Katakana
Imagine you're a Japanese Buddhist monk roughly a thousand years ago.
You're annotating texts and need a quick phonetic shorthand.
So you take pieces of existing Chinese characters and use those pieces for their sounds.
Eventually, congratulations, you've helped create Katakana.
Katakana developed from abbreviated components of Kanji and today is used heavily for foreign loanwords, foreign names, technical terminology, emphasis, animal and plant names in some contexts, and sound effects. It's also all over manga.
For example, a passing or revving car might go ブロロロロ, while screeching brakes might be キキーッ.
Which means that if you want to write "America":
ア = a
メ = me
リ = ri
カ = ka
America becomes Amerika.
France:
Furansu.
Coffee:
Kōhī.
And that last one is important.
See those horizontal lines?
That's the prolonged sound mark.
In ordinary Katakana, it tells you to stretch the preceding vowel.
So:
is roughly:
ko + long o + hi + long i
Kōhī.
Coffee.
That is completely normal Japanese Katakana.
Remember it. We're going to abuse it later.
Using all three at once
Now suppose you want to say:
"French food is delicious."
You can write:
And congratulations, you've just used all three Japanese writing systems in one perfectly ordinary sentence.
Add a few thousand Kanji, learn the two kana syllabaries, spend the next ten thousand hours discovering that every rule has an exception, and congratulations!
You can now read Japanese.
Sort of.
BUT YOU THOUGHT THAT WAS ENOUGH TO ENTER THE JAPANESE PATENT OFFICE?
No.
Of course not.
Welcome to trademark data.
Pronunciation matters a lot in trademarks.
Imagine somebody files a trademark written entirely in Kanji.
How do you search for similar-sounding marks?
That's already a problem because Japanese Kanji don't uniquely tell you how a word is pronounced.
Names are particularly entertaining in this respect. The same characters can have multiple legitimate readings, and sometimes you're simply not going to guess correctly without being told.
Then add trademarks written in:
- Kanji
- Hiragana
- Katakana
- Latin letters
- invented words
- foreign names
- mixtures of several scripts
And now you have to build a search system where pronunciation matters.
JPO's solution is to maintain a pronunciation field, 称呼.
When searching J-PlatPat by pronunciation, JPO explicitly tells users to enter the reading in full-width Katakana.
That gives them one script in which to represent pronunciation.
Very sensible.
Then I opened the actual data.
Meet コオヒイ
Remember coffee?
Normal Japanese:
Now suppose you don't want the character ー to mean "make the previous vowel longer".
Instead, you want the pronunciation represented using actual kana vowels.
The first long o becomes:
The final long i becomes:
So:
Same intended pronunciation.
Very different string.
If you've ever used a Japanese dictionary, this idea may actually be familiar. Japanese dictionary collation can treat the long-vowel mark as the corresponding vowel when determining sort order. So コーヒー can effectively be treated as コオヒイ for ordering purposes.
But seeing this kind of representation sitting in a trademark pronunciation field is still wonderfully cursed.
And JPO itself gives us an even better example.
The Latin letters:
have the pronunciation field:
Break that apart:
Again, the long vowels are represented explicitly rather than with ー.
And this is where my perfectly innocent little romanization code gets punched in the face.
Why this matters if you're actually processing the data
If all I wanted to do was display the JPO record exactly as supplied, none of this would matter.
Print ジェイピイオオ.
Go to lunch.
But BrandCat has to search and compare trademarks.
A user searching for a Latin trademark isn't necessarily going to know that a JPO pronunciation field has been normalized in this particular way.
They'll search for something like:
JPO
or perhaps a romanized Japanese pronunciation.
Likewise, if a trademark contains コーヒー, I don't want my search system to consider コーヒー and コオヒイ completely unrelated strings just because one uses ordinary modern Katakana orthography and the other represents the long vowels explicitly.
So somewhere in the ingestion and matching pipeline, you need to know what you're actually looking at.
You can't just say:
"Katakana detected. Romanize string. Job done."
Because garbage in, garbage out.
Or more precisely:
government-standardized phonetic Japanese in, hilariously wrong Latin string out.
And Google almost made it even better
While trying to work out exactly what JPO was doing, I asked Google about it.
Google confidently explained that JPO's old systems supposedly:
- removed long-vowel marks
- expanded small Katakana characters such as ャ, ュ and ョ into full-sized characters
- did all this because some ancient Japanese mainframe couldn't handle modern Katakana properly
Fantastic story.
There was only one small problem.
JPO's own example is:
Look carefully at:
That ェ is a small Katakana character.
So the "JPO doesn't use small kana" explanation falls apart immediately.
This is why "an AI told me so" is not quite the same thing as documentation.
The long-vowel behaviour is real.
The sweeping explanation about eliminating small kana isn't.
As for exactly why JPO settled on this particular representation and how far back the convention goes, I'm still digging.
If somebody reading this has an old JPO specification, mainframe manual, trademark-search standard, or has spent the last 30 years maintaining Japanese IP databases for reasons I probably don't want to know about, please get in touch.
Seriously.
There are at least four of us now.
So, WTF is with JPO spelling?
The short version:
JPO needs a standardized pronunciation representation because trademark comparison depends heavily on sound, while Japanese trademarks can be written using several scripts and can have readings that aren't obvious from the written mark.
So JPO provides a Katakana pronunciation field called 称呼.
But that field isn't always written the way you'd normally write the same word in everyday Katakana.
In particular, long vowels can be represented explicitly with vowel kana:
which makes perfect sense as normalized phonetic/search data and looks absolutely deranged if you encounter it expecting ordinary Japanese spelling.
Which is exactly what happened to me 😁
So, I hope you've learned something new today. If not about trademarks, then at least about Japanese. Maybe you'll even start a Japanese Duolingo course, or go completely off the rails and install five Anki decks.
Either way, you've been warned 😁