A voice cloned from an English sample can speak Japanese. The language is a setting on the generation, not a property of the sample - so one clone covers all ten.
Language, voice and engine all sit in the same box as the script. Change the language, generate again, and the same voice reads in the new one - which is the point of cloning once rather than recording a speaker ten times.
Dictation works the other way round: it transcribes what you recorded, and the optional cleanup that drops fillers and restores punctuation stays in the language you actually spoke rather than translating it.