kotoba-tech
/

kotoba-whisper-bilingual-v1.0

@@ -47,7 +47,7 @@ We compare our kotoba-whisper-bilingual with OpenAI whisper models, kotoba-whisp
 OpenAI whisper is not trained for English to Japanese speech-to-text translation, and other models are specific to the Task (eg. kotoba-whisper is Japanese ASR and
 distil whisper is English ASR only).
-### Speech2Text Translation (Japanese->English)
 | model                                                                                                                                                                                                     |   [CoVoST2 (Ja->En)](https://huggingface.co/datasets/japanese-asr/ja2en.s2t_translation)|   [Fleurs (Ja->En)](https://huggingface.co/datasets/japanese-asr/ja2en.s2t_translation) |
 |:----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------------------------------------------------------------------------------------:|------------------------------------------------------------------------------------------------------:|
@@ -65,7 +65,7 @@ distil whisper is English ASR only).
 | [openai/whisper-tiny](https://huggingface.co/openai/whisper-tiny)                                                                                                                                         |                                                                                                  377.2 |                                                                                                 474   |
-### Speech2Text Translation (English->Japanese)
 | model                                                                                                                                                                                                     |   [CoVoST2 (En->Ja)](https://huggingface.co/datasets/japanese-asr/en2ja.s2t_translation)|   [Fleurs (En->JA)](https://huggingface.co/datasets/japanese-asr/en2ja.s2t_translation) |
 |:----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------------------------------------------------------------------------------------:|------------------------------------------------------------------------------------------------------:|
@@ -83,7 +83,7 @@ distil whisper is English ASR only).
 | [openai/whisper-tiny](https://huggingface.co/openai/whisper-tiny)                                                                                                                                         |                                                                                                  185.2 |                                                                                                 200.5 |
-### ASR (Japanese)
 | model                                                                                                                                             |   [CommonVoice 8 (Japanese test set)](https://huggingface.co/datasets/japanese-asr/ja_asr.common_voice_8_0) |   [JSUT Basic 5000](https://huggingface.co/datasets/japanese-asr/ja_asr.jsut_basic5000) |   [ReazonSpeech (held out test set)](https://huggingface.co/datasets/japanese-asr/ja_asr.reazonspeech_test) |
 |:--------------------------------------------------------------------------------------------------------------------------------------------------|------------------------------------------------------------------------------------------------------------:|----------------------------------------------------------------------------------------:|------------------------------------------------------------------------------------------------------------:|
@@ -101,7 +101,7 @@ distil whisper is English ASR only).
-### ASR (English)
 | model                                                                                                           |   [ESB](https://huggingface.co/datasets/japanese-asr/en_asr.esb_eval) (ami) |   [ESB](https://huggingface.co/datasets/japanese-asr/en_asr.esb_eval) (earnings22) |   [ESB](https://huggingface.co/datasets/japanese-asr/en_asr.esb_eval) (librispeech) |   [ESB](https://huggingface.co/datasets/japanese-asr/en_asr.esb_eval) (tedlium) |   [ESB](https://huggingface.co/datasets/japanese-asr/en_asr.esb_eval) (voxpopuli) |
 |:----------------------------------------------------------------------------------------------------------------|----------------------------------------------------------------------------:|-----------------------------------------------------------------------------------:|------------------------------------------------------------------------------------:|--------------------------------------------------------------------------------:|----------------------------------------------------------------------------------:|

 OpenAI whisper is not trained for English to Japanese speech-to-text translation, and other models are specific to the Task (eg. kotoba-whisper is Japanese ASR and
 distil whisper is English ASR only).
+### Speech2Text Translation (Japanese->English): WER
 | model                                                                                                                                                                                                     |   [CoVoST2 (Ja->En)](https://huggingface.co/datasets/japanese-asr/ja2en.s2t_translation)|   [Fleurs (Ja->En)](https://huggingface.co/datasets/japanese-asr/ja2en.s2t_translation) |
 |:----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------------------------------------------------------------------------------------:|------------------------------------------------------------------------------------------------------:|
 | [openai/whisper-tiny](https://huggingface.co/openai/whisper-tiny)                                                                                                                                         |                                                                                                  377.2 |                                                                                                 474   |
+### Speech2Text Translation (English->Japanese): CER
 | model                                                                                                                                                                                                     |   [CoVoST2 (En->Ja)](https://huggingface.co/datasets/japanese-asr/en2ja.s2t_translation)|   [Fleurs (En->JA)](https://huggingface.co/datasets/japanese-asr/en2ja.s2t_translation) |
 |:----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------------------------------------------------------------------------------------:|------------------------------------------------------------------------------------------------------:|
 | [openai/whisper-tiny](https://huggingface.co/openai/whisper-tiny)                                                                                                                                         |                                                                                                  185.2 |                                                                                                 200.5 |
+### ASR (Japanese): CER
 | model                                                                                                                                             |   [CommonVoice 8 (Japanese test set)](https://huggingface.co/datasets/japanese-asr/ja_asr.common_voice_8_0) |   [JSUT Basic 5000](https://huggingface.co/datasets/japanese-asr/ja_asr.jsut_basic5000) |   [ReazonSpeech (held out test set)](https://huggingface.co/datasets/japanese-asr/ja_asr.reazonspeech_test) |
 |:--------------------------------------------------------------------------------------------------------------------------------------------------|------------------------------------------------------------------------------------------------------------:|----------------------------------------------------------------------------------------:|------------------------------------------------------------------------------------------------------------:|
+### ASR (English): WER
 | model                                                                                                           |   [ESB](https://huggingface.co/datasets/japanese-asr/en_asr.esb_eval) (ami) |   [ESB](https://huggingface.co/datasets/japanese-asr/en_asr.esb_eval) (earnings22) |   [ESB](https://huggingface.co/datasets/japanese-asr/en_asr.esb_eval) (librispeech) |   [ESB](https://huggingface.co/datasets/japanese-asr/en_asr.esb_eval) (tedlium) |   [ESB](https://huggingface.co/datasets/japanese-asr/en_asr.esb_eval) (voxpopuli) |
 |:----------------------------------------------------------------------------------------------------------------|----------------------------------------------------------------------------:|-----------------------------------------------------------------------------------:|------------------------------------------------------------------------------------:|--------------------------------------------------------------------------------:|----------------------------------------------------------------------------------:|