Whisper is an open-source automatic speech recognition system trained on 680,000 hours of multilingual and multitask supervised data collected from the web. It is designed to be robust to accents, background noise and technical language, and can transcribe and translate speech in multiple languages into English. It is a simple end-to-end approach, implemented as an encoder-decoder Transformer. It is also capable of performing language identification and phrase-level timestamps. It is designed to be easy to use and have high accuracy, allowing developers to add voice interfaces to more applications.

openai.comGitHub
Upvotes
128
Pricing
GitHub
Category
Speech-To-Text
In the database since
Dec 2022
Overview
About Whisper (OpenAI)
Newsletter
Tools like this change fast
New Speech-To-Text tools launch every week — the newsletter keeps you ahead of what's worth trying.
GitHub Repository
Note: This is a GitHub repository, meaning that it is code that someone created and made available for others to use. It typically requires some technical knowledge to set up and run.
AI news twice a week
Join 250,000+ readers getting the most important AI news and coolest tools every Wednesday and Friday.