MP3 ਨੂੰ ਟੈਕਸਟ ਵਿਚ transcribe ਕਰੋ।ਸਪੀਕਰ ਲੇਬਲ, 100+ ਭਾਸ਼ਾਵਾਂ।

ਕਿਸੇ ਵੀ bitrate 'ਤੇ 64 ਤੋਂ 320 kbps ਤੱਕ MP3 ਫ਼ਾਈਲ ਡ੍ਰੌਪ ਕਰੋ। 99 ਭਾਸ਼ਾਵਾਂ ਵਿਚ timestamp ਕੀਤਾ, speaker-labeled transcript ਪ੍ਰਾਪਤ ਕਰੋ — ਕੋਈ format conversion, ਕੋਈ re-encoding, ਕੋਈ queue 'ਤੇ ਇੰਤਜ਼ਾਰ ਨਹੀਂ।

ਆਪਣਾ ਆਡੀਓ ਜਾਂ ਵੀਡੀਓ ਛੱਡੋ

MP3 · WAV · M4A · MP4 · MOV · MKV · OGG · OPUS · FLAC · WEBM — up to 100 MB anonymously

Paste a link, we’ll fetch the audio

YouTube · TikTok · Vimeo · Twitter · SoundCloud · Spotify · 50+ more

ਸਿੱਧਾ ਆਪਣੇ ਬ੍ਰਾਊਜ਼ਰ ਤੋਂ ਰਿਕਾਰਡ ਕਰੋ

ਸਾਈਨ ਅੱਪ 30 ਸਕਿੰਟ ਲੈਂਦਾ ਹੈ — ਉਸ ਤੋਂ ਬਾਅਦ ਡੈਸ਼ਬੋਰਡ ਵਿੱਚ ਸਿੱਧਾ ਰਿਕਾਰਡਿੰਗ ਖੁੱਲ੍ਹ ਜਾਂਦੀ ਹੈ।

No card required~90s per 60-min fileSRT · VTT · DOCX · TXTਫ਼ਾਈਲਾਂ 24 ਘੰਟਿਆਂ ਵਿੱਚ ਆਪੇ-ਡਿਲੀਟ

↓ ਵੇਖੋ ਕਿ ਕੀ ਨਿਕਲਦਾ ਹੈ

MP3 ਅੰਦਰ। Diarized transcript ਬਾਹਰ।

ਅਸੀਂ MP3 frame headers ਸਿੱਧੇ ਪੜ੍ਹਦੇ ਹਾਂ — VBR, CBR, joint-stereo, ਕੋਈ ਵੀ encoder (LAME, Fraunhofer, FFmpeg)। ਜੇ ਫ਼ਾਈਲ ਅਲੱਗ-ਅਲੱਗ ਚੈਨਲਾਂ 'ਤੇ ਸਪੀਕਰ ਦੇ ਨਾਲ true stereo ਹੈ, ਅਸੀਂ voices ਨੂੰ ਵੰਡਣ ਲਈ ਇਸ ਦੀ ਵਰਤੋਂ ਕਰਦੇ ਹਾਂ। Mono mix-down acoustic diarization 'ਤੇ ਵਾਪਸ ਜਾਂਦਾ ਹੈ।

interview-tape-04.mp3REC 192 kbps · stereo · 38:42
auto-detected en-GB44.1 kHz · LAME 3.100
~90s
Transcript · streaming95% ਸ਼ੁੱਧਤਾ
S1

ਤੁਸੀਂ ਪਹਿਲਾਂ ਕਦੋਂ ਸਮਝਿਆ ਕਿ ਆਰਕਾਈਵ ਅਧੂਰੀ ਸੀ?

S2

ਸ਼ਾਇਦ 2019 ਦੇ ਆਸ-ਪਾਸ, ਜਦੋਂ ਮੇਂ reel-to-reels ਨੂੰ digitise ਕਰਨਾ ਸ਼ੁਰੂ ਕੀਤਾ।

S1

ਅਤੇ ਗੁਆਪਤ ਟੇਪਾਂ — ਕੀ ਉਹ ਕਹੀਂ ਵੀ catalogue ਕੀਤੀਆਂ ਗਈਆਂ ਸਨ?

S2

'78 ਤੋਂ ਵੀ ਇੱਕ paper index ਹੈ, ਪਰ ਅੱਧੀ ਪਾਣੀ ਦੀ ਨੁਕਸਾਨ ਹਾ।

192 kbps stereo 'ਤੇ 95%SRT · DOCX · TXT · JSON · VTT

↓ This is the dashboard

This is what loads when the job finishes.

Same layout as the real dashboard — Summary, full Transcript, Speakers tab, Exports. Key points and action items extracted automatically. Auto-tags on every job.

Try it on your own file — it's free

ਤਿੰਨ ਅਸਲੀ ਵਿਕਲਪ · ਸ਼ੇਮ ਅਨੁਸਾਰ

ਮੁਫ਼ਤ local Whisper। Otter ਜਾਂ Sonix। ਜਾਂ ਅਸੀਂ।

ਤੁਸੀਂ ਮੁਫ਼ਤ ਵਿੱਚ ਆਪਣੇ laptop 'ਤੇ Whisper ਨੂੰ ਚਲਾ ਸਕਦੇ ਹੋ ਜੇ ਤੁਸੀਂ ਤਕਨੀਕੀ ਹੋ। Otter ਅਤੇ Sonix subscription dashboards ਵਿਚ MP3 uploads ਸਵੀਕਾਰ ਕਰਦੇ ਹਨ। ਅਸੀਂ ਫ਼ਾਈਲ ਲੈਂਦੇ ਹਾਂ, transcript ਵਾਪਸ ਕਰਦੇ ਹਾਂ, ਅਤੇ ਤੁਹਾਨੂੰ UI ਵਿਚ ਰਹਿਣ ਲਈ ਕਹਿ ਨਹੀਂ।

Option 01

Whisper local / open source

ਮੁਫ਼ਤ ਜੇ ਤੁਕੱਲ ਕੋਲ GPU ਅਤੇ ਇੱਕ afternoon ਹੈ। Speaker diarization box ਤੋਂ ਬਾਹਰ ਨਹੀਂ।

SetupPython + CUDA + 10 GB models
Speaker diarizationInclude ਨਹੀਂ ਕੀਤਾ (pyannote add-on)
Speed · 1 hr MP3consumer GPU 'ਤੇ 5–40 ਮਿੰਟ
Languages99, ਪਰ tiny model 80% ਨਿਵੇਂ ਤੇ ਆਉਂਦਾ ਹੈ
ExportTXT / SRT / VTT / JSON
Costਮੁਫ਼ਤ + ਤੁਹਾਡੀ ਬਿਜਲੀ
Best forEngineers ਜਿਨ੍ਹਾਂ ਕੋਲ ਪਹਿਲਾਂ ਤੋਂ GPU ਹੈ, speaker ਲੇਬਲ ਦੀ ਲੋੜ ਨਹੀਂ, ਅਤੇ ਪੂਰੀ local privacy ਚਾਹਿੰਦੇ ਹਨ।
Option 02

Transcription.Solutions

MP3 ਡ੍ਰੌਪ ਕਰੋ। Speaker-labeled ਲਿਖਤ ਵਾਪਸ ਪ੍ਰਾਪਤ ਕਰੋ roughly real-time × 0.025 ਵਿਚ।

SetupDrag-and-drop, ਅਜ਼ਮਾਉਣ ਲਈ account ਦੀ ਲੋੜ ਨਹੀਂ
Speaker diarizationBuilt in (Pro & Business ਪਲੈਨ)
Speed · 1 hr MP3~90 ਸੈਕਿੰਡ
Languages99, auto-detected
ExportSRT · VTT · DOCX · TXT · JSON
Cost · per min$0.03
Best forਕੋਈ ਵੀ MP3 ਦੇ ਨਾਲ — journalist ਟੇਪ, podcast export, voice memo, archival dub — ਜੋ ਸਿਰਫ਼ ਸ਼ੁੱਧ ਲਿਖਤ output ਚਾਹਿੰਦਾ ਹੈ।
Option 03

Otter / Sonix

Polished dashboard, monthly minutes cap, English-tuned। File upload ਪਾਸੇ ਦੀ feature ਲੱਗਦਾ ਹੈ।

SetupAccount + paid ਪਲੈਨ
Speaker diarizationAcoustic, EN-leaning
Speed · 1 hr MP3queue ਵਿਚ 5–10 ਮਿੰਟ
LanguagesOtter EN-only; Sonix ~40
ExportLocked ਪੇ paid tiers ਪਿੱਛੇ
Cost$17+/mo ਜਾਂ $10+/hr (Sonix)
Best forTeams ਜੋ clean API-style file→text flow ਨਾਲ ਖ਼ਾਲੀ transcript editor ਅਤੇ collaboration UI ਸਾਂਝਾ ਪਸੰਦ ਕਰਦੇ ਹਨ।

May 2026 ਤੋਂ ਸਹੀ pricing ਅਤੇ ਫ਼ੀਚਿਅਰ availability। Whisper ਪ੍ਰਦਰਸ਼ਨ model ਸਾਈਜ਼ ਅਤੇ hardware ਅਨੁਸਾਰ ਵੱਖਵੱਖ।

MP3 ਨੂੰ Specific

ਤਿੰਨ ਚੀਜ਼ਾਂ ਜੋ 'ਤੇ ਲੋਕਾਂ ਨੂੰ ਕਾਟਦੀ ਹਨ। generic transcription tools.

MP3 ਇੱਕ format ਹੈ, ਰਿਕਾਰਡਿੰਗ style ਨਹੀਂ — ਜੋ ਮਤਲਬ ਹੈ failure modes encoder ਤੋਂ ਆਉਂਦੇ ਹਨ, ਬੋਲੀ ਤੋਂ ਨਹੀਂ।

ਕੀ ਗਲ਼ਤ ਹੋ ਸਕਦਾ ਹੈ

  1. 1VBR headers mis-parsed। ਕੁਝ tools variable-bitrate MP3s ਨੂੰ fixed-rate ਦੇ ਤੌਰ 'ਤੇ ਪੜ੍ਹਦੇ ਹਨ ਅਤੇ duration ਦੀ غ ਮਿਸ ਕਰਦੇ ਹਨ — timestamps ਘੰਟੇ ਲੰਬੀ ਫ਼ਾਈਲ 'ਤੇ ਮਿੰਟ ਨਾਲ drift।
  2. 2Joint-stereo upload preprocessing ਦੌਰਾਨ mono ਵਿੱਚ flatten। ਤੁਹਾਨੂੰ per-speaker channel separation ਮਿਲ ਜਾਂਦਾ ਹੈ ਜੋ ਅਸਲ ਵਿਚ ਫ਼ਾਈਲ ਵਿਚ ਸੀ।
  3. 3Embedded ID3 album art ਕੁਝ uploaders ਨੂੰ ਤਕਲੀਫ ਵਿਚ ਪੈਲਦਾ ਹੈ — ਉਹ ਫ਼ਾਈਲ ਨੂੰ 'not pure audio' ਦੇ ਰੂਪ ਵਿਚ ਰੱਦ ਕਰਦੇ ਹਨ ਜਾਂ ਇਸ ਨੂੰ strip ਕਰਦੇ ਹਨ ਅਤੇ re-encode ਕਰਦੇ ਹਨ, ਗੁਣਵੱਤਾ ਨੂੰ ਹੋਰ drop ਕਰਦੇ ਹਨ।

ਅਸੀਂ ਕੀ ਕਰਦੇ ਹਾਂ

  1. 1ਅਸੀਂ Xing/LAME header ਵਰਤਦੇ ਹਾਂ ਜਦੋਂ ਮੌਜਮਪੂਦ ਹੋ ਅਤੇ frame-count fallback ਜਦੋਂ ਨਹੀਂ। VBR timestamps ਮਲਟੀ-ਘੰਟੇ ਦੀ ਫ਼ਾਈਲ 'ਤੇ ±0.1 s ਤੱਕ ਸਹੀ ਰਹਿੰਦੀ ਹੈ।
  2. 2Joint-stereo ਅਤੇ true-stereo MP3s ਨੂੰ diarization ਤੋਂ ਪਹਿਲਾਂ L/R PCM ਵਿਚ decode। ਜੇ ਤੁਹਾਡੇ speakers ਨੂੰ panned ਸੀ, ਅਸੀਂ ਉਨ੍ਹਾਂ ਨੂੰ split ਰੱਖਦੇ ਹਾਂ।
  3. 3ID3v1, ID3v2, APE tags, embedded art — ਸਾਰਾ untouched pass ਕੀਤਾ। ਅਸੀਂ ਤੁਹਾਡੇ MP3 ਨੂੰ ਕਦੇ re-encode ਨਹੀਂ ਕਰਦੇ।

MP3 uploads ਲਈ ਸਿਫਾਰਸ਼ ਕੀਤੀ job settings

Defaults ਜੋ ~80% MP3 ਫ਼ਾਈਲਾਂ ਨੂੰ fit ਕਰਦੀਆਂ ਹਨ। Form ਤੋਂ ਪ੍ਰਤੀ-job override।

Decoder
Frame-accurate, no re-encode
Diarization
Channel split ਜੇ stereo, ਨਹੀਂ tο acoustic
Speaker model
Auto · 1-12 speakers
Language
Auto-detect ਪਹਿਲੇ 30 s ਤੋਂ
Filler words
Removed (toggle ਨੂੰ ਰੱਖਣ ਲਈ)
Export bundle
DOCX + SRT + timestamped TXT

Accuracy · real-world numbers

95%+ 192 kbps stereo 'ਤੇ। Usable down to 64 kbps mono।

MP3 ਸ਼ੁੱਧਤਾ ਨੂੰ bounds ਜੋ encoder ਰੱਖਿਆ, ਸਾਨੂੰ ਨਹੀਂ। Perceptual compression ~96 kbps ਤੋਂ ਉੱਪਰ speech intelligibility ਨੂੰ ਬਹੁਤ ਚੰਗੀ ਤਰ੍ਹਾਂ ਸੁਰੱਖਿਅਤ ਰੱਖਦਾ ਹੈ; 64 kbps ਤੋਂ ਘੱਟ, sibilants ਅਤੇ consonants dissolve ਹੋਣ ਲਗਦੇ ਹਨ। ਹੇਠਲੀ ਸੰਖਿਆਵਾਂ production ਵਿਚ ਅਸਲੀ customer MP3s ਤੋਂ ਹਨ।

96%
320 kbps stereo, studio ਜ਼ਾਂ

Near-lossless ਬੋਲੀ ਲਈ। Podcast masters, dictation app exports, professional interview rigs। Diarization ਸਾਫ਼ ਹੈ ਜੇ speakers ਅਲੱਗ channels 'ਤੇ ਹਨ।

95%
192 kbps stereo, 2-3 speakers

ਸਭ ਤੋਂ ਆਮ bitrate spoken-word MP3s ਲਈ। Zoom exports, Riverside downloads, voice recorders default। Compression artifacts recognizer ਲਈ inaudible।

91%
128 kbps mono, conversational

ਜ਼ਿ voice memo defaults ਆ ਜ਼ਿਭਪਟਭ phones 'ਤੇ। Acoustic diarization 2-4 speakers ਨੂੰ ਸੰਭਾਲਦੀ ਹੈ। ਨੰਬਾਂ ਅਤੇ proper nouns ਨੂੰ ਕਦੇ ਇੱਕ ਨਜ਼ਰ ਦੀ ਲੋੜ।

84%
64 kbps mono, archival / phone-dump

ਪੁਰਾਣੀ answering-machine rips, lecture archives, narrow-band ਸੋਤ। High-frequency consonants (f/s/sh) blur। ਫਿਰ ਵੀ legible — proofread ਦੀ ਯੋਜਨਾ ਬਣਾਓ।

Common ਸਵਾਲ

8 ਚੀਜ਼ਾਂ ਲੋਕ ਬਾਰੇ ਪੁੱਛਦੇ ਹਨ। MP3 transcription.

01MP3 bitrate ਘੱਟ ਤੋਂ ਘੱਟ ਕੀ ਹੋ ਸਕਦਾ ਹੈ ਜੋ ਫਿਰ ਵੀ usable transcript ਦਿੰਦਾ ਹੈ?+
64 kbps ਅਸਲੀ ਫ਼ਰਸ਼ ਹੈ। ਉਸ ਤੋਂ ਘੱਟ, sibilants (s, sh, f) noise ਵਿਚ compress ਹੋ ਜਾਂਦੇ ਹਨ ਅਤੇ word error rate 20% ਤੋਂ ਉੱਪਰ ਚਢ jਾਂਦਾ ਹੈ। ਜੇ ਤੁਸੀਂ ਨਿਆ ਰਿਕਾਰਡ ਕਰ ਰਹੇ ਹੋ, 128 kbps mono ਜਾਂ 192 kbps stereo ਲੀਲਾ ਕਰੋ — ਕੋਈ ਵੀ ਇਸ ਤੋਂ ਵੱਧ ਬੋਲੀ ਲਈ overkill।
02ਕੀ ਮੈਨੂੰ ਪਹਿਲਾਂ MP3 ਨੂੰ WAV ਵਿਚ ਤਬਦੀਲ ਕਰਨਾ ਚਾਹੀਦਾ ਹੈ?+
ਨਹੀਂ। MP3 → WAV re-encoding ਕੋਈ ਵੀ ਸ਼ੁੱਧਤਾ ਜੋੜਦਾ ਹੈ ਕਿਉਂਕਿ ਡਾਟਾ ਜੋ encoder ਨੇ discard ਕੀਤਾ ਹੈ ਉਹ ਚਲਾ ਗਿਆ ਹੈ। MP3 ਸਿੱਧਾ upload ਕਰੋ। ਅਸੀਂ frames ਨੂੰ memory ਵਿਚ decode ਕਰਦੇ ਹਾਂ ਅਤੇ PCM ਨੂੰ recognizer ਨੂੰ feed ਕਰਦੇ ਹਾਂ।
03ਕੀ stereo MP3 mono ਨਾਲੋਂ ਮੈਨੂੰ ਬਿਹਤਰ speaker ਲੇਬਲ ਦੇਵੇਗਾ?+
ਸਿਰਫ਼ ਜੇ speakers ਨੂੰ ਅਸਲ ਵਿਚ ਅਲੱਗ channels 'ਤੇ ਰਿਕਾਰਡ ਕੀਤਾ ਗਿਆ — ਜ਼ਿ stereo MP3s ਦੋਵਾਂ ਪਾਸਿਆਂ 'ਤੇ ਇਕੋ ਆਡੀਓ ਹੈ ('dual mono') ਅਤੇ ਕੁਝ ਵੀ gain ਨਹੀਂ ਕਰਦੇ। True channel-split (ਮਸਲਾ Riverside exports, two-mic ਫ਼ੀਲਡ rigs) ਸਾਨੂੰ acoustic diarization ਨੂੰ skip ਕਰਨ ਦਿੰਦਾ ਹੈ ਅਤੇ speakers ਨੂੰ near-perfectly label।
04ਤੁਸੀਂ ਸਭ ਤੋਂ ਵੱਡੀ MP3 ਫ਼ਾਈਲ ਕੀ accept ਕਰਦੇ ਹੋ?+
5 GB per upload, ਜੋ roughly 192 kbps 'ਤੇ 60 ਘੰਟੇ ਜਾਂ 128 kbps 'ਤੇ 90 ਘੰਟੇ ਹੈ। ਜੇ ਤੁਹਾਡੀ ਫ਼ਾਈ��� ਵੱਡੀ ਹੈ ਅਸੀਂ chunked upload ਦਿਖਾਵਾਂਗੇ — ਤੁਹਾਨੂੰ ਇਸ ਨੂੰ ਖ਼ੁਦ split ਕਰਨ ਦੀ ਲੋੜ ਨਹੀਂ।
0560-ਮਿੰਟ MP3 transcribe ਕਰਨ ਵਿਚ ਕਿੰਨਾ ਵਕਤ ਲਗਦਾ ਹੈ?+
ਆਮ 'ਤੇ 90 ਸੈਕਿੰਡ upload-complete ਤੋਂ transcript-ready ਤੱਕ, bitrate ਕਿਮ ਧਿਆਨ ਦਿਏ। MP3 frames ਨੂੰ decode ਕਰਨਾ ਤੇਜ਼ ਹੈ; ਸਮੇਂ recognizer ਵਿਚ ਹੈ। Diarization multi-speaker ਫ਼ਾਈਲਾਂ 'ਤੇ 5-10 ਸੈਕਿੰਡ jੋੜਦੀ ਹੈ।
06ਮੇਰੀ MP3 ਦੇ background ਵਿਚ ਸੰਗੀਤ ਹੈ — ਕੀ transcript ਬਰਬਾਦ ਹੋਵੇਗਾ?+
ਬੋਲੀ ਤੋਂ ਘੱਟ ਬਿਸਤਰਾ ਸੰਗੀਤ ਠੀਕ ਹੈ। ਜ਼ੋਰ ਵਾਲਾ ਸੰਗੀਤ ਜੋ ਆਵਾਜ ਨਾਲ ਮੁਕਾਬਲਾ ਕਰਦਾ ਹੈ (intro stings, scoring ਅੰਤਰਵਿਅ)) ਕਦੇ ਓਵਰਲੈਪਿੰਗ syllables 'ਤੇ misrecognitions ਟਰਿਗਰ ਕਰਦਾ ਹੈ। Job form 'ਤੇ music suppression toggle ਕਰੋ pre-filter ਲਈ��
07ਕੀ ਤੁਸੀਂ phone voicemail ਜਾਂ answering machines ਤੋਂ ripped MP3s ਲਾ ਸਕਦੇ ਹੋ?+
ਹਾਂ, ਭਾਵ ਇਹਨਾਂ ਨੂੰ acronym 8 kHz narrow-band re-encoded ਦੇ ਤੌਰ 'ਤੇ MP3 ਵਿਚ — ਆਡੀਓ quality ceiling original PSTN capture ਦਾ ਤੈ ਬਲਾਕ ਹੈ, MP3 wrapper ਦਾ ਨਹੀਂ। Expect 78-85% accuracy ਇਸ ਕਿਸਮ ਦੋ ਮਾਪ 'ਤੇ, ਜੋ ਸਮਾਨ ਹੈ ਜੋ ਅਸੀਂ underlying call 'ਤੇ ਪਾਏ।
08ਕੀ ਤੁਸੀਂ transcript ਹਨ ਤੋਂ ਬਾਅਦ ਮੇਰੀ MP3 ਰੱਖਦੇ ਹੋ?+
ਫ਼ਾਈਲਾਂ default ਮੁਤਾਬਕ 30 ਦਿਨਾਂ ਬਾਹਰ ਡਿਲੀਟ ਹੋ ਜਾਂਦੀਆਂ ਹਨ, ਜਾਂ dashboard ਦੁਆਰੇ ਤੱਤ 'ਤੇ ਬਿਆਨ। Transcript ਤੁਹਾਡੇ account ਵਿਚ ਰਹਿੰਦਾ ਹੈ ਤੁਸੀਂ ਇਸ ਨੂੰ ਡਿਲੀਟ ਨਹੀਂ ਕਰ। ਅਸੀਂ ਕਿਸੇ ਮਾਡਲ ਨੂੰ ਲੰਮੇ ਪਲੇ ਚਾਲਾ ਕਰਨ ਲਈ customer audio ਵਰਤ ਨਹੀਂ ਕਰਦੇ — ਕਦੇ ਨਹੀਂ।

ਆਪਣੀ MP3 ਡ੍ਰੌਪ ਕਰੋ। 90 ਸੈਕਿੰਡ ਵਿਚ ਟੈਕਸਟ ਵਾਪਸ ਪ੍ਰਾਪਤ ਕਰੋ।

ਹਰ ਮਹੀਨਾ 30 ਮੁਫ਼ਤ ਮਿੰਟ। ਕਾਰਡ ਦੀ ਲੋੜ ਨਹੀਂ। Speaker ਲੇਬਲ, 99 ਭਾਸ਼ਾਵਾਂ, ਹਰ export format include।

ਮੁਫ਼ਤ ਸ਼ੁਰੂ ਕਰੋ