Freelance › Projects › Software development › Speech-to-Text Transcription Service with API (Russian & English)
Speech-to-Text Transcription Service with API (Russian & English)

Employer
Dmytro
Project parameters
Type of cooperationOne-time project
SectionSoftware development
Prepaymentwithout prepayment
Payment methodsCash, Bank transfer
Acceptance of requestsfrom today, 21:46 until Aug 30, 2026
Project description
We are a company that processes a large volume of recorded audio every week — customer support calls, remote interviews, and internal team meetings — and we need a reliable speech-to-text transcription service built for us. The core task is to take audio files in common formats and automatically convert them into accurate, readable text. Each transcript must clearly separate who said what by labelling individual speakers, and every segment should carry precise timestamps so that any line in the text can be traced back to the exact moment in the original recording. The system has to handle both Russian and English audio, including mixed recordings where participants switch between the two languages during the same conversation.
Beyond raw accuracy, we care about how the service fits into our existing workflow. The transcription engine must be exposed through a clean, well-documented API so our internal applications can submit audio, poll for processing status, and retrieve finished transcripts programmatically without any manual steps. We expect sensible handling of long recordings, background noise, and overlapping speech, along with structured output that we can store and search later. The ideal contractor will build the backend in Python, wire up a proven recognition model, and deliver a solution that is stable under regular daily load rather than a one-off script. Please include examples of similar transcription or audio-processing work you have delivered, and describe how you would approach speaker separation and accuracy tuning for our language pair.
— API endpoints for uploading audio and fetching transcripts
— Speaker labels (diarization) and word- or segment-level timestamps
— Support for Russian, English, and mixed-language recordings
— Python backend with a documented, testable deployment
Beyond raw accuracy, we care about how the service fits into our existing workflow. The transcription engine must be exposed through a clean, well-documented API so our internal applications can submit audio, poll for processing status, and retrieve finished transcripts programmatically without any manual steps. We expect sensible handling of long recordings, background noise, and overlapping speech, along with structured output that we can store and search later. The ideal contractor will build the backend in Python, wire up a proven recognition model, and deliver a solution that is stable under regular daily load rather than a one-off script. Please include examples of similar transcription or audio-processing work you have delivered, and describe how you would approach speaker separation and accuracy tuning for our language pair.
— API endpoints for uploading audio and fetching transcripts
— Speaker labels (diarization) and word- or segment-level timestamps
— Support for Russian, English, and mixed-language recordings
— Python backend with a documented, testable deployment