Licensed to be used in conjunction with basebox, only.
// documents
Using speech-to-text
What is speech-to-text in basebox?
With speech-to-text in basebox, you can automatically convert spoken content into text. This saves time typing and lets you work on the go or in situations where writing is impractical.
Two options are available:
- Live recording via your microphone
- Audio file upload for recordings you already have
Method 1: Live speech input via microphone
When it's useful: For spontaneous input, notes, or whenever you'd rather speak than type.
How to use direct speech input:
- Find the microphone icon – Click the microphone icon next to the send button in the chat
- Grant browser permission – Allow basebox to access your microphone (appears the first time)
- Select a language – Choose between German or English
- Start recording – Recording begins automatically after you select a language
- Speak clearly – Speak clearly and at a normal pace
- End recording – Recording stops automatically after about 1 minute, or you can stop it manually
- Check the text – The spoken content is automatically inserted as text in the chat
What happens next: You can use the transcribed text just like normal text input – edit it, add to it, or send it directly.
Understanding browser permissions:
The first time, your browser will ask for microphone access:
- Chrome/Edge: Popup in the top left with "Allow" or "Block"
- Firefox: Notification in the address bar
- Safari: Permission in the browser settings
Important: Live recording will not work without microphone permission.
Method 2: Upload and transcribe an audio file
When it's useful: For meetings, interviews, talks, or other content you've already recorded.
How to transcribe audio files:
- Prepare the file – Make sure your audio file is in a supported format
- Start the upload – Drag and drop the file into the chat area
- Wait for processing – basebox analyzes the file automatically (can take a few minutes depending on length)
- Get the transcription – The text appears in the chat
- Work with the text – Use the transcribed text for further analysis or editing
Supported audio formats:
- WAV – Uncompressed quality (best results)
- MP3 – Compressed, widely used
- FLAC – Lossless compression
- OGG – Open-source format
Maximum file size: 10 MB per file
Tips for optimal recognition quality
For live recordings:
Optimize your environment:
- Quiet environment – Minimize background noise
- Good microphone – Use a headset or external microphone if possible
- Stable distance – Keep about 20–30 cm from the microphone
Adjust how you speak:
- Articulate clearly – Speak clearly and not too fast
- Normal volume – Don't whisper, don't shout
- Take pauses – Short pauses between sentences help recognition
For audio files:
Recording quality:
- High audio quality – At least 16 kHz sample rate
- Mono or stereo – Both are supported
- Low compression – WAV or FLAC for best results
Optimizing content:
- One speaker – Works best with a single person
- Clear speech – Dialects can affect accuracy
- Short segments – Split up very long recordings
Supported languages
Currently available:
- German – Optimized for German language and terminology
- English – For English-language content
Language selection: The language is detected automatically or can be selected manually.
Common problems and solutions
Microphone doesn't work:
Problem: No recording is possible Solutions:
- Check browser permissions and grant them again
- Enable the microphone in your system settings
- Try a different browser
- Reload the page and try again
Poor recognition quality:
Problem: Text is inaccurate or incomplete Solutions:
- Speak more clearly and slowly
- Reduce background noise
- Adjust the distance to the microphone
- Select a different language
Audio file isn't processed:
Problem: Upload doesn't work Solutions:
- Check the file format (WAV, MP3, FLAC, OGG)
- Keep the file size under 10 MB
- Convert the file to a different format
- Clear the browser cache
Practical use cases
For meetings and discussions:
- Live notes during video conferences
- Creating meeting minutes from recordings
- Quickly dictating action items
For content creation:
- Dictating blog articles instead of typing
- Collecting ideas on the go via voice memo
- Transcribing interviews for further editing
For documentation:
- Documenting work steps verbally
- Speaking project reports instead of writing them
- Following up on customer conversations
After transcription: making further use of the text
What you can do with the transcribed text:
- Edit it directly – Correct errors and add to it
- Have it summarized – basebox can analyze the text
- Structure it – Organize it into bullet points or chapters
- Translate it – Convert it into other languages
- Process it further – Use it as a basis for other documents
Important note
Recognition quality depends heavily on audio quality and how you speak. Always check the transcribed text for accuracy, especially for important documents.
Tip
Start with short test recordings to get a feel for the optimal settings.
Problems with speech recognition or questions about audio quality? Contact support