Skip to content

// documents

Using speech-to-text

What is speech-to-text in basebox?

With speech-to-text in basebox, you can automatically convert spoken content into text. This saves time typing and lets you work on the go or in situations where writing is impractical.

Two options are available:

  • Live recording via your microphone
  • Audio file upload for recordings you already have

Method 1: Live speech input via microphone

When it's useful: For spontaneous input, notes, or whenever you'd rather speak than type.

How to use direct speech input:

  1. Find the microphone icon – Click the microphone icon next to the send button in the chat
  2. Grant browser permission – Allow basebox to access your microphone (appears the first time)
  3. Select a language – Choose between German or English
  4. Start recording – Recording begins automatically after you select a language
  5. Speak clearly – Speak clearly and at a normal pace
  6. End recording – Recording stops automatically after about 1 minute, or you can stop it manually
  7. Check the text – The spoken content is automatically inserted as text in the chat

What happens next: You can use the transcribed text just like normal text input – edit it, add to it, or send it directly.

Understanding browser permissions:

The first time, your browser will ask for microphone access:

  • Chrome/Edge: Popup in the top left with "Allow" or "Block"
  • Firefox: Notification in the address bar
  • Safari: Permission in the browser settings

Important: Live recording will not work without microphone permission.

Method 2: Upload and transcribe an audio file

When it's useful: For meetings, interviews, talks, or other content you've already recorded.

How to transcribe audio files:

  1. Prepare the file – Make sure your audio file is in a supported format
  2. Start the upload – Drag and drop the file into the chat area
  3. Wait for processing – basebox analyzes the file automatically (can take a few minutes depending on length)
  4. Get the transcription – The text appears in the chat
  5. Work with the text – Use the transcribed text for further analysis or editing

Supported audio formats:

  • WAV – Uncompressed quality (best results)
  • MP3 – Compressed, widely used
  • FLAC – Lossless compression
  • OGG – Open-source format

Maximum file size: 10 MB per file

Tips for optimal recognition quality

For live recordings:

Optimize your environment:

  • Quiet environment – Minimize background noise
  • Good microphone – Use a headset or external microphone if possible
  • Stable distance – Keep about 20–30 cm from the microphone

Adjust how you speak:

  • Articulate clearly – Speak clearly and not too fast
  • Normal volume – Don't whisper, don't shout
  • Take pauses – Short pauses between sentences help recognition

For audio files:

Recording quality:

  • High audio quality – At least 16 kHz sample rate
  • Mono or stereo – Both are supported
  • Low compression – WAV or FLAC for best results

Optimizing content:

  • One speaker – Works best with a single person
  • Clear speech – Dialects can affect accuracy
  • Short segments – Split up very long recordings

Supported languages

Currently available:

  • German – Optimized for German language and terminology
  • English – For English-language content

Language selection: The language is detected automatically or can be selected manually.

Common problems and solutions

Microphone doesn't work:

Problem: No recording is possible Solutions:

  • Check browser permissions and grant them again
  • Enable the microphone in your system settings
  • Try a different browser
  • Reload the page and try again

Poor recognition quality:

Problem: Text is inaccurate or incomplete Solutions:

  • Speak more clearly and slowly
  • Reduce background noise
  • Adjust the distance to the microphone
  • Select a different language

Audio file isn't processed:

Problem: Upload doesn't work Solutions:

  • Check the file format (WAV, MP3, FLAC, OGG)
  • Keep the file size under 10 MB
  • Convert the file to a different format
  • Clear the browser cache

Practical use cases

For meetings and discussions:

  • Live notes during video conferences
  • Creating meeting minutes from recordings
  • Quickly dictating action items

For content creation:

  • Dictating blog articles instead of typing
  • Collecting ideas on the go via voice memo
  • Transcribing interviews for further editing

For documentation:

  • Documenting work steps verbally
  • Speaking project reports instead of writing them
  • Following up on customer conversations

After transcription: making further use of the text

What you can do with the transcribed text:

  • Edit it directly – Correct errors and add to it
  • Have it summarized – basebox can analyze the text
  • Structure it – Organize it into bullet points or chapters
  • Translate it – Convert it into other languages
  • Process it further – Use it as a basis for other documents

Important note

Recognition quality depends heavily on audio quality and how you speak. Always check the transcribed text for accuracy, especially for important documents.

Tip

Start with short test recordings to get a feel for the optimal settings.

Problems with speech recognition or questions about audio quality? Contact support