Skip to content

// documents

Working with images

Overview

Photos, screenshots and scans: what basebox reads from images, where the limits are and how to use text recognition deliberately. In short: basebox reads the text in an image via text recognition (OCR) and makes it analysable like a document.

Which images you can upload

Format Typical for
JPG Photos, phone snapshots of documents
PNG Screenshots, graphics with text
TIFF Scans from document scanners
HEIC Photos from Apple devices
GIF Simple graphics
SVG, EMF Vector graphics and Windows metafiles

You upload images directly as standalone files – you do not have to embed them in a PDF or Word document first. Image-only PDFs and Word files that contain nothing but images are also read via text recognition.

What happens to an image

  1. You upload the image into the chat (paperclip or drag and drop) or into an app's knowledge base.
  2. basebox recognises the contained text via OCR. Where a GPU is available, this is accelerated; otherwise it runs on the CPU automatically.
  3. The recognised text is available to the assistant – you can summarise it, translate it, turn it into a table or ask questions about it.

Text recognition works the same everywhere: in the chat as in the knowledge base.

Typical use cases

  • Photographed form – "Transfer the details from the form into a table."
  • Screenshot of an error message – "What does this message mean and what can I do?"
  • Scanned letter – "Summarise the letter and state the deadline."
  • Whiteboard photo – "Turn the bullet points into clean minutes." (Handwriting is recognised only to a limited extent.)
  • Business card or receipt – "Extract name, company, amount and date."

Limits

basebox reads text – it does not describe images

An image without text – say a photo of a landscape or a diagram without labels – yields no content that basebox can process further. Do not expect an image description.

Further limits:

  • Handwriting is recognised only to a limited extent; printed text works considerably more reliably.
  • Poor quality – blurry, skewed, dark, small print – degrades the result.
  • Multilingual images work; mixed writing systems can produce errors.

Tips for good recognition

  • Photograph straight from above, in even light, without shadows.
  • When scanning, use 300 dpi and black-and-white or greyscale for text documents.
  • Crop away margins and background; the document should fill the image.
  • For several pages: prefer a multi-page PDF over many single images – the order is preserved.

Notes

Note

  • In the chat, the limit of 10 MB per file applies to images too. Phone photos are often larger than necessary – downsizing rarely hurts text recognition.
  • Videos are not processed. If you need the sound, extract the audio track; see Using speech-to-text.
  • Check recognised numbers – 0 and O, 1 and l are classic OCR errors.

Frequently asked questions

Can I paste an image directly into the chat (Ctrl+V)? Upload images via the paperclip or by drag and drop.

Why do I get nothing back for my image? It probably contains no recognisable text, or the quality is too poor. Check whether you can read the text well yourself.

Are images in a knowledge base stored permanently? Yes, like every other file of the app – until you remove them. Make sure you only upload images whose content all users of the app may see.

Need help? Contact support