Licensed to be used in conjunction with basebox, only.
// documents
Working with images
Overview
Photos, screenshots and scans: what basebox reads from images, where the limits are and how to use text recognition deliberately. In short: basebox reads the text in an image via text recognition (OCR) and makes it analysable like a document.
Which images you can upload
| Format | Typical for |
|---|---|
| JPG | Photos, phone snapshots of documents |
| PNG | Screenshots, graphics with text |
| TIFF | Scans from document scanners |
| HEIC | Photos from Apple devices |
| GIF | Simple graphics |
| SVG, EMF | Vector graphics and Windows metafiles |
You upload images directly as standalone files – you do not have to embed them in a PDF or Word document first. Image-only PDFs and Word files that contain nothing but images are also read via text recognition.
What happens to an image
- You upload the image into the chat (paperclip or drag and drop) or into an app's knowledge base.
- basebox recognises the contained text via OCR. Where a GPU is available, this is accelerated; otherwise it runs on the CPU automatically.
- The recognised text is available to the assistant – you can summarise it, translate it, turn it into a table or ask questions about it.
Text recognition works the same everywhere: in the chat as in the knowledge base.
Typical use cases
- Photographed form – "Transfer the details from the form into a table."
- Screenshot of an error message – "What does this message mean and what can I do?"
- Scanned letter – "Summarise the letter and state the deadline."
- Whiteboard photo – "Turn the bullet points into clean minutes." (Handwriting is recognised only to a limited extent.)
- Business card or receipt – "Extract name, company, amount and date."
Limits
basebox reads text – it does not describe images
An image without text – say a photo of a landscape or a diagram without labels – yields no content that basebox can process further. Do not expect an image description.
Further limits:
- Handwriting is recognised only to a limited extent; printed text works considerably more reliably.
- Poor quality – blurry, skewed, dark, small print – degrades the result.
- Multilingual images work; mixed writing systems can produce errors.
Tips for good recognition
- Photograph straight from above, in even light, without shadows.
- When scanning, use 300 dpi and black-and-white or greyscale for text documents.
- Crop away margins and background; the document should fill the image.
- For several pages: prefer a multi-page PDF over many single images – the order is preserved.
Notes
Note
- In the chat, the limit of 10 MB per file applies to images too. Phone photos are often larger than necessary – downsizing rarely hurts text recognition.
- Videos are not processed. If you need the sound, extract the audio track; see Using speech-to-text.
- Check recognised numbers – 0 and O, 1 and l are classic OCR errors.
Frequently asked questions
Can I paste an image directly into the chat (Ctrl+V)? Upload images via the paperclip or by drag and drop.
Why do I get nothing back for my image? It probably contains no recognisable text, or the quality is too poor. Check whether you can read the text well yourself.
Are images in a knowledge base stored permanently? Yes, like every other file of the app – until you remove them. Make sure you only upload images whose content all users of the app may see.
Need help? Contact support