File documents by what they contain
Reads each document that arrives by email and files it by what it says - or leaves it for you when it cannot tell.
What it does
The integration steps this workflow runs, in the order it first runs them.
- 1
GmailNew email trigger - 2
GmailList emails - 3
GmailFind email - 4
GmailDownload attachments - 5
GmailAdd label to message - 6
OpenAIQuery model - 7
BoxUpload file
How it works
Everything the template sets up, and what to fill in before the first run.
Mail gets filed by who sent it, which is why the contract from your accountant ends up in the accounts folder. This reads each PDF that arrives by email and files it by what the document itself says - so an invoice is an invoice whoever sent it, and a contract from the same address goes somewhere else.
You describe up to three folders: a name, the Box folder to put things in, and a line about what belongs there. Each document is asked what kind of document it is, what it is about, what organisation or product it concerns, its reference number, its date and any total. Those answers - and nothing else - are what the folder is chosen from. The sender's address and the subject line are never shown to the step that decides, because they are exactly the thing this is meant to ignore.
When it cannot tell, it does not guess. The answer has to be one of your folder names, word for word - only stray quote marks and a single full stop are forgiven, and the comparison ignores capitals. Anything else - a hedge, a folder you did not configure, two names at once, the word UNCLEAR, or nothing at all - means the document is left exactly where it is, nothing is copied anywhere, and the email is labelled 'Needs sorting' so you have a short queue to work through. A misfiled document is worse than an unfiled one, because nobody goes looking for it.
The order here is the other way round from most filing: the mail is marked before anything is copied, not after. Copying is not repeatable in the way a spreadsheet row is - a second run reads the document again and can land on a different folder, and the same contract sitting in two folders is a misfile by another name. So if a run dies part way, a copy is missing and the run log says which one, rather than a document turning up twice in two different places.
A document that could not be read at all goes the same way: nothing is filed, the mail is labelled for you, and the run log says which document it was.
Filed copies keep their original file name. Only PDF attachments are treated as documents - a signature logo is an image a document reader will happily describe, and it is not something to file. A run reads at most 25 messages and handles at most 'maxDocumentsPerRun' documents, which can be anything from 1 to 25.
New mail is what wakes it, but the sweep is what does the work, so several documents landing together are all sorted. Gmail only hands over the attachments of a sender's newest message in the inbox, so an older one from the same person cannot be opened - the run log names it rather than quietly passing over it - and the mail has to still be in the inbox when the run happens.
Reading a document is a smart action, so an occasional run pauses for review before it finishes.
Fill in 'documentSearch' with a Gmail search that finds the mail carrying documents, then at least one category: its name, its Box folder (the id, or the folder's address in Box) and a line saying what belongs in it. The two label steps and 'maxDocumentsPerRun' have working defaults.
Start from a workflow that already works.
Add "File documents by what they contain" to your workspace, connect its apps, and make it yours. No credit card required.
Free plan available · No credit card required
