A crowded service inbox is rarely only an email problem. It exposes unclear rules for urgency, ownership and the next operational step. AI can distribute that ambiguity faster, but it cannot repair it on its own. A sound pilot therefore starts with a limited task: recognise the type of incoming message, propose a transparent category and put it in front of a responsible person.
That boundary matters for a Swiss SMB. A quote request needs speed, a complaint needs judgement, and a suspicious attachment needs a security check. All three may arrive in the same mailbox, but they must not follow the same automation path.
Classification is not a business decision
In the first phase, inbox triage may identify a message type, suggest urgency, flag missing information and propose an owner. It must not confirm prices, promise deadlines or formulate a legal position. This separation makes results auditable and prevents a language model from quietly becoming the company policy.
A reliable flow has two exits. Routine cases reach the responsible team with a short summary and an optional draft. Uncertain cases go to a separate review queue without a draft. The model must be allowed to say that it is unsure; forcing a confident result for every email is a design defect.
A category matrix built for daily work
| Message | AI may prepare | A person decides |
|---|---|---|
| Appointment or callback | Extract contact details, preferred time and topic | Availability and final confirmation |
| Quote request | Identify service, location, deadline and missing inputs | Price, scope and commercial terms |
| Existing job | Match customer number and responsible project | Changes that affect cost or delivery |
| Complaint | Flag urgency and forward the complete context | Tone, goodwill, liability and response |
| Suspicious email | Apply a risk label and isolate the message | Opening links or attachments and any further action |
The matrix must be tested with real examples from the company. Ten theoretical labels are less valuable than one hundred anonymised messages from two normal working weeks. That sample reveals the words customers actually use and the places where current responsibilities contradict each other.
A worked pilot example
Assume a service team receives 120 messages in ten working days. Forty-two concern appointments, 31 ask for quotes, 27 relate to ongoing work, 12 are complaints and eight do not fit any category. The pilot works on copies and has no sending permission. A person confirms or corrects every proposed assignment.
After two weeks, the quality of the prose is not the primary result. The team checks whether fewer messages remain ownerless, whether internal forwarding has fallen and whether urgent cases become visible earlier. If 18 of 120 emails are assigned incorrectly, the answer is not full automation. Categories, examples and routing rules need another iteration.
Data protection starts before the model
The Swiss Federal Data Protection and Information Commissioner states that the FADP applies directly to AI-supported processing and requires transparency about purpose, operation and data sources. For an inbox project, document which fields are processed, where content travels, how long logs remain and who can see the output.
A responsible pilot minimises data. Signatures, long email histories and attachments should not be sent to a model when the subject and latest message are enough for classification. Permissions must also be separated: reading, classifying, drafting and sending are four different capabilities. The first pilot needs only the first two.
Phishing is not another customer category
Switzerland's National Cyber Security Centre warns about links, attachments and messages that pressure recipients into taking action. A triage system should never open an unknown attachment merely to produce a better summary. Security scanning and quarantine happen before language classification.
A convincing writing style is not a trust signal. Sender details can be forged. Payment changes, passwords and new bank details must be verified through a second channel regardless of the AI score. This control belongs in the workflow rather than in a prompt that can be overlooked.
Turning the test into an operating process
- Capture a baseline: for ten working days, measure time to first owner, internal forwards and open messages at close of business.
- Define five to seven categories: each category receives examples, an owner and explicit exclusion cases.
- Run read-only: the AI writes labels and suggestions to a separate view but does not change production email.
- Record corrections: store the category, error type and correct decision rather than retaining unnecessary message content.
- Gate permissions: drafts are enabled only after assignment is stable; automatic sending remains a separate later decision.
This flow can connect process automation with a narrowly scoped AI assistant. The value does not come from the model brand. It comes from visible responsibility at every handoff.
Four measurements instead of a demo
- share of messages assigned to the correct owner;
- median time to the first human action;
- number of unnecessary internal forwards;
- share of critical cases incorrectly treated as routine.
The final measure is a stop condition. A pilot can save time and still be unsuitable if complaints or security cases slip through. A valid two-week decision may be to automate sorting only for appointments and callbacks while leaving every other category manual.
Primary sources used for this framework
- FDPIC: AI and data protection
- NCSC: handling email securely
- Swiss SME Portal: the new Federal Act on Data Protection
Inbox triage is useful when it exposes responsibility and removes repetitive handling without hiding decisions. That narrow and measurable design is a stronger starting point for a Swiss service team than an autonomous agent with broad mailbox permissions.
FAQ
Should an AI send email automatically in the first pilot?
No. Start read-only: classify, flag missing details, suggest an owner and at most prepare a draft. Sending and binding commitments remain human decisions.
Which emails should never be auto-approved?
Complaints, legal or financial commitments, suspicious links or attachments, and messages containing sensitive personal data should always receive human review.
Which metric shows value fastest?
For a small service team, the share of open messages without a clear owner at the end of the day is often more useful than the number of generated replies.
Does Swiss data-protection law apply to AI triage?
Yes. The FDPIC states that the technology-neutral FADP applies directly to AI-supported processing.