Skip to main content

Have something to say?

Tell us how we could make the product more useful to you.
Completed

Exported PDFs show US-format dates in the header

The metadata header on an exported PDF renders the date in US format — Date: 1/27/2025, 7:05:46 PM — while dates inside the item body are in the original UK long form, e.g. On Sun, 26 Jan 2025 at 18:13. For a UK disclosure that's confusing at best and ambiguous at worst: 1/2/2025 reads as 2 January to the recipient and 1 February to us. Likely cause The header date is formatted at conversion time; a locale defaulting to en-US rather than en-GB would produce exactly this. Reported in support conversation 197.

Completed

Never-redact list: protect specific names from every redaction pass

A list of names, addresses or terms that must never be redacted, regardless of what a word list, Auto-Redact or Search & Redact would otherwise match. Why The requester's own name has to survive redaction, and so do things like the organisation's name or a case reference. Today a broad word list or Auto-Redact can catch them, and the only fix is to spot it and remove the redaction by hand — on every item. The ask A project-level "never redact" list that every redaction method respects. Simpler and independent of automatic name discovery, and useful on its own. Asked twice in support ("How do I use a whitelist of names that should not be redacted ever?" and "How do I whitelist certain names and email addresses which I don't want to redact?"). Logged from conversations 67 and 66.

Completed

Redact every name except the requester's

For a SAR, the job is almost always the same: keep the requester's name, redact everyone else's. Today that means building a word list of every other name in the mailbox first — the inverse of what people actually want to say. The ask Tell RedactBox who the requester is, and have it find and redact all the other names across the project, leaving the requester's alone. How often it comes up Asked five different ways by five separate customers in support: "Can your app automatically redact all names save for that of the SAR requester?" "I'd like the tool to automatically identify any names not listed in the SAR and redact them" "Bulk remove names not associated with the person in question" "If a file contains the requester's name but also names I want to redact, how do I do that?" "Can we flag specific personal data before upload and have it redact everything?" It's the most-requested capability in the support inbox after demos. Current workaround Word Lists + Search & Redact — put every name to redact in a list. Works, but the customer has to discover the names themselves, which is the hard part. Logged from support conversation 537 (and 11, 31, 94, 127).

Wide tables are silently cut off in exported PDFs

When an item contains a table wider than the printed page, the exported PDF silently loses the columns that overflow the right-hand margin. The customer gets a document that looks complete but isn't, with no warning anywhere. What happens We render items to PDF at A4 with 1cm margins. Chromium shrinks over-wide content to fit up to a bounded factor and then simply clips whatever is left. The clipped text is not drawn on the page and is not in the PDF's text layer, so it cannot be found, redacted, searched or recovered. Real example A delivery-exception report pasted from Excel into Outlook: 14 columns, roughly 1850px wide. The exported PDF prints as far as "Address 2" plus two characters of the postcode. Category, Issue, Repeat, the free-text Comment and Closed are gone from every row. One cell that reads Admin-Mitcham-039-Customer contacted for delivery instructions not followed on 23/01... prints as Admin-Mitcham-0 and stops. Why it matters This is a data-completeness problem on a disclosure product. Items are exported for DSARs and legal disclosure, so handing over a document with content missing is worse than failing loudly — and nothing currently tells anyone it happened. It also produces a confusing second-order effect: a term can be redacted successfully in the interface while no black box ever appears, because the text it matched is not on the page at all. Suggested direction Detect and flag first — cheap, and turns a silent loss into a visible one. At conversion, compare the rendered text against the source and warn on the item when content did not fit. Then remedy — render items that overflow in landscape, or scale the page down further for wide content, so the columns actually print. Raised from support conversation 583, where the missing text was found while investigating a separate export-gate problem (now fixed).

Harry ElliottinBugs26 days ago
Completed

Create a way to un-triage from the redact screen

Sometimes, from the redact screen you realise that something should not have been triages (e.g. an email chain that contains duplicates of the previous email copies). It would be great to be able to untriage those directly from the redact screen instead of having to try and find them again in the triage screen.

Export the attachments with the PDFs

Export the attachments with the PDFs as an option, unredacted, so you can then do something about them afterwards. Somehow link them in the PDF.

Attachment Manager

Add a view to view all attachments at a glance in emails, and eventually import them once this is supported.

Completed

Deleted files leave source uploads and attachment blobs in storage

When a file is deleted, the database rows are removed and the logs report success, but the underlying objects stay in Cloud Storage indefinitely. Evidence (production, 16 July deletion, still present 20 July) Logs show the full happy path: Attachments cleaned up and unreferenced blobs queued for reap, Starting async PDF cleanup, File deleted successfully. Despite this, thousands of attachment blobs remain under the project's attachments/ prefix in the uploads bucket. The original uploaded source files also remain under the user's uploads/ prefix — they do not appear to be removed by the delete path at all. Four days later none of these objects have been reaped. Why it matters This is a retention and data-protection gap, not just wasted storage. A customer who deletes their content — often the whole point of a redaction workflow — still has the original documents and attachments held on our infrastructure, including personal data belonging to third parties. Deletion must actually delete. Suggested next step Verify the reap queue is receiving rows and is actually being drained, confirm the delete path removes the source upload as well as attachments, and backfill a cleanup for objects already orphaned by past deletions.

Harry ElliottinBugs3 months ago
2
Completed

File deletion intermittently returns 500 before eventually succeeding

Deleting a file from a project intermittently fails with a 500 before a later retry succeeds. Observed repeatedly on 15-16 July in production. Evidence One file returned 500 on four consecutive delete attempts, then succeeded. A second file returned 500 twice before succeeding. A third returned 500 twice the previous day. One delete returned 200, and an immediate repeat of the same request returned 404 — so the client cannot tell a real failure from an already-completed delete. Why it matters Deletion is a data-protection operation. A customer clearing their own content sees an error and cannot tell whether the data was removed. Repeated retries also mean the cleanup path runs more than once for the same file. Suggested next step Capture the underlying error behind the 500 (the response body is generic), make the delete idempotent so a repeat returns success rather than 404, and surface a clear outcome to the user.

Harry ElliottinBugs3 months ago
2
Completed

Blank Pages at the beginning of PDF Exports

Fixed and deployed. Blank pages at the start of an exported PDF were caused by how documents authored in Word/Outlook declare their page layout: it forced a page break that left the document header alone on the first page, with the content pushed onto the next page (and, for multi-section messages such as bounce notifications, over several pages). The exporter no longer honours that author-supplied page layout, so the header and content now start together on page one. Genuine, deliberate page breaks are still respected. Note: this applies to documents processed from now on. PDFs that were already exported keep their existing layout — re-exporting the affected documents produces the corrected version.

Harry ElliottinBugs3 months ago
High Priority
Completed

Redaction comment quick selection

Having to type similar comments while redacting is very tiresome. It would be useful to have some presets to quickly click/choose from: TPD: Third party personal data NPD: Not the requester’s personal data SEC: Security-sensitive information LPP: Legal professional privilege CON: Confidential information OOS: Out of scope DUP: Duplicate already provided This would speed up the redaction process considerably, and keep things consistent across files. It would be great if you could quickly add your own redaction reasons as well?

Completed

Background PDF Generation

We have several thousand emails that are triaged for redaction, none of which have PDFs pre-generated for redaction. This means we cannot do word list across all files, and each time we choose an email it takes several seconds to generate and we are getting a lot of lag/failures.

Completed

Ability to mark files/emails as "Complete" in redact mode

We are going to be redacting several thousand emails and documents. We need to know whether we’ve finished processing each file - sometimes files might have no redactions, so we need to be able to mark these as ‘complete’ to be able to track our progress. Add to that, the ability to filter by “Complete” or “Incomplete”

Completed

Wordlist redact numbers are all over the place

When doing a redaction using a wordlist, the numbers it shows are all over the place. For example, for one keyword it shows 250 matches across 100 items. Then, when inspecting, it shows “Redact 4223 matches” Then, when clicking the “Redact 4233 matches” it changes to a confirm button: “Redact 4223 matches of 40”

Apply wordlist per file

It’d be great to be able to quickly apply word list redaction on a per mail/file basis. As we’re processing large mailboxes, the PDFs have not been rendered for most items, so we are having to do them one at a time.

Completed

Increase upload rate limit

When trying to upload many files, I get the error: Upload Rate Limit You're uploading too quickly. Please wait 30 seconds before uploading more files. This really shouldn’t be happening - the whole point of this application is to upload many files and emails! Either the rate limiting should be handled transparently, or the limit should be removed?

1High Priority
Completed

Make email body links open in new tab by default

When clicking links in email body, they should probably open in a new tab by default - otherwise it’s very frustrating to have to come back find your place again when going back to triage/redaction.

Open and import password-protected PDFs

Password-protected PDF attachments now show a clear "password-protected" notice instead of locking the page up — but there's still no way to get at the contents. Add a way to unlock these files: prompt for the password when previewing or importing, then store an unlocked copy so the document behaves like any other — preview, search, redact and export. Raised off the back of the locked-out display bug report.

Completed

[bug] Password protected PDF attachments cause locked out display

When a PDF attachment is password protected, the page prompts you for a password. When this is cancelled, it just prompts again, soft locking the page until you refresh.

Completed

Implement stable sorting algorithm / Multiple column sorting

When you sort emails by a column that has a lot of duplicate values (e.g. sorting by “From”) and then take an action on an email, the order of the emails can jump around randomly. The emails should remain in the same order, even if one is triaged. This could be solved by including a final ordering param (e.g. always have a final order by ID or date). It would also help to be able to sort by multiple columns, e.g. From ASC subject ASC