{yourcompany}os
Login

{yourcompany}os

How to Get AI-Ready Data Without a Cleanup Project

You don't sort your way to trustworthy data - you build a process that never lets the mess happen

Published

Simon Dilhas (Owner, abstract ag and digitalbau gmbh)

Everyone repeats the same line: data is the foundation for AI in AEC. Fair enough - but open any real project's file storage and you'll see what that foundation actually looks like. Old versions sitting next to new ones. Nobody, not even the people who work on the project daily, can say with confidence which file is current. Duplicates, orphaned folders, three naming conventions in one directory. The instinct is to fix this with a cleanup: sort it, tag it, archive the rest. That instinct is wrong. This guide covers why file storage ends up this way, why sorting after the fact doesn't hold, and how {yourcompany}os makes the process produce trustworthy artefacts instead.

What "clean data" quietly assumes

The "data is the foundation" pitch assumes a static problem: a pile of files that's disorganized because nobody has gotten around to organizing it yet. Sort it once, keep it tidy, and AI has what it needs. That assumption treats the mess as a backlog. It isn't. It's the visible output of dozens of uncoordinated people and tools writing into the same folders over years, with no shared record of what happened, when, or why one file replaced another. A cleanup project sorts a snapshot. The moment work resumes, the same forces that created the mess start writing the next one.

Why the file storage is always a mess

Nobody sits down and decides to create duplicate, outdated files. It happens because the process has no memory. A plan gets revised in a meeting and re-exported under a slightly different name. A subcontractor emails a PDF that never makes it into the shared folder. Two people fix the same clash independently and save two "final" versions. None of this is negligence - there's simply no system tracking which action produced which file, so nothing enforces that the old one gets archived or the new one gets flagged as current. Sorting after the fact treats the symptom. The folder will look the same way again in three months, because the thing that generated the disorder - an unrecorded, ad hoc process - is still running underneath it.

Example: a submittal review, not a folder cleanup

The difference is concrete. The cleanup approach looks like: "someone goes through the shared drive, renames files, deletes obvious duplicates, and writes a naming convention doc" - a one-time pass that starts decaying the day it's finished. With process-first thinking, the same submittal review becomes three visible, recorded steps:

  1. Submit - a contractor uploads a document through a defined step, not into an open folder. The system captures who submitted it, when, and what it's replacing, if anything.
  2. Review - a reviewer approves, rejects, or requests revision inside the process, not in a side email thread. The decision and its reasoning are attached to that specific file, not implied by its filename.
  3. Publish - only an approved submittal becomes "current." Superseded versions aren't deleted or renamed by hand; the process marks them as superseded automatically, because the process knows what replaced what.

Nobody had to sort anything. The current version is current because the process made it so.

Where to put the effort: process over sorting

A practical rule of thumb: don't spend effort making a messy artefact archive look clean - spend it on the step that creates the artefact. Any document that matters for downstream decisions - a submittal, a change order, an as-built drawing, a clash report - should come out of a defined process step with a recorded actor, timestamp, and decision, not out of "someone saved it to the project folder." A file's trustworthiness isn't a property you can add later by sorting; it's a property of how the file came to exist. Get that right at the source, and there's nothing left to clean up.

How this looks in practice

On {yourcompany}os, every artefact a process produces carries its own audit trail: who submitted it, who reviewed it, what decision was made, and what it superseded, logged automatically as part of the workflow - not reconstructed afterward from filenames and folder dates. When an AI step later reads that document (to summarize a submittal, flag a discrepancy, or draft a response), it isn't guessing whether the file is current. The process already answered that question the moment the file was created. That's what makes the data trustworthy enough for AI to build on: not that a human sorted it, but that a process governed how it got there.

Key takeaways

The mess in AEC file storage isn't a backlog to clear - it's the ongoing output of processes with no memory, and sorting a snapshot of it fixes nothing for long. {yourcompany}os's position is that trustworthy data comes from where the process defines and records how each artefact was created, reviewed, and superseded, not from cleaning up after the fact.

Stop treating "clean data for AI" as a filing project. Put the artefacts that matter (submittals, change orders, as-builts) behind a defined process step with a recorded actor and decision, and the audit trail that produces is what makes the data trustworthy, not a naming convention.

← All guides

Contact

abstract ag

Imprint · Picassoplatz 4, 4052 Basel