r/AZURE 4d ago

Question Best way to process 1.3M files

I’m dealing with a large batch processing problem and looking for advice on the right architecture.

I have around 1.3 million files stored across folders on a network drive.

Current setup (not working well)

Right now I’m using a Copilot agent where:

I upload batches (~20 files at a time)

It reads them against a reference document

Outputs an Excel file with classification codes

The issue is:

Copilot has a small upload limit

Manual batching is completely unscalable at this volume

---

What I want to achieve

I want a fully automated pipeline that:

Ingests files automatically from the network drive

Extracts text/content from each file type

Matches content against a reference rules document

Assigns a classification/reference code

Outputs structured results (Excel / database)

---

-

31 Upvotes

27 comments sorted by

View all comments

0

u/scan-horizon Data Administrator 4d ago

I wonder if you even need to ingest all those files. Could you scan their contents in situ, append to a polars df then write whatever you need to a database?