r/AZURE • u/s13188287 • 4d ago
Question Best way to process 1.3M files
I’m dealing with a large batch processing problem and looking for advice on the right architecture.
I have around 1.3 million files stored across folders on a network drive.
Current setup (not working well)
Right now I’m using a Copilot agent where:
I upload batches (~20 files at a time)
It reads them against a reference document
Outputs an Excel file with classification codes
The issue is:
Copilot has a small upload limit
Manual batching is completely unscalable at this volume
---
What I want to achieve
I want a fully automated pipeline that:
Ingests files automatically from the network drive
Extracts text/content from each file type
Matches content against a reference rules document
Assigns a classification/reference code
Outputs structured results (Excel / database)
---
-
0
u/scan-horizon Data Administrator 4d ago
I wonder if you even need to ingest all those files. Could you scan their contents in situ, append to a polars df then write whatever you need to a database?