How To Scale Unstructured Data Classification Without Manual Overload

Direct Source Verification: This story is aggregated from Forbes (forbes.com). Full reporting rights and copyright belong to the primary publisher.
Effective classification can help businesses manage security, compliance and access, but relying on employees to review and label content manually doesn’t scale well.

gettyAs organizations accumulate growing volumes of unstructured data, understanding what they have and how it should be handled can quickly become resource-intensive. Effective classification can help businesses manage security, compliance, access and other data governance needs, but relying on employees to review and label content manually doesn’t scale well.

Organizations therefore need practical ways to make classification more efficient without sacrificing accuracy or oversight. Below, members of Forbes Technology Council share strategies for classifying large volumes of unstructured data while keeping the manual workload manageable.

One practical approach is to classify data by its actual usage. Instead of classifying every stored file, prioritize information accessed by employees, applications and AI agents. Usage patterns can reveal where business, security and compliance risks are concentrated. This focuses classification resources on active, higher-value data, while rarely used information can follow simpler archival and retention policies, significantly reducing manual workload. - Salice Thomas, Wipro Limited

Let AI do the first pass and give humans the last word. I have seen teams use LLMs to auto-tag documents into a draft taxonomy, then route only the low-confidence cases to people. That flips the workload: Instead of classifying everything, your experts review maybe 10%. The trick is starting with a small, ugly taxonomy and letting real usage refine it, not designing the perfect one up front. - Sarah Choudhary, Ice Innovations

Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?

Use AI-assisted tagging with semantic indexing capabilities to classify data by meaning, owner, sensitivity and lineage. For instance, an organization could use AI to classify complex documents like contracts, invoices and purchase orders; emails; case files; and so on, routing only uncertain or high-risk cases for human review while governance controls ensure accountability. - Faruk Muratovic, Deloitte Consulting LLP

Don’t get stuck trying to fit everything into detailed categories. Begin with a handful of categories that actually matter for your business. For example, in payments, it’s not about perfectly labeling every single piece of data; it’s about quickly spotting what needs attention, what’s risky and what can run on autopilot. This way, teams can focus on the decisions that truly need a human touch, making the whole process more scalable and effective. - Serhii Zakharov, PayDo

Use AI-assisted auto-classification with human-in-the-loop validation. Train a model on a representative sample of your data that subject matter experts have already labeled, then let it classify at scale, flagging only low-confidence items for human review. This flips the economics: Instead of humans touching everything, they spend time correcting edge cases and continuously improving the model. - Prashant Darisi, Octave

The practical answer is AI, but the elephant in the room is this: Many organizations are data-rich and insight-poor because their data is scattered and inconsistent. When data lives in siloed systems with no shared definitions, AI cannot magically eliminate inconsistencies, and employees can’t leverage it. Classification at scale requires a consistent framework and clear governance rules. - Alex Saric, Ivalua

Don’t treat classification as a stand-alone job. It’s really the most scalable task of all: Bake it into the ingestion process, tagging the data as it enters the environment and leaves. When classification is part of data workflows and not a post-facto process, the burden on your people is minimal and ongoing coverage continues. The backlog that makes the work unmanageable exists only because, well, you’ve already let the data go everywhere. - Maitrik Patel, Apple

The answer, of course, is to leverage AI. Use AI to scan the content of every file and understand its meaning and intent. Use this semantic context alongside other metadata to automatically determine the nature of the document and an appropriate sensitivity level. But do not conflate AI’s capability with a product that’s effective at scale. Deploy a data security governance tool that uses AI securely, without training on your data, and scales to handle your entire data estate. - Madhu Shashanka, Concentric AI

Put human judgment where it compounds: defining the use case and the labeling rules. Refine them until two people applying them agree, then let AI handle what it’s confident about and route the rest to a person. Human workload then scales with category ambiguity, not volume. Measure it: Relabel a random sample independently and track where the two disagree. Disagreement means the definition isn’t clear enough yet. Fixing it reduces manual work. - Chhaya Methani, Microsoft Corporation

Use AI to discover natural groupings before people classify the data. Build a similarity graph, cluster related documents into domains, and use an LLM to generate a domain-specific ontology and extract relevant entities and relationships. Experts can validate the ontology and review exceptions instead of tagging every document, making classification scalable while retaining human oversight. Start with one set of related documents, then expand. - Shekhar Iyer, Arango

Do not read every file. Classify the pipe, not the pile. Tag data at ingestion by source, owner and destination: public, internal or secret. That is enough for most risk. Use a model only to propose a label for new unstructured blobs. A person reviews the unclear slice, never the full lake. Store the label next to the object so the next job does not have to think again. If a team cannot say where a file came from and who may see it, no classifier will save them. Origin and access do most of the work. - Andrii Stetsenko, Wyllo

We have created a context layer that enriches content with metadata, classification, relationships and domain knowledge. This business context foundation maps the relationships across content, people, projects, systems and domain knowledge, then assembles the unique context each AI agent needs to do its job. The context layer helps us deploy AI faster, reduce retrieval costs and maintain data governance policies. - Amrit Jassal, Egnyte

Classify at the point of creation, not in a cleanup project. Attach an AI classifier where documents, tickets and logs enter the system and are labeled with a confidence score, with only low-confidence cases routed to people. People then improve the model by correcting a small sample each week rather than tagging everything. The manual work becomes a bounded review, not an endless task. - Mayank Bhola, Testmu AI (Formally Lambdatest)

Let AI do the first pass at the moment someone acts on the document: a save, upload, print or scan. Let AI classify right there, using what the action captures and a file scan can’t: the user, the intent and where it’s headed. Route only low-confidence cases to people. The backlog never forms because classification rides an action already taken. Classify at the action, not the archive. - Corey Ercanbrack, Vasion

Classify for the decision, not for the taxonomy. Most teams build a 200-node ontology first and never finish labeling against it. Start with the few labels that actually change what happens to a document—retention, access and redaction—and let everything else stay unclassified. A small taxonomy tied to a real control is finishable. A complete one never is. - Kiran Kodithala, N2N Services, Inc.

As AI adoption accelerates, knowing where sensitive data resides is essential. AI-driven data classification uses machine learning to continuously discover, label and prioritize unstructured data by sensitivity and business value. Automating classification while reserving human review for high-risk cases strengthens governance, improves compliance and reduces the risk of sensitive data exposure. - Hakan Ekmen, P3 communications

The most practical way to classify large volumes of unstructured data without creating an unsustainable workload is to use AI to understand the data. AI can help build trust in the data that will fuel projects by answering key questions along the way: Do we need this data? Is it sensitive information? Is it resilient? What is in this data? Without these answers, organizations put AI projects and other initiatives at risk by relying on data they do not fully understand. - Rick Vanover, Veeam

Start by questioning whether the data needs classifying at all. A customer data platform can work with unstructured data alongside structured enterprise systems while keeping each dataset in its original form and location. Data is pulled together only when needed and written back to the source, reducing any need to move data or create new structures while reducing the attack surface for bad actors. - Martin Taylor, Content Guru

Original Source
https://www.forbes.com/councils/forbestechcouncil/2026/10/05/how-to-scale-unstructured-data-classification-without-manual-overload/
Visit Forbes ↗
SHARE STORY:
𝕏 f in

Related Coverage in Business