You are a records professional. Your organisation has in its possession a legacy repository to which no content has been added in the past five years. The repository contains over 100 million items (documents/messages etc.). No retention rules have been set on this content. You estimate that only 30% of the content has any ongoing value.
The repository might be an old email system, an old SharePoint system, or a legacy set of network file shares left over from the days of on-premise computing.
Let us imagine that you have decided to undertake an intervention, with the help of AI, to identify and remove the 70% of content within the repository that has no ongoing value. You have a cross disciplinary project team containing both records management and data science expertise. You have the backing of your organisation.
How would you go about making the intervention?
Design choices
You have three key design choices to make when setting up the intervention. You have choices as to:
- what level of aggregation you target your intervention at
- what task you set your AI model/agent/set of agents
- whether you use supervised or unsupervised learning
You could target your intervention:
- directly at individual items – the documents, messages etc. that make up the repository OR
- at containers – the sites, accounts or drives that the items sit within
You could task task the AI model with:
- classifying – assigning each item or container to one of a set of pre-defined categories OR
- clustering – in which the model looks for similarities between items (or between containers, or between items within each container) and groups like items with like
You could use:
- supervised learning – where you transfer some of your records knowledge to the AI model so that the AI model can assign content to retention categories OR
- unsupervised learning – where the AI model transfers information about content (or content groupings) to you so that you can use your records knowledge to assign that content to retention categories
In theory that gives you nine potential approaches, one for every permutation of these choices. In practice these choices are linked, and the choice in effect boils down to:
- developing a classifier that looks at every item in the repository and assigns it to a retention category under your supervision OR
- developing dashboards that can present you with analytical information about each container within the repository, and through which you can partition or cluster content within each container. This would enable you to understand and make retention decisions on each container. It would also enable you to intervene within any container that you choose to retain
Both these approaches scale up to repositories of digital heap scale. Both approaches enable you to wrest control of repositories that were previously seen to be unmanageable.
Lets see how you might use these two approaches to tackle a repository such as the one described at the start of this post.
Applying a classifier directly to every item
This set of approaches targets the intervention at the 100 million items that make up the repository. The aim is to develop a classifier that can evaluate each item and assign it to a retention category. You can define the scope of each category, and how many categories there are. For example:
- you might define a set of retention categories that correspond to different areas of activity of your organisation, each with its own accompanying retention rule OR
- you might simply define two categories, one for content of ongoing value and the other for content without ongoing value
This approach uses supervised learning. Your aim is to teach the model the retention categories that you have defined. You might do this through one or more of the following:
- writing a set of rules that define the features likely to be present (and/or absent) in items responsive to each category
- collating a training pack of documents responsive to each category (and of documents that are not responsive, but are superficially similar)
- prompting a large language model with your rule set for each category, and providing the model with access to your training pack
- iteratively working with a large language model to identify and build a training pack.
You will deploy the model once you have satisfied yourself that, when presented with a sample of items from the repository, it correctly categorises enough of the items to meet whatever threshold you have set for success.
Faced with a repository of 100 million items, the model would make 100 million decisions, with each decision assigning an item to a category. You might aim to check a statistically significant sample of those decisions, on an ongoing basis, to ensure that the model continues to meet or exceed the success threshold set even when faced with new content.
Analysing every container and intervening within those you wish to retain
When using the second set of approaches you intervene through the containers that content sits within. A container is a site, account or drive that an organisation provides to a team or individual to enable them to accumulate content within a corporate system. All corporate multi-purpose collaboration, document management and messaging systems use some form of container:
- in a SharePoint repository every item/document sits within a site
- in an email repository every message sits within an email account
- in a shared drive repository every document sits within a drive
With a container based approach you might split your AI intervention into two stages:
- in stage 1 you might use AI to help you find the most valuable containers in the repository (so that you can remove the rest)
- in stage 2 you might use AI to look for the least valuable content within each of those valuable containers (so that you can remove trivial content from the containers you have chosen to keep)
So in our imaginary scenario, with our notional target of removing 70% of content:
- you might, in stage 1, seek to identify the most valuable 35% of containers (and delete the rest)
- in stage 2 you might identify the least valuable 14% of content within each of the remaining containers, and remove it
After completing these two stages you would be left with approximately 30% of the content of the original legacy repository.
This set of approaches uses unsupervised learning. You would use a dashboard to look at each container. On the basis of the analysis available to you in the dashboard, you make a retention decision on the container, and in phase 2 you intervene in those containers that you chose to retain.
The first role of AI in this approach is to extract or reconstitute key information about the container. This might usefully include:
- contextual information about the container – for example the identity and role of the team or individual to whom the container was provisioned
- summary information about the container– for example the date range within which content was added, the entities (places, organisations etc.) frequently mentioned within it, numbers of items within the container, volumes of different filetypes within the container etc
The second role of AI in this approach is to provide you with a means to divide the container into smaller chunks. For example AI could be used to
- partition into sections – if there is significant structure within the container, for example a folder structure, the AI could suggest partitions of that structure into sections
- cluster into clusters – if there is no significant structure within the container (as is often the case within email accounts) then AI could be used to cluster the contents, so that similar items are grouped together
- create time based divisions – it could partition messages within an email account by year depending on the date on which each message was sent or received. It could partition folders (or groupings of folders) within a folder structure depending on the year in which content was last added to the folder/folder grouping
The process of arriving at a set of sections/clusters that are suitable is an iterative one. You will almost certainly want to change the partitions/clusters initially suggested by the model, for example by changing the boundaries of the partitions, or changing the numbers of clusters generated. Once you are happy with the sections/clusters, you could identify the clusters/sections that relate to the less important aspects of a team’s work, and remove them.
Combining both sets of approaches
The two sets of approaches have different strengths. This makes them complementary to each other – each can compensate for the weaknesses of the other.
As a general rule developing a classifier is at its most effective when:
- the categories of content that you want to find are relatively evenly distributed across the repository AND
- you need to treat all content of a category in the same way regardless of the role of the team or individual that created or received it
In such circumstances it should be both feasible and useful to develop a classifier that can be used across an entire repository.
In contrast working through containers is at its most effective when:
- the categories of content that you wish to find are not uniformly distributed, but instead are ‘clumped’ within certain places within the repository, for example in the sites/drives/accounts of individuals or teams playing particular roles AND
- your treatment of a category of content would vary depending on the role of the team or individual that created or received it.
If you adopt one approach and find there is an obstacle in your way, then incorporating some aspect of the thinking and the methods of the other approach might help you overcome that obstacle.