SharePoint Syntex: Microsoft rolls out AI that automatically categorises documents
- Reference: 1601051167
- News link: https://www.theregister.co.uk/2020/09/25/sharepoint_syntex_microsoft_uses_ai/
- Source link:
SharePoint Syntex, currently in preview but with general availability promised for 1 October, is the first product based on a wider technology unveiled at the 2019 Ignite event called Project Cortex.
The core idea is to use AI to parse content stored in Microsoft's cloud, drawing not only on the words, images, and links in the documents, but also on other signals in the Microsoft Graph, such as who is engaging with the content and what departments they are in.
[1]
Syntex can drive document workflows such as approval after categorising documents though AI-powered analysis, presented in the Content Center
Microsoft said that after seeing how Project Cortex was used with preview customers, it has decided to have multiple projects based on the technology, rather than one. SharePoint Syntex is the first, a premium add-on for SharePoint online which is focused on using AI to automate content understanding and automation, such as routing a document to the right person for approval.
This is not the first time we have seen AI applied to SharePoint content. Microsoft introduced Office Delve in 2014, also based on the Office Graph, the theory being that it automatically shows users the documents that are most relevant to them. Delve has had little impact – will Syntex be different?
It is early days, but Syntex is more ambitious than Delve. Delve was focused on surfacing relevant content for a user, whereas Syntex can add metadata to documents that in theory could save substantial manual effort. Syntex could parse a purchase order, for example, work out the monetary value, the customer, and the region where the customer is based, and another process could forward it to the appropriate team to progress the order.
[2]According to general manager Seth Patton, Syntex processes three different types of content: images, forms, and unstructured documents. It will tag images with "thousands of commonly recognized objects", make tags by recognising handwritten text, and read the fields in forms including parsing of dates, numbers, names, and addresses.
Syntex documents are surfaced in a new Content Center, which sorts documents into libraries and shows the metadata it has extracted as columns. Syntex tagging can also be used for compliance, adding retention or sensitivity labels, and setting things like encryption, sharing restrictions, and conditional access policies.
[3]
Creating a custom model in Syntex by training based on files which have labelled content, identifying the metadata
The most intriguing part of Syntex is the ability to train new models for extracting metadata from documents. Every business has its own terms and categories. Syntex has a model creation feature where you can define entities, such as "Contractor" or "Fee amount", mark existing documents with labels identifying the values for these entities, and submitting these to train a model that will enable AI to extract them automatically from new documents.
As few as five files to train the model
Naomi Moneypenny, director of program management for Syntex, said at Ignite that as few as five files could be sufficient for training, particularly if users supply both positive and negative examples of a particular content type. Form processing, which should be the easiest type of content from which to extract metadata, has a specific form processing engine.
Content processed by Syntex does not have to live in SharePoint, but can also be sucked in from other sources via Microsoft Graph [4]content connectors . Examples of such sources include file shares, Azure SQL, Box, Amazon S3, Google Drive, SharePoint on-premises, and Salesforce.
Microsoft spoke at Ignite about new features planned for Syntex early next years, which include expanded model types, central model management, Syntex-based solutions for business processed, and more integration between Syntex and "knowledge improvements across Microsoft 365".
All a bit vague, but you get the impression that the company sees AI-driven content analysis as a significant piece in its 365 offering.
Whereas Delve was free for licensed SharePoint users, Syntex is a paid-for service available to E3 or E5 subscribers to Microsoft 365. The pricing looks complex, being per-user and limited to "500 items indexed by content connector, pooled", according to a slide presented at Ignite. Customers also get credits for form processing. Presumably additional fees apply if these limits are exceeded.
The problem with all the above is whether the company is over-promising when it comes to the benefits of Syntex. Considering the complexity of the underlying data science, the company's ability to simplify the usage of AI services, whether in Syntex or its other Cognitive Services portfolio, is not in doubt.
AI is an inherently imperfect technology, though, which is worrying in a business context if organisations depend on it too much, for example, to decide whether or not a document is confidential. As a paid-for service, Syntex will have to deliver high enough accuracy to justify its cost.
Whether or not Syntex flies, you can bet Microsoft, like others in the document management area, will continue to apply AI technology in the hope of making better sense of these repositories of unstructured data. ®
Get our [5]Tech Resources
[1] https://regmedia.co.uk/2020/09/25/approval.jpg
[2] https://techcommunity.microsoft.com/t5/project-cortex-blog/announcing-sharepoint-syntex/ba-p/1681139
[3] https://regmedia.co.uk/2020/09/25/createmodel.jpg
[4] https://docs.microsoft.com/en-us/MicrosoftSearch/connectors-overview
[5] https://whitepapers.theregister.com/
I take it that this will ensure retrieval has all the specificity of, say Bing searches. Or any of the other usual suspects - Google, Amazon, eBay (who suddenly seem to have got as crap as Amazon) etc.
Ahh... Sharepoint
What goes in seldom comes out.
At one employer we had to use sharepoint for all our docs. We duly started adding them to Sharepoint only to find a week or so later, that the URL that we'd lovingly saved when the doc was uploaded no longer worked. Welcome to the world of 404.
After an immense struggle, we found the docs. Only to find a month later that they'd moved again.
We carried in loading Sharepoint with all the docs but we also put them on a local (as not in the cloud) GIT repository.
After a year or so, even our management and this included the ones that had made using Sharepoint mandatory were using Git.
I moved on not long after that so I can't say how it all panned out but from then on, I avoided Sharepoint at all costs.
Surely …
… most corporate documents can be classified as bullshit, crap or waffle, except for marketing documents which are all classified as blatant lies?