You can catalog, discover, and classify unstructured documents alongside your structured data assets in Data Governance and Catalog to gain complete visibility across your data landscape.
This release includes the following key capabilities:
•Extract unstructured documents, including PDFs, Word documents, and text files from the following catalog sources:
- Amazon S3
- Databricks
- File System
- Google Cloud Storage
- Hadoop Distributed File System
- Microsoft OneDrive
- Microsoft SharePoint Online
- Microsoft Azure Blob Storage
- Microsoft Azure Data Lake Storage Gen2
- Microsoft Fabric OneLake
- Oracle Cloud Object Storage
- SFTP File System
•Automatically classify PDFs and text files into categories defined by you. You can classify unstructured documents using the new Document Classification capability in Metadata Command Center for the following catalog sources:
- Amazon S3
- Hadoop Distributed File System
- Microsoft Azure Blob Storage
- Microsoft Azure Data Lake Storage Gen2
- Microsoft SharePoint Online
- Oracle Cloud Object Storage
- SFTP File System
For information about unstructured data discovery and classification, see the corresponding catalog source help.
Handle unrecognized files during metadata extraction
You can now configure how Metadata Command Center handles files with unrecognized formats during metadata extraction from the following catalog sources:
•Amazon S3
•Databricks
•File System
•Google Cloud Storage
•Hadoop Distributed File System
•Microsoft Azure Blob Storage
•Microsoft Azure Data Lake Storage Gen2
•Microsoft Fabric OneLake
•Microsoft OneDrive
•Microsoft SharePoint Online
•Oracle Cloud Object Storage
•SFTP File System
You can configure the catalog sources to treat all unrecognized file formats as either structured or unstructured files. Optionally, you can define specific path-level exceptions to override whether the unrecognized files should be treated as structured or unstructured files.
For information about handling unrecognized files, see the corresponding catalog source help.
Enhanced catalog sources
This release includes the following enhancements to catalog sources:
Microsoft Power BI
You can now extract PaginatedReportDataset objects that connect to Power BI datasets which use DAX and MDX queries.
When you run a metadata extraction job on Microsoft Azure Data Factory pipelines, the job uses operational metadata from trigger runs to generate lineage for pipelines.