AWS re:Invent 2025 - Accelerate data discovery with object metadata in Amazon S3 (STG357)
AWS Events
1 views • 7 months ago
Video Summary
Short Highlights
- Keywords
S3 metadata, data discovery, data management, Apache Iceberg, SQL queries, journal table, live inventory table, AWS, AI, data scientists, security, compliance
Key Details
-
Short Keypoints
-
S3 stores over 500 trillion objects, with GenAI revolutionizing data value.
- S3 metadata provides automatic metadata extraction from S3 objects, queryable via SQL.
- Two tables are created: a journal table (audit log, refreshed within minutes) and a live inventory table (snapshot, refreshed hourly).
- Metadata tables are in Apache Iceberg format and stored in S3 table buckets, managed by AWS.
- Use cases include finding sensitive data, tracking deletions, ensuring encryption compliance, and managing storage with actions like restoring from Glacier.
- S3 metadata can be queried using AWS analytics services (Athena, Redshift), open-source engines (Spark, Trino), and natural language via AWS re:Invent tools like MCP for S3 tables.
- Real-world impacts include reduced processing times for medical imaging customers and confident data migration for digital content companies.