AWS re:Invent 2025: Five S3 Updates That Signal the Future of Enterprise AI Storage
AWS shared some big announcements at re:Invent 2025, signaling something important about tomorrow’s enterprise workloads and what storage platforms must deliver to support them. The announcements stand out for what they collectively say about the convergence of AI, analytics, and object storage.
For organizations evaluating storage strategies for the next decade of AI evolution, let’s examine what AWS announced, what it means for enterprise infrastructure, and why a hybrid approach could be the architectural insurance policy against the next market disruption.
1. S3 Vectors goes GA: AWS’s Bet on Vector Storage
Initially previewed in July 2025, S3 Vectors represents Amazon’s answer to a question enterprises are increasingly asking: Where should we store the embeddings, vectors, and feature stores that power our AI applications?
What S3 Vectors Delivers
- Up to 20 trillion vectors per bucket (40x increase from preview capacity)
- Up to 2 billion vectors per index
- Support for the massive-scale RAG (retrieval-augmented generation) applications enterprises are building
- Performance for production AI workloads:
- 2-3x faster query performance for frequently-accessed vectors compared to preview
- Low-latency retrieval enabling real-time semantic search and AI agent applications
- Integration with AWS Bedrock Knowledge Bases and OpenSearch Service
Why This Matters
Vector databases are critical infrastructure for retrieval-augmented generation (RAG) which requires storing and searching billions of high-dimensional vectors representing semantic meaning. Until now, enterprises faced a choice: deploy specialized vector databases or build custom solutions. AWS’s decision to embed vector storage directly into S3 signals object storage and AI infrastructure are converging.
For infrastructure teams, this convergence creates an opportunity: simplify architecture by consolidating storage platforms.
Learn what we expect to see in the world of cloud computing in 2026.
2. 50TB Objects: Preparing for AI-Scale Data
AWS increased the maximum S3 object size from 5TB to 50TB—a 10x jump that could address three industry use cases.
High-Resolution Video and Media
The challenge: 8K video, 360-degree footage, and ultra-high-resolution content for film, broadcasting, and virtual production exceed 5TB per file.
The impact: Media organizations can now store complete projects as single objects, simplifying version control and reducing multi-part upload complexity. For studios working with uncompressed 8K RAW footage, a single day of shooting can generate 40-50TB—previously requiring multi-file splits that complicated workflows. [Link to film/movie blog]
Seismic and Geospatial Data
The challenge: Oil and gas companies, climate research institutions, and mapping organizations work with massive geospatial datasets. A single high-resolution seismic survey can exceed 20TB.
The impact: Scientists and engineers can work with complete datasets as atomic units, improving data integrity and reducing the operational overhead of managing thousands of file chunks. For climate modeling, where simulations generate petabytes of continuous data, larger object sizes mean fewer metadata operations and better performance.
AI Training Datasets
Modern AI models train on consolidated datasets—think Common Crawl at 400TB, or massive image datasets for computer vision. Previously, these required elaborate sharding strategies.
The impact: Data scientists can structure training data as fewer, larger objects, reducing API call costs (fewer objects = fewer list operations) and simplifying data pipeline architecture.
This increase is AWS acknowledging that enterprise data (not just hyperscalers) is growing faster than previous assumptions. Massive objects will be the rule, not the exception. Infrastructure teams need to evaluate whether their storage platforms scale to match cloud providers’ expanding capabilities.
3. 10x Faster Batch Operations: Scaling Data Management
AWS announced that S3 Batch Operations now run up to 10x faster and scale to 20 billion objects per job—addressing a pain point that becomes acute at enterprise scale. Previously, batch jobs at scale could take days or weeks. The 10x performance improvement means operations that took 10 days now complete in 24 hours.
This will be great for data mobility and meeting compliance requirements, but managing 20 billion objects can be incredibly complex. With scale comes new challenges.
The question for hybrid architectures: Can your on-premises storage management tools operate at this scale? Many enterprise storage systems with traditional management interfaces struggle with datasets exceeding millions of objects, let alone billions.
FlashBlade recently surpasses 3 trillion stored
4. S3 Tables with Intelligent-Tiering: The Iceberg Revolution
Last year, AWS introduced S3 Tables, bringing Apache Iceberg table format natively to S3. This year, they announced two enhancements: automatic cross-region replication and Intelligent-Tiering for cost optimization.
What Are S3 Tables and How Have They Evolved?
Apache Iceberg provides ACID transactions, schema evolution, and time travel capabilities on top of object storage—bringing database-like features to data lakes. Now, AWS automatically moves table data across three access tiers based on usage patterns.
The Iceberg enhancements signal AWS’s commitment to making S3 the foundation for modern data platforms—not just raw object storage, but the transactional data lake layer that competes with traditional data warehouses.
For enterprises, this begs a few strategic questions:
- Do we invest in AWS-native Iceberg on S3, knowing it deepens AWS dependency?
- Can we run Iceberg workloads on-premises or other clouds with equivalent features?
- What’s our exit strategy if AWS pricing or capabilities change?
5. S3 Storage Lens: Visibility for Better FinOps
AWS introduced three new S3 Storage Lens features to provide deeper visibility into storage usage and performance. It’s a boon for FinOps teams, who can query storage metrics alongside application logs and cost data to build chargeback models and find cost optimization opportunities.
- Performance metrics. Organizations with complex S3 architectures (hundreds of buckets, thousands of prefixes) can identify bottlenecks and optimize workload placement. For AI/ML pipelines where storage latency directly impacts GPU utilization, this helps pinpoint slow storage as the root cause.
- Analysis of billions of prefixes. Large S3 deployments organize data with prefix structures (think /customer-id/year/month/day/), creating billions of distinct prefixes. Storage Lens can now analyze billions of these for fine-grained cost allocation, lifecycle policy optimization, and compliance reporting.
- Direct export to S3 Tables. Storage Lens metrics export directly to S3 Tables (Iceberg format), enabling SQL-based analysis with tools like Athena, Redshift Spectrum, or Spark.
| The Everpure approach: On-premises storage systems often lack this level of detailed, queryable telemetry data. For enterprises pursuing hybrid strategies, this creates a challenge: maintaining operational parity between cloud and on-premises storage requires investing in observability infrastructure.Pure1, the Everpure cloud management platform, provides unified visibility across all Pure arrays—whether on-premises, at edge, or in cloud (via Everpure Cloud)—addressing the observability gap for on-premises storage. |
What These Updates Mean for Enterprise Storage Strategy
Viewed individually, each S3 update addresses a specific technical requirement. Viewed collectively, they reveal strategic trends that infrastructure leaders must address:
The era of “storage is storage, compute is compute” is ending
S3 Vectors, 50TB objects, and new integrations all point to the same reality: storage platforms must natively understand AI workloads, not just store their data. Modern storage platforms need to:
- Handle vector embeddings and semantic search queries
- Support massive training datasets as atomic objects
- Integrate with AI/ML frameworks through standard APIs
- Deliver the low latency and high throughput that GPU-accelerated workloads demand
For infrastructure teams, this means evaluating storage not just on capacity and cost, but on AI readiness.
Assumptions About Scale Are Shifting Dramatically
20 billion objects per batch job. 50TB maximum object size. 20 trillion vectors per bucket. AWS is acknowledging that enterprises operate at scales that previous storage architectures weren’t designed for. Organizations need storage that assumes:
- Petabyte-scale datasets are normal, not exceptional
- Billions of objects are standard, not edge cases
- AI workloads will demand 10x more performance and capacity every few years
For procurement and architecture decisions, this means asking: Does my storage platform scale to where our workloads will be in 2030, not where they are today?
Everpure FlashBlade delivers S3-compatible object storage designed for the exact hybrid scenarios AWS announcements implicitly acknowledge enterprises require.
S3-Compatible Hybrid Data Platform — Technical Architecture Overview
| Domain | Architecture / Capability | Technical Details |
| S3 Implementation | Native S3 (No Gateway) | FlashBlade speaks S3 natively—no proxy, translation, or gateway layers |
| AWS S3 API Compatibility | Applications written for AWS S3 run unchanged | |
| Tooling Support | AWS CLI, Boto3, Terraform S3 backend, MinIO client | |
| Performance Architecture | Low-Latency Access | Sub-millisecond latency for inference and real-time workloads |
| Parallel Throughput | Massive parallel I/O to sustain GPU training pipelines | |
| Predictable Performance | No throttling, no noisy-neighbor effects | |
| AI & GPU Integration | GPU Feeding at Scale | Designed to supply hundreds of GPUs without storage bottlenecks |
| NVIDIA Validation | Validated with NVIDIA DGX SuperPOD and BasePOD | |
| Roadmap | (Becoming available this year) S3 over RDMA (2026): ~250 GB/s per 5-chassis system | |
| Hybrid Data Mobility | On-Prem AI Training | Avoid cloud data ingress costs |
| Cloud Replication | Native replication for DR or cloud-native services | |
| Burst & Repatriation | Move compute without relocating primary datasets | |
| Unified Protocol Support | Object + File | S3, NFS, and SMB on the same platform |
| Infrastructure Simplification | Eliminates “file vs. object” architectural trade-offs | |
| Data Protection & Efficiency | Immutable Snapshots | Ransomware-resistant, even with admin compromise |
| Data Reduction | Global deduplication + compression | |
| Capacity Predictability | Guaranteed 4:1 data reduction minimum | |
| Operations & Lifecycle | Non-Disruptive Upgrades | Controller and storage upgrades without data migration |
| Unified Management | Pure1 cloud-based monitoring and management | |
| Cloud Integration | Cloud Block Store | Extends Everpure into AWS and Azure |
| Consistent Data Services | Same capabilities across on-prem, edge, and cloud |
Hybrid S3 Architecture Checklist
The AWS re:Invent announcements make clear that S3 API is the standard. The question isn’t whether to support S3, but where your S3-compatible storage lives and who controls it.
Validate that your chosen platform delivers:
- Native S3 API support (not gateway or translation layer that adds latency/complexity)
- Performance matching or exceeding cloud (sub-millisecond latency, massive parallel throughput)
- File and object protocol support (eliminate silos between file and object workloads)
- Cloud replication and integration (seamless data movement when you need cloud benefits)
- Immutable snapshot protection (ransomware resilience without separate backup infrastructure)
- Non-disruptive upgrades (no maintenance windows as data and workloads grow)
- Unified management across environments (on-premises, edge, and cloud through single pane)
- Proven at scale (validated with AI frameworks, backup tools, analytics platforms you use)
- Flexible consumption model (CapEx, subscription, or pay-as-you-grow to match financial strategy)
Everpure FlashBlade checks these boxes—but the framework applies to any hybrid storage evaluation. The goal is architectural independence that survives the next decade of infrastructure evolution.
The Future Is Hybrid, Not Binary
AWS re:Invent 2025 demonstrated that S3 is evolving to support the most demanding enterprise workloads—AI at scale, petabyte-sized objects, sovereign cloud requirements, and management complexity at billions of objects.
The question leaders should be asking is “How do we build S3-compatible infrastructure that serves our business requirements without creating the next vendor dependency that constrains our future?” Because, in an era where infrastructure markets shift rapidly—VMware acquisitions, cloud pricing changes, new AI paradigms every 18 months—architectural flexibility is the most valuable capability you can build.