AWS Storage Blog

Category: Amazon EC2

Amazon S3 Tables

How Tubular Labs reclaimed 50% of engineering capacity by rebuilding their 70TB pipeline on Apache Iceberg and Amazon S3 Tables

Customer Story | Amazon S3 Tables – Learn how Tubular Labs (part of Chartbeat Inc.) reclaimed 50% of engineering capacity by replacing fragile, file-based data pipelines with a Common Pipeline Runtime (CPR) built on Apache Iceberg and Amazon S3 Tables. By Tubular Labs Engineering (part of Chartbeat Inc.), in collaboration with the AWS Solution Architecture team […]

Amazon S3 Express One Zone thumbnail

Run Spark 31% faster and optimize compute costs with Amazon S3 Express One Zone on Amazon EMR

As your Spark datasets grow, storage latency often becomes the constraint, impeding application performance. Query runtimes stretch, and the bottleneck shifts from compute to how fast each node can read from Amazon Simple Storage Service (Amazon S3). We benchmarked this directly on Amazon EMR with TPC-DS at 3 TB scale. On an 8-node Graviton4 cluster […]

Orchestrating multi-agent AI architectures with Amazon S3 Files

​​​​Organizations are moving beyond single-model AI toward multi-agent architectures. In these systems, agents offload intermediate results to files rather than carrying everything in the prompt, because a large prompt inflates cost and degrades quality. A model’s context window is finite, so files become working memory that persists after a session ends. In multi-agent systems, a […]

FSxZ featured image

Optimize your self-managed PostgreSQL data warehouse with Amazon FSx for OpenZFS

A data warehouse is the analytical backbone of a modern enterprise, consolidating data from disparate sources into a single, authoritative view that enables complex queries, trend analysis, and confident decision-making. In financial services, this means sharper regulatory reporting, faster fraud detection, and deeper customer understanding. The operational reality is demanding. Enterprises juggle multiple source databases […]

Amazon S3 Express One Zone thumbnail

How Turso built a transactional database using Amazon S3 Express One Zone

Turso is a transactional, cloud database platform built on SQLite, serving developers and enterprises that need lightweight, high-density databases at the edge and in the cloud. Because each database is just a file, a single Turso compute node can host millions of them with no cold start. Turso delivers these databases in a Bring Your […]

Scalable cross-cloud data migration to Amazon S3 with distributed rclone

Migrating petabytes of data across cloud providers is one of the most operationally demanding tasks an organization can take on. At this scale, simple transfer approaches break down. Teams lose track of what has been copied and what has failed. Transfers stall and require constant manual intervention to restart. In some cases, teams need to […]

Automatically decompress files in Amazon S3 using AWS Step Functions

Every day, AWS customers process millions of compressed files in Amazon S3, from small ZIP archives to multi-gigabyte datasets. While decompressing a single file is straightforward, processing thousands of files efficiently requires complex orchestration, error handling, and infrastructure management. Consider this scenario: Your organization receives over 10,000 compressed files daily from partners, ranging from 5 […]

Amazon FSx for OpenZFS

Getting started with self-managed Oracle in AWS using Amazon FSx for OpenZFS

Organizations of all sizes run their enterprise applications and databases in the cloud. These organizations may choose to run self-managed databases on Amazon Elastic Compute Cloud (Amazon EC2) rather than using the fully-managed Amazon Relational Database Service (Amazon RDS) due to internal policies, Amazon RDS service maximums, and other reasons. When running self-managed databases in the cloud, there […]

Amazon S3 Multi-Region Access Points

How to use Amazon S3 Multi-Region Access Points to streamline and reduce the cost of writing across AWS Regions

Large global organizations often struggle to efficiently manage data copies across different geographic regions when using distributed object storage services. Although several approaches exist for cross-region data writing, common solutions such as data replication or streaming can be costly and introduce latency issues. Many customers have core services deployed globally across multiple Amazon Web Services […]

Amazon S3 Tables

Faster threat detection at scale: Real-time cybersecurity graph analytics with PuppyGraph and Amazon S3 Tables

Modern cybersecurity teams are facing unprecedented challenges in data analysis by the scale, complexity, and velocity of data. Cloud environments continuously generate massive amounts information in form of access logs, configuration changes, alerts, and telemetry. Traditional analysis methods of looking at these data points in isolation can’t effectively detect threats such as lateral movement and […]