The News: Direct Data Lake Queries Come to Aurora

The AWS News Blog recently announced that Amazon Aurora PostgreSQL can now query data directly from your data lake in Apache Iceberg and Parquet formats. This means you can combine live operational data with massive historical datasets without building complicated data pipelines.

This is significant because it removes the need for extract, transform, and load (ETL) processes for many common workloads. For years, businesses have had to copy data from a data lake into a database just to run a simple query, which costs time and money.

Why This Changes the Data Architecture Game

This update is a major step toward simplifying cloud computing for small and mid-sized businesses. The engine behind this capability is embedded directly into the database, allowing it to scan data in Amazon S3 without sending it over the network first. This keeps query processing close to the data, which improves speed and reduces data transfer costs.

From an expert perspective, the real value here is the ability to ask questions of all your data at once. Previously, you had to decide which data was "hot" (operational) and which was "cold" (archived), then manage the movement between them. Now, you can keep recent transactions in Aurora and historical records in the data lake, but query them together as if they were in one place.

This is particularly useful for building AI agents or real-time dashboards. These tools often need to reference both live events and historical context to make decisions, and this feature allows them to do so through a single, familiar PostgreSQL interface.

What This Means for Australian SMBs

For Australian SMBs, this is a practical solution to a common problem. Many businesses keep years of data in low-cost storage like Amazon S3 but find it difficult to use that data for daily reporting. This update lets you query that old data using your existing tools, making cloud migration projects simpler and less risky.

This also helps with the unique challenges of the Australian market, such as data sovereignty and managing tight IT budgets. You can keep sensitive operational data in Aurora while storing large archives in S3, and you only pay for the compute you use when you actually run a complex query.

What You Can Do Now

If you are currently running Aurora PostgreSQL or planning to, here are some immediate steps to consider:

  • Review your data storage: Identify which large datasets are sitting unused in S3 or data lakes. If they are accessed less than once a day, they might be perfect candidates for this new query model.
  • Map out high-value queries: List the reports that require combining current data (like today's sales) with historical data (like the last three years). These are the workflows that will benefit most.
  • Check your version: This feature is only available on specific major versions of Aurora PostgreSQL. Verify that your current database cluster can be upgraded to a compatible version.
  • Start small: Instead of migrating everything at once, pick one use case—like a monthly financial report—and test it against a Parquet file in S3 to see the performance gains.
  • Check your IAM roles: Ensure your database can securely access your S3 buckets by setting up the correct permissions, which is a prerequisite for this feature to work.

For Australian businesses looking to reduce operational overhead, this is a strong argument for modernising your data stack. If you are unsure how this applies to your current setup, MS&VG can help you assess your data landscape and identify where this new capability could unlock value.