POV | What We Believe, On the Record | Mactores

AWS Data Lake: Architecture, Services & Build Guide

Written by Ammar Nizami | Oct 9, 2026, 2:13:43 PM

Key takeaways

  • Governance and architecture decide whether a lake succeeds. Gartner expects 80% of D&A governance initiatives to fail by 2027, so design for governance from day one.
  • S3 Tables avoid a warehouse copy. Registered with the Glue Data Catalog and Lake Formation, S3 Tables let Athena, Redshift, EMR and Spark query the same Iceberg data.
  • Governance turns a schema change into a same-day fix. Cataloged schemas and lineage show what a change affects before it breaks anything downstream.
  • Query cost depends on partitioning and storage format. Athena charges $5/TB scanned, and the 72 TB worked example runs about $950–$1,000/month.
  • Build order and timeline matter. A specialist gets a first governed dataset live in about 3–5 weeks, against 8–14 weeks in-house.

A team stands up an S3 bucket, points every new pipeline at it, and eighteen months later three groups in the same company are computing "monthly active customers" three different ways from three different tables in that bucket, with nobody able to prove which number is right. That's not a data problem. It's the default outcome of treating an AWS data lake as a place to put things rather than a system that needs the same design discipline as anything else running in production.

An AWS data lake, done properly, is a centralized store built primarily on Amazon S3 that holds structured, semi-structured, and unstructured data at any scale, governed well enough that analytics, BI, and machine learning teams can trust what they're querying.

The gap between that definition and the scenario above is almost entirely architecture and governance, not storage. Gartner has predicted that 80% of data and analytics governance initiatives will fail by 2027, largely because they get built as generic, one-size-fits-all programs instead of being scoped to a specific business outcome. That's precisely the failure mode that turns a lake into the ungoverned bucket above.

This piece is a working reference: what an AWS data lake actually is, how it differs from a warehouse or lakehouse, which AWS services map to which architectural layer, the patterns teams use in production, a build sequence, and when building in-house makes sense versus when it doesn't.

What Is an AWS Data Lake?

An AWS data lake is a centralized repository, built primarily on Amazon S3, that stores data in its native format: CSV, JSON, Parquet, Avro, log files, images, sensor streams, without forcing a schema at write time. Structure gets applied later, when the data is read, which is why data lakes are described as "schema-on-read" rather than "schema-on-write."

That's the core distinction from a traditional data warehouse, and it's worth being precise about, since the three terms get used loosely.

  • Data warehouse. A structured, schema-on-write system (think Amazon Redshift) optimized for fast, consistent SQL analytics over curated, modeled data. Good for BI dashboards and financial reporting; expensive and rigid for raw, high-volume, or unstructured data.
  • Data lake. A schema-on-read repository (Amazon S3 plus a catalog) optimized for storing everything cheaply and flexibly. Good for raw ingestion, data science, and ML training data; historically weaker on transactional consistency, governance, and query performance without extra tooling.
  • Data lakehouse. A newer pattern that adds warehouse-like features (ACID transactions, schema enforcement, indexing, time travel) directly on top of lake storage, usually through an open table format like Apache Iceberg. On AWS this is delivered through Amazon S3 Tables (managed Iceberg storage), the AWS Glue Data Catalog, and Amazon Redshift's zero-ETL and Spectrum integrations, which let a warehouse query lake data directly.

It also helps to be precise about what a data lake isn't, since the term gets stretched to cover things it doesn't mean:

  • Not just cheap storage. An S3 bucket full of files without a catalog and a permission model isn't a data lake. It's a bucket, and it becomes a liability at roughly the same rate it becomes an asset.
  • Not a warehouse replacement. A lake is the raw, flexible layer underneath governed BI, not a competitor to a warehouse. Most mature environments run both, plus a lakehouse pattern to avoid duplicating data between them.
  • Not a one-time migration. Treating the build as a project with an end date is the most common reason zones and permissions stop being maintained six months after go-live. A data lake is infrastructure, not a deliverable.

 

Core Architecture Components, Mapped to AWS Services


An AWS data lake isn't one service. It's a set of layers that work together, and how those layers are assembled is what separates a reusable data platform foundation from another bucket the next reorganization has to clean up. Here's how each layer typically maps to the AWS platform.

Table format layer

Storing files in S3 isn't the same as storing tables. Without a transactional table format, concurrent writes, schema changes, and record-level updates get messy fast. Apache Iceberg, delivered on AWS through Amazon S3 Tables as a fully managed option, gives lake data ACID transactions, schema evolution, and time travel.

The mechanism is worth understanding, because it answers the question most evaluations get stuck on. When you enable S3 Tables integration, AWS registers your table buckets with the AWS Glue Data Catalog as a federated catalog (s3tablescatalog), and Lake Formation registers that catalog as a governed data location.

That single registration is what lets Athena, Redshift, EMR, and Spark all query the same Iceberg tables directly, under one set of Lake Formation permissions, without anyone copying the data into a second system first. So: do you need a warehouse copy of this data? Usually no, provided the table is registered this way.

Deciding early whether a given dataset needs Iceberg-backed tables or plain S3 objects is one of the higher-leverage architecture calls in the whole build. See the FAQ below for when that tradeoff is worth making.

Storage layer

Amazon S3 is the foundation almost every AWS data lake is built on: durable, cheap, and elastic, with storage classes and Intelligent-Tiering to manage cost as data ages. Most designs split S3 into zones, commonly raw (or "bronze"), curated (or "silver"), and consumption-ready (or "gold"), using prefixes or separate buckets per zone.

Ingestion layer

Getting data into the lake reliably is where most early pain lives.

  • AWS Glue and AWS Database Migration Service (DMS) for batch loads from relational databases and SaaS sources. DMS is also the usual first stop when a source system is a legacy database being retired outright rather than just replicated; that broader migration is its own body of work, covered under by application and database modernization.
  • Amazon AppFlow for no-code ingestion from SaaS platforms such as Salesforce and ServiceNow.
  • Amazon Kinesis Data Streams, Amazon Data Firehose, and Amazon MSK (Managed Streaming for Apache Kafka) for real-time and event-driven ingestion.
  • AWS Transfer Family for SFTP and FTPS-based file transfers where legacy systems require it.

Cataloging and metadata layer

A data lake without a catalog is just a bucket. The AWS Glue Data Catalog holds the technical metadata: schemas, partitions, and table definitions that let query engines find and interpret data.

Above that sits a business-facing layer that moved fast over the past year. Amazon SageMaker Catalog, built on Amazon DataZone and accessed through Amazon SageMaker Unified Studio, adds discovery, ownership, and an approval-based subscription model for requesting access to a dataset, including S3 Tables specifically since SageMaker Catalog added governance for S3 Tables in May 2025. Lake Formation still enforces the underlying permissions; SageMaker Catalog is the layer people browse to find and request a dataset, particularly in data mesh-style setups with multiple domains and accounts.

Processing and transformation layer

This is where raw data becomes usable data.

  • AWS Glue (Spark-based ETL jobs) and AWS Glue DataBrew (visual data prep) for most transformation workloads.
  • Amazon EMR for large-scale, custom Spark, Hive, or Presto/Trino processing where Glue's managed model is too constrained.
  • Amazon Managed Service for Apache Flink for continuous stream processing.

Governance and security layer

AWS Lake Formation is the control plane for fine-grained, table- and column-level permissions across the lake, tied into IAM for identity and AWS KMS for encryption. Amazon Macie helps discover and classify sensitive data such as PII and financial data at scale, and AWS CloudTrail provides the audit trail governance and compliance teams will ask for.

Consumption and query layer

Amazon Athena provides serverless, pay-per-query SQL directly against S3 data. Amazon Redshift Spectrum and Redshift's zero-ETL integrations let warehouse users query lake data without moving it. Amazon Quick (formerly QuickSight) handles BI and dashboarding, and Amazon SageMaker consumes lake data for model training and feature engineering.

Orchestration layer

AWS Step Functions, Amazon Managed Workflows for Apache Airflow (MWAA), and Amazon EventBridge coordinate the pipelines that move data through the earlier layers on schedule or in response to events.

04

What Governance Buys You

The abstract version of "governance matters" is easy to nod along to and hard to act on. Here's the concrete version, drawn from the pattern that repeats across data lake builds.

A source system, a CRM say, renames a customer identifier field, or starts sending it in a slightly different format, without telling anyone downstream. That happens constantly and is rarely announced in advance. What occurs next depends entirely on how the lake was built.

In an unzoned bucket with no catalog, the change is invisible until a downstream job either fails with a cryptic error or, worse, keeps running and silently produces wrong joins. Someone notices weeks later when a report doesn't reconcile, and the investigation starts from zero: nobody can say which tables consumed the bad data, or for how long. 

With a Glue Data Catalog in place but no Lake Formation permissions, a crawler eventually picks up the schema drift and the technical metadata updates. But because access wasn't governed, a team may already have built a dashboard directly against the old schema without anyone knowing that table had a consumer, so fixing the source doesn't fix the dashboard, and finding out who's using what becomes its own investigation.

In a zoned lake with Lake Formation permissions and Iceberg-backed tables through S3 Tables, three things are true that weren't true before. The schema change can go through Iceberg's schema evolution instead of a full rewrite.

Every consumer of that table is a registered grant, not a mystery, so the affected team can be notified directly instead of found after the fact. And the curated-zone transform job, if it was built against an explicit schema contract rather than a permissive one, fails immediately in a controlled environment instead of downstream in a report someone had already trusted.

None of this stops schema drift from happening. Source systems will keep changing without warning regardless of how well the lake is built. What governance buys is the difference between a same-day fix and a two-week forensic exercise.

At a program level, that difference compounds into fewer duplicate copies of the same table, faster onboarding of new sources once the ingestion and cataloging pattern exists, query costs that scale with actual usage instead of idle warehouse capacity, and a more reliable foundation for machine learning and production AI agents that depend on consistently structured, governed data.

None of it fixes bad source data on its own, and none of it removes the need for someone to own governance decisions on an ongoing basis. A data lake is infrastructure, not a substitute for that job.

Common AWS Data Lake Architecture Patterns

Most production AWS data lakes fall into one of three patterns, or a combination of them.

  1. Batch, layered (medallion) architecture

    Data lands in a raw zone unchanged, gets cleaned and conformed into a curated zone, and is aggregated into a consumption zone optimized for specific use cases. Typically built with Glue jobs or EMR for transformation, the Glue Data Catalog for schema management, and Athena or Redshift Spectrum for querying. This is the default starting point for most teams because it's well understood, easy to reason about, and forgiving of imperfect early decisions.

  2. Streaming / near-real-time architecture

    Kinesis or MSK ingest events continuously, Amazon Managed Service for Apache Flink processes them in flight, and results land in S3, often in Iceberg tables through S3 Tables, for both real-time dashboards and historical analysis. This pattern matters for fraud detection, IoT telemetry, clickstream analytics, and anywhere the value of data decays within minutes rather than days.

  3. Governed, multiaccount (data mesh) architecture

    Larger organizations increasingly split the lake across business-domain-owned AWS accounts, with Lake Formation handling cross-account permissions and SageMaker Catalog providing a shared, business-facing catalog for discovering and requesting access to domain-owned data products. This pattern trades some simplicity for organizational scalability. It's the right call once a single central data team becomes the bottleneck.

Batch and streaming aren't mutually exclusive. Many lakes run both, landing streaming data into the same zoned structure a batch pipeline uses, so downstream consumers don't need to know or care how the data arrived.

How to Build an AWS Data Lake: A Step-by-Step Guide

This is the sequence at a high level. None of these steps are trivial in practice; each has its own set of decisions. But the order matters more than most teams expect.

  1. Start from use cases, not infrastructure.

    Identify the two or three analytics, BI, or ML use cases that justify the build. A lake designed around "store everything" instead of "answer these questions" tends to grow storage faster than the catalog, until nobody can say who owns a given table or how fresh it is.

  2. Design the zone structure and data domains

    Decide on raw/curated/consumption zones (or your organization's equivalent), how domains are separated, and your S3 bucket/prefix and partitioning strategy up front. This is expensive to retrofit later.

  3. Stand up storage and access controls first

    Provision S3 buckets, apply encryption (KMS), and configure Lake Formation permissions before data starts flowing, not after.

  4. Build ingestion pipelines for priority sources

    Start with the two or three highest-value source systems rather than trying to onboard everything simultaneously.

  5. Catalog everything as it lands

    Use Glue crawlers or explicit schema definitions so every table is discoverable in the Glue Data Catalog from day one.

  6. Transform and model for consumption

    Build the Glue/EMR jobs that move data from raw to curated to consumption zones, applying data quality checks along the way.

  7. Expose data through the right consumption path

    Athena for ad hoc SQL, Redshift for governed BI, SageMaker for ML; match the tool to the consumer, not the other way around.

  8. Operationalize: orchestration, monitoring, and cost controls

    Add Step Functions or MWAA for scheduling, CloudWatch for pipeline monitoring, and S3 lifecycle policies and Athena query limits to keep costs predictable as usage grows.

     

The teams that get stuck usually skip step 1 or step 3; they built storage and pipelines before agreeing on governance, or before anyone could say precisely which questions the lake was meant to answer.

What an AWS Data Lake Costs

There's no single number for this, but the cost drivers are well defined, and the pricing is public. List pricing below is for US East (N. Virginia), as of September 2026; check the AWS Pricing Calculator for other regions or current rates.

AWS data lake cost drivers by layer, us-east-1 list rates
Layer Service Pricing Model list rate (use-east-1)
Storage

S3 Standard

per GB-month

$0.023

Storage S3 Standard-IA

per GB-month

$0.0125

Storage

S3 Glacier Deep Archive

per GB-month $0.00099

Ingestion

Amazon MSK, kafka.m5.large broker

per broker-hour, plus storage

$0.21/hr + $0.10/GB-month

Processing AWS Glue (ETL jobs, crawlers)

per DPU-hour, billed per second (1-minute minimum for ETL jobs, 10-minute minimum for crawlers)

$0.44
Processing Amazon EMR EC2 instance price plus the Amazon EMR price $0.192/hr EC2 for m5.xlarge, plus $0.048/hr for EMR
Query
Amazon Athena per TB scanned
$5.00
Query Amazon Redshift Serverless

      per TB scanned

$5.00
Query Amazon Redshift Serverless
per RPU-hour $0.375
Governance AWS Lake Formation

 permissions and           governance

$0 for permissions and access control (Storage API and governed tables are billed separately; standard S3 and Glue rates apply underneath)

 

Two caveats before budgeting from this table. Redshift Serverless does not bill Spectrum scans separately, so the two query lines should not be added together for a serverless deployment. And the S3 tiers carry minimums: Standard-IA has a 128 KB minimum billable object size and a 30-day minimum duration, and Glacier Deep Archive has a 180-day minimum duration plus retrieval charges.

A Worked Example

Take a single-domain lake ingesting 2 TB of new data a month and retaining three years online (roughly 72 TB accumulated), with a daily batch transform job and moderate analyst query volume. Assume a blended storage mix of 20% still in S3 Standard, 50% aged into Standard-IA, and 30% aged into Glacier Deep Archive via lifecycle rules:

  • Storage: 72,000 GB × blended rate (~$0.011/GB) ≈ $795/month
  • Processing: a daily Glue job at 10 DPUs for 1 hour, 30 days ≈ 300 DPU-hours × $0.44 ≈ $132/month
  • Query: analysts scanning roughly 300 GB/day of well-partitioned curated data ≈ 9 TB/month × $5 ≈ $45/month
  • Total: roughly $950–$1,000/month

That excludes streaming ingestion, Redshift, EMR, data transfer, and support plans, and it assumes the data is partitioned well. Run the same query pattern against unpartitioned raw data, and the query line alone can be 10 to 50 times higher, since Athena and Redshift Spectrum bill by bytes scanned, not by query.

Where the leverage actually is:

  • Partition and compress. Columnar formats (Parquet, or Iceberg tables through S3 Tables) plus sensible partitioning is the single biggest lever on query cost, since you pay per byte scanned.
  • Automate tiering. An S3 lifecycle rule moves data through storage classes without anyone remembering to do it:

{

"Rules": [{

"ID": "age-curated-zone",

"Filter": { "Prefix": "curated/" },

"Status": "Enabled",

"Transitions": [

{ "Days": 90, "StorageClass": "STANDARD_IA" },

{ "Days": 365, "StorageClass": "GLACIER" }

]

}]

}

  • Right-size Glue and turn on job bookmarks so re-runs don't reprocess data that hasn't changed.
  • Cap runaway queries. Athena workgroups can enforce a per-query data-scanned limit, which catches an unpartitioned scan before it turns into a four-figure line item.
  • Let S3 Tables handle compaction. Iceberg tables through S3 Tables auto-compact small files, which otherwise inflate both storage and per-request costs on high-frequency streaming writes.
  • Use Redshift Serverless for spiky BI load rather than a provisioned cluster; it scales to zero when nobody's querying.

When to Build In-House vs. Partner With a Specialist

Neither path is automatically right. The real question is how much of the work is genuinely novel to your business versus well-trodden AWS pattern-matching, and the gap between the two paths is easiest to see in numbers rather than adjectives.

In-house build vs. specialist partner for an AWS data lake
Dimension Pilot Production
Time to first governed dataset 8–14 weeks on a first build, most of it spent learning Lake Formation and Glue patterns 3–5 weeks, using patterns already validated on prior builds
Engineering capacity assumed 2–4 dedicated data engineers for the first quarter, on top of existing platform work 1 engineer from your team paired with a partner pod; your engineer owns the system after handoff
Rework cost when zone or permission decisions are unwound Commonly 20–30% of the original build effort re-spent on re-architecture once a zone boundary or permission model has to change Low; zone and permission models are checked against production precedent before anything is built
AWS service depth required High: Glue job tuning, Lake Formation permission models, and Iceberg/S3 Tables design have real failure modes Brought in from day one, transferred to your team over the engagement
Best fit when Your team has already shipped one or two AWS data lakes and has a quarter of slack capacity You're on a fixed deadline, building your first lake, or your engineers are better spent on domain-specific work

None of these is universally "better." A team with prior AWS data lake experience and slack capacity can build in-house successfully; a team on a fixed deadline, or building its first lake, usually loses more time to rework than it saves by keeping the work internal. Knowing which situation you're actually in, honestly, matters more than the choice itself.

Outcomes to Expect From a Well-Built AWS Data Lake

A data lake isn't the goal; it's infrastructure in service of decisions people actually need to make faster. It's worth being specific about what changes when the build is done well, because vague promises ("better data culture," “data-driven decision making”) are exactly what makes engineering leaders skeptical of this category in the first place.

Done well, an AWS data lake changes the shape of a few specific problems, not everything at once:

  • Fewer duplicate, drifting copies of the same data. A single governed source reduces the "whose number is right" arguments between teams.
  • Faster onboarding of new data sources. Once the ingestion and cataloging pattern exists, adding a new source is incremental work, not a new project.
  • Query costs that scale with actual usage. Athena's pay-per-query model and Redshift Spectrum mean you're not paying for idle warehouse capacity to query infrequently accessed data.
  • Governed self-service. Analysts and data scientists can discover and request access to data through Lake Formation permissions and SageMaker Catalog, rather than filing tickets and waiting on a central team.
  • A real foundation for ML and AI initiatives. Feature engineering and model training need consistent, discoverable, well-labeled data, which is exactly what a properly zoned, cataloged lake provides.

What it doesn't do on its own: fix bad source data, replace the need for data quality processes, or eliminate the need for someone to own governance decisions on an ongoing basis. Those stay true regardless of how the lake gets built.