KnowBite KnowBite

Data Engineering

Data Engineering knowledge cards on KnowBite — curated from specialist publications, updated throughout the day.

  1. Discover and govern Snowflake data using SageMaker Unified Studio

    Connect Snowflake to Amazon SageMaker Unified Studio to build a unified data catalog. Query federated Snowflake tables without moving data, publish enriched assets to SageMaker Catalog, and validate data quality with AWS Glue Data Quality, all while keeping data in Snowflake.

    AWS Big Data Blog · 2026-09-16T15:19:51Z

  2. 10 Years of MongoDB Atlas: Built for What’s Next

    Nearly a decade ago, I joined MongoDB as a Senior Product Manager to help build the company’s new cloud product, MongoDB Atlas. Our customers had been telling us they wanted to bring MongoDB’s familiar developer experience to the cloud, with the reliability and confidence teams needed to run in production. Atlas was our answer. Today, we’re celebrating 10 years of MongoDB Atlas, the generational data platform for AI…

    MongoDB Blog · 2026-06-25T17:28:40Z

  3. Build Trust in Agentic AI: From POC to Production

    The enterprise adoption of artificial intelligence has reached an inflection point. Organizations are rapidly moving into the era of agentic AI, autonomous systems capable of executing complex reasoning and making operational decisions independently. Yet as executives attempt to transition agents from sandbox environments into mission-critical production channels, they inevitably collide with an AI trust gap. Unlike…

    MongoDB Blog · 2026-06-23T15:27:00Z

  4. Production-Ready Agents Need A Production-Ready Data Platform

    There’s a common theme to the conversations I’ve been having with AI teams lately: change. Constant, head-spinning change. Teams across industries are evaluating and re-evaluating model providers, agent frameworks, and harnesses on a continuous basis. At MongoDB, we believe that your choice of technology partner—specifically, your data platform—should simplify how you build with AI. It should deliver performance at…

    MongoDB Blog · 2026-06-11T19:46:11Z

  5. Agentic Supplier Management with MongoDB Atlas, Voyage AI, and Multi-Modal Search

    Retail supply chains are not a back-office logistics function; they are a high-stakes, board-level concern. Imagine learning suddenly that shipment rerouting surcharges have doubled due to new regional escalations; the impact on competitive differentiation and consumer trust is immediate. As a result, a long-standing focus on linear efficiency and lean inventory is being disrupted by a mandate for resilience and…

    MongoDB Blog · 2026-06-03T19:51:09Z

  6. Scale down Kinesis Data Streams on-demand capacity with ODA warm throughput

    Amazon Kinesis Data Streams now supports scaling down ingest capacity for on-demand Advantage streams with warm throughput. Learn how the scale-down works, how to monitor stream behavior with Amazon CloudWatch, and best practices for releasing excess capacity after transient traffic bursts.

    AWS Big Data Blog · 2026-09-14T15:36:36Z

  7. Fighting Tool Sprawl: The Case for AI Tool Registries

    As enterprise AI agent adoption scales, the absence of centralized, organization-level tool infrastructure is producing compounding costs. When adoption is built around optimizing for deployment speed, enterprises expose themselves to a combination of risks: duplicated engineering effort, security exposure, and operational opacity. Every enterprise needs its own shared tool registry, one that reflects its specific…

    MongoDB Blog · 2026-05-11T14:35:00Z

  8. How to migrate from Amazon CloudSearch to Amazon OpenSearch Serverless

    Learn how to migrate an Amazon CloudSearch domain to Amazon OpenSearch Serverless: assess your configuration, create a collection with explicit index mappings, convert your documents and queries to the OpenSearch query DSL, configure security, load data with Amazon OpenSearch Ingestion, and validate before cutover.

    AWS Big Data Blog · 2026-09-10T16:06:42Z

  9. AI Is Changing What Customers Need From a Database. MongoDB 8.3 Is Built for It

    Today, we announced at .local London that MongoDB 8.3 is built for the speed AI demands—and our customers can't afford to wait. The data layer has to move at AI speed The old contract between databases and the applications on top of them was simple: databases improve slowly, and architectures evolve around them. AI has changed that contract. The workloads our customers are shipping today—agents retrieving at…

    MongoDB Blog · 2026-05-07T14:33:00Z

  10. Accelerating Spark queries with Iceberg materialized views

    Accelerate slow, repetitive Apache Spark analytical queries on Apache Iceberg tables without rewriting any SQL. This post shows how automatic query rewrite in Amazon EMR and AWS Glue uses Iceberg materialized views in the AWS Glue Data Catalog to transparently substitute matching query plans, and how to design materialized views for the best speedup.

    AWS Big Data Blog · 2026-09-10T16:05:45Z

  11. New Research Reveals Overcoming Legacy Tech Issues Key to AI Success

    This guest post comes from IDC’s Dr. William Lee, Senior Research Director, Service Provider and Core Infrastructure Research. MongoDB commissioned IDC to explore the connection between legacy infrastructure, data challenges, and AI across Asia Pacific, and today we’re happy to share that work. For more, see the full MongoDB-sponsored IDC InfoBrief, Modernizing Legacy: Winning in the Age of AI, Doc #AP242555-IB…

    MongoDB Blog · 2026-04-14T17:01:07Z

  12. Every team is a data team — bring Amazon Redshift analytics to ChatGPT Work

    AWS is announcing the AWS Data Analytics plugin for the new Data agent in ChatGPT Work. Teams can ask questions in natural language, analyze governed data across their Amazon Redshift data warehouse and data lakes, and build shareable dashboards, all from a conversation in ChatGPT Work.

    AWS Big Data Blog · 2026-09-10T15:14:01Z

  13. MongoDB Predictive Auto-Scaling: An Experiment

    You can often predict a load spike before it arrives. Maybe it happens at the same time every day, or there’s always a spike at midnight on a Friday when you run a certain batch job. Or maybe it’s not cyclical, but load is rising steadily, and it’s a reasonable guess that it will keep rising for a while. MongoDB Atlas’s reactive auto-scaler handles these spikes, but scaling to the right size takes several minutes…

    MongoDB Blog · 2026-04-07T17:03:00Z

  14. Build declarative ETL pipelines with AWS Glue 6.0

    AWS Glue 6.0 introduces Spark Declarative Pipelines. In this post, you build a single declarative AWS Glue 6.0 job that turns raw order records into validated, aggregated tables through a bronze, silver, and gold sequence, without writing any orchestration logic.

    AWS Big Data Blog · 2026-09-09T17:38:22Z