Skip to main content

r/bigquery


September 2026 - BigQuery Feature Summary
September 2026 - BigQuery Feature Summary

Hey everyone, here's the 09/2026 edition! (Apologies for the late post!)

🔤 GoogleSQL Language Features & Functions

🧠 AI, Machine Learning & Foundation Models

💻 Developer Experience (DX) & BigQuery Tooling

🗄️ Lakehouse Architecture, Apache Iceberg & Open Data Formats

⚡ Core Engine Performance, Indexing & Optimization

🔌 Data Integration, Pipelines & Ingestion (ELT)

🔒 Security, Governance & Workload Management

As usual, let us know if there's anything: comments, questions, concerns!


Advertisement: Skip the tutorial. Go straight to boss fights, blazing showdowns, epic combat.
Skip the tutorial. Go straight to boss fights, blazing showdowns, epic combat.
media poster


Open source conversational analytics tool for BigQuery?
Open source conversational analytics tool for BigQuery?

Hi all,

Has anyone found a decent open source tool for asking questions in plain English on top of BigQuery? Something self hosted.

I tried Gemini in BigQuery and had a look at the conversational analytics API. The API gives you back SQL and a chart spec, so you still end up building the front end yourself, which is the part I was hoping to avoid.

My main worry is the numbers. If two people ask for revenue in slightly different words I'd like them to get the same answer, and I don't see how that works unless the metrics are defined somewhere.

The other thing is cost. I have no feel for how many queries a single question turns into once the model starts looking around the schema.


Replaced a weekly full reload (MySQL → BigQuery) with direct binlog reads, no Kafka or Debezium: first production pilot, numbers inside
Replaced a weekly full reload (MySQL → BigQuery) with direct binlog reads, no Kafka or Debezium: first production pilot, numbers inside
Replaced a weekly full reload (MySQL → BigQuery) with direct binlog reads, no Kafka or Debezium: first production pilot, numbers inside

I just finished the first full pilot of rivet, an open-source Rust CLI that moves data from OLTP databases into a warehouse.

Stack: MySQL → rivet (reads the binlog directly) → Parquet → Google Cloud Storage → BigQuery. No Kafka, no Debezium, no Airflow: one VM and cron.

The setup: 154 tables, 3.2B rows, 8 refreshes a day. The existing pipeline pulls changes with a date-window query that overlaps the previous day, and fully reloads every table once a week. Without the weekly reload it would never see rows changed without an `updated_at` bump.

Results, against the existing pipeline on the same database:

- Correctness. On day one we found 880 rows that had been changed after the fact without an updated timestamp. The window pipeline would not have seen them until the weekly reload; the binlog did right away. We checked some of them against the source by hand: they really were stale balances.
- Read load on the MySQL replica: −99%. About 4 TB of reads a month (~95% of it the weekly full reloads) becomes about 24 GB of binlog stream.
- AWS → GCP egress: −98%. Only changes cross the wire: ~150 MB of compressed Parquet a day instead of a full snapshot every week.
- BigQuery queries. 25% fewer MERGE jobs and 30% fewer bytes per MERGE, because each change is merged once instead of on every run while it sits inside the window.
- Total infra cost: about −50% at the same refresh frequency (BigQuery + GCS + egress). Part of that saving ships in the next release: on a 50M-row test table it cut 80% of the bytes per compaction cycle, with an identical result. Still to be confirmed on the pilot.
- BigQuery storage: about the same (±2%).

Being honest about speed. A full cycle takes about 68 minutes today:
- reading the binlog for all 154 tables takes ~9 minutes;
- the rest is loading and merging into BigQuery, one table at a time.

The job logs show BigQuery is busy only 25–30% of that time; the rest is waiting between jobs. Newer releases process up to 16 tables in parallel. My estimate is 10–25 minutes per cycle; I'll measure after upgrading the pilot.

Scale: the client has 5–6 databases like this one. Moving all of them projects to −63…−73% in cost, mostly from egress. That still has to be confirmed against the AWS bill.

- Repo: https://github.com/panchenkoai/rivet
- Cheat sheet: https://panchenkoai.github.io/rivet/cheat-sheet.html

Questions, criticism and "why not just Debezium?" are all welcome in the comments..

upvotes comments