Skip to content

V1.6.4 - #804

Merged
flarco merged 68 commits into
mainfrom
v1.6.4
Sep 27, 2026
Merged

V1.6.4#804
flarco merged 68 commits into
mainfrom
v1.6.4

Conversation

@flarco

@flarco flarco commented Sep 19, 2026 •

Copy link
Copy Markdown
Collaborator

New Features

  • New database connectors: Firebolt, OpenSearch, DynamoDB (scan, batch write and delete_missing), LanceDB (with the DuckDB lance extension) and dBase .dbf files (read-only).
  • Databricks targets: New databricks-volume file connection for Unity Catalog Volumes. New copy_method: zerobus on databricks connections for direct Delta ingestion through Arrow Flight.
  • ADBC for ClickHouse: ClickHouse now supports ADBC ingest over HTTP. The MySQL source and the ClickHouse target can use the Arrow Lane.
  • DuckDB copy_format: copy_format (arrow | csv) replaces copy_method / DUCKDB_USE_ARROW. Arrow is now the default, and Sling falls back to CSV when the arrow extension does not load.
  • Iceberg merge strategies: Iceberg targets now support insert, update, update_insert and delete_insert. Cloudflare R2 is now a supported Iceberg storage.
  • StarRocks schema migration: StarRocks targets now support schema migration, with dialect-specific DDL (quoted defaults, inline comments, FK table property).
  • Connection tests: API spec tests are now cancellable and isolated per run. They report the request index, the iteration state and the final record shape per request.
  • Idle connection reaper: Long-running sling serve sessions close cached connections after 10 minutes of inactivity, so expired credentials get refreshed.
  • env.yaml editing: sling conns set/unset now edit env.yaml line by line and keep comments, formatting and key order. Literal secret values are allowed. Writes stop when env.yaml does not fully parse.

Bug Fixes

  • GCS / BigQuery gc_bucket: Fixed the multiple credential options provided error with Application Default Credentials.
  • MongoDB filters: Incremental and backfill filters now parse mongosh syntax (ISODate, ObjectId) and Extended JSON ($date, $oid, $numberLong, etc.).
  • SQL Server BCP: Fixed 3-part table names with the -d option. An interrupt during a BCP load no longer drops the _tmp table under bcp.
  • DuckDB sidecar: Sling now retries once when the DuckDB process stops during Parquet writes. Ctrl-C now cancels the query cleanly. Arrow output now goes through the same session, so temp tables and attached databases stay visible.
  • DuckDB type fidelity: hugeint/ubigint map to decimal, timestamp_s/ms/ns are recognized and ±infinity values decode correctly. checksum_decimal now truncates in DuckDB, DuckLake, LanceDB and MotherDuck.
  • Cloudflare D1: Query results are paged to survive memory-limit resets, and the limit option is respected.
  • MySQL ADBC URI: Special characters in credentials now decode correctly.
  • ClickHouse native connection: Fixed a hang caused by http_port in the native connection string.
  • Snowflake / S3 on EC2: Upgraded gosnowflake to v1.19.1. Sling no longer disables AWS IMDS, so EC2 instance-profile credentials work.
  • Connection URL passwords: Spaces are now encoded as %20 instead of +.
  • Environment loading: Existing process environment variables are no longer overwritten. Tab and non-breaking-space indentation in env files is repaired with a warning.
  • Incremental state: String update keys now save the correct maximum value.
  • Delete missing: Target table columns are refreshed before the delete runs.
  • Hooks: File-locked embedded databases now close correctly.
  • sling setup / assist: Fixed terminal detection when stdin is /dev/null. The error now points to --non-interactive.

flarco and others added 30 commits September 13, 2026 23:05
Incremental/backfill templates render values as mongosh syntax
(e.g. ISODate("...")) which the Go driver cannot parse. Add
normalizeFilterValue to strip shell constructors (ISODate/ObjectId),
trim quotes, and route _id fields through ObjectId conversion.

Also add parseExtendedJSON to convert single-key Extended JSON
wrappers ($date, $oid, $numberLong, $numberInt, $numberDouble,
$numberDecimal) into native BSON values, and normalize
map[any]any filters to map[string]any before processing so
wrappers are detected consistently.
Phase 1 of the connection-management UI:
- EnvFile.SetConnectionNode/DeleteConnectionNode splice only the touched
  entry in the raw node tree, so expanded ${VAR} values can never be
  written and comments/order/unmanaged keys survive.
- DeleteConnectionNode keeps a block-trailing comment when the last
  connection is removed (reattaches it to the preceding entry).
- EnvFile.ConnectionNames/EnvKeys/RawConnections/ParseEnvFileConnections
  parse without interpolation; ValidateKey/ValidateEnvKey/ExpandRef.
- EnvFileConns.Get/SetValidated with SetOptions (overwrite guard, literal
  secret refusal, atomic env: updates); PromoteLiteralSecrets and
  PreserveRefs; EnvVarNameOf sanitizes dots/dashes in env var names.
…rite path

- EnvFile.ParseEnvFileKeys / RawEnv / LookupConnectionBody / BodySha /
  ConnectionEditorEnabled; SetConnectionNode writes env updates in the same
  save; LookupConnection refactored onto a shared body-based lookup.
- EnvFileConns.SetValidated checks env overwrite against raw values;
  PromoteLiteralSecrets skips fields whose on-disk ref already expands to the
  typed value.
Shared by the master and agent EnvironmentSet guards (agent cannot import the
master package).
- Implement Databricks Volume target via Files REST API for streaming Parquet uploads to Unity Catalog Volumes
- Implement Databricks Zerobus target using Arrow Flight / Arrow IPC streaming with configurable batch_size, max_inflight_batches, and compression (none, lz4, zstd)
- Register databricks-volume and zerobus connection types and filesystem/database handlers in core dbio packages
- Add unit tests for URL parsing, REST streaming, and Arrow IPC serialization
Zerobus is no longer a standalone connection type. Direct Delta
ingestion now uses copy_method: zerobus on a type: databricks
connection, while Unity Catalog Volumes remain a file connection
(type: databricks-volume).

- Remove TypeDbZerobus URL templating and ZerobusConn dispatch
- Consolidate zerobus driver handling under the databricks connection
- Update README to reflect the new configuration and link to docs
- Document CGO_ENABLED=1 build requirement (SQLite, Zerobus)
- Add tests for seekable upload retry on HTTP 500
- Enable merge strategy tests 26-29 (insert/update/update_insert/delete_insert) for Iceberg
- Replace env var-based S3 credential hack with applyIcebergS3Props property mapping
- Add attachIcebergAwsConfig to attach AWS session config to the catalog
- Register gocloud IO for iceberg-go v0.6+ (S3/GCS/Azure blob access)
- Update dependencies (gonum, google.golang.org/api, etc.)
`sling build run` previously had no machine-readable output, so callers
(CI, editors, wrappers) had to parse the human summary to learn per-node
results. The pipeline `type: build` step already returns node results in
its state, so this reuses that shape instead of inventing a second one.

- Add buildOutputFlags to the `run` subcommand and PrintRunJSON, which
  prints RunResultsPayload, or SubProjectsPayload when the directory holds
  independent sub-projects (node names are not unique across them).
- Share RunResultsPayload between the CLI and hook_runner so
  `sling build run --json` and hooks report one contract.
- Emit the JSON payload before returning on failure: a failed run still
  hands back per-node errors alongside the non-zero exit code.
- Collect failed sub-builds and aggregate their row/byte counts, sorting
  by project dir since goroutines finish in arbitrary order.
- Route error diagnostics and the "No models selected"/summary blank line
  to stderr when --json is set, keeping stdout parseable.
- Add fixture project plus tests 613-616 covering success, per-model
  errors, `--help`, and `-R` sub-project grouping.
LanceDB is reachable through DuckDB's lance extension, so it reuses the
DuckDB driver, dialect and temp-table handling. Register the type, map
`lancedb://{path}` URLs (path is the namespace root and may itself be an
`s3://`, `gs://` or `az://` URI), and let the connection round-trip `path`.

- Exclude LanceDB from CREATE OR REPLACE TABLE: the lance extension serves
  a stale projection after a schema-changing replace, so drop first.
- Mark it index-less like DuckLake and adjust the shared DB suite for
  JSON-as-varchar; skip the five cases that depend on views, since the
  lance extension keeps views in-session only.
- Add a local, path-backed CLI suite covering full refresh, merge and CDC.
- LanceDBConn attaches the namespace root as an in-memory DuckDB catalog
  (TYPE lance), so `<table>.lance` datasets are plain tables and DuckDB's
  dialect, type mapping and merge strategies apply unchanged.
- Resolve the namespace root from `path` (alias `instance`) and clear
  `instance`, so it never becomes DuckDB's own database file; reject
  `read_only`, which neither the CLI nor the lance attach accepts.
- Register scoped lance secrets for object-store roots (s3/gs/az families,
  with credential-chain fallback), pass unknown schemes through so the
  extension reports them, and skip secrets for stores that self-resolve.
- Add the lancedb template: metadata queries pinned to `current_database()`
  so the temp-table catalog is excluded, and every merge strategy written as
  a single MERGE INTO since the extension cannot plan DELETE/UPDATE that
  read a second table. Matched rows are updated in place for delete_insert
  (same end state), and views stay session-only as the namespace has no view
  operations.
- Cover namespace-root handling, attach quoting, scheme/scope mapping and
  secret construction in database_test.go, plus full-refresh, merge and
  synthetic CDC replications.
storage.NewClient internally appends option.WithAuthCredentials for the
credentials it resolves itself. With google.golang.org/api v0.258.0 any
additional credential option sling passed (WithCredentials,
WithCredentialsFile, WithTokenSource) was counted as a second one,
failing GCS/BigQuery gc_bucket connections with "dialing: multiple
credential options provided".

Pass an oauth2-backed HTTP client instead: an authenticated HTTP client
is not treated as a credential option, and Application Default
Credentials still flow through unchanged.

Add TestFileSysGoogleADC covering a GCS client built with no explicit
credential props, which is the path BigQuery with gc_bucket takes.
- The workflow-dispatch step pushed builds for any branch, so
  non-version branches were pushed to slingdata-io/sling.
- Add a check that only proceeds for `v<digit>...` branch names,
  matching the repo's version-branch push convention.
…ket-set-multiple-credential

Fix GCS ADC credential clash by passing authenticated HTTP client
Wide tables spend most of a run building []any rows that are then
cast, merged and serialized right back. The lane carries Arrow record
batches instead: an ADBC reader feeds a RecordStream the sink pulls
from, so no row is materialized and no DuckDB merge runs.

- Gate eligibility before the read: config, connections, token and
  target columns decide, then the real reader schema is re-checked
  before the stream starts, so a stream never changes mode after
  Start. Every decline logs its reason and falls back to the row path.
- Targets: ADBC ingest, staged Parquet (Snowflake, Databricks,
  Redshift) and Parquet/Arrow files; Parquet/Arrow files also read as
  lane sources, with a shared-schema check over the file footers.
- Staged loaders and writers take records directly, keeping columns,
  metadata columns and string update keys so incremental state,
  stats and part naming match the row path.
- SLING_ARROW_LANE=false/force opt out or fail a decline; the open
  build keeps the stub engine and always takes the row path.
Add Arrow-native dataflow lane for ADBC sources and sinks
- Replace the global SpecEventChn with a typed SpecEvent sent through the
  request context, so each spec test run is isolated, cancellable, and reports
  the request index plus iteration state before and after each request. Legacy
  wire fields stay for released spec inspectors.
- Add Connection.TestWithOptions and ConnEntries.TestWithOptions (endpoints,
  limit, max requests, context, trace, OnEvent) plus a spec-file overlay that
  tests a draft spec without mutating the connection entry. Test() still reads
  the SLING_TEST_* env vars, so the CLI is unchanged.
- Emit error events and return context.Canceled promptly on cancellation
  instead of reporting success.
- Add the workbench block to env.yaml, run `sling serve workbench
  --no-browser` as a brew service, and cover the command with CLI smoke tests.
Add context-aware connection tests, spec events, and workbench env config
url.QueryEscape renders spaces as "+", but the userinfo section of a
URL is not form-encoded, so drivers received a literal "+" and auth
failed for passwords containing spaces. Apply the "+" -> "%20"
replacement in both URL slug generation (connection.go) and the
driver parse fallback (database.go) so the encoded password round-trips
correctly.
- Run the sidecar in its own process group, so a console Ctrl-C reaches only
  sling, which then cancels the query instead of reporting a duckdb failure.
- Kill registered child processes on CLI exit, since they are outside sling's
  process group and get no console signal.
- Retry once when the sidecar dies mid-query: file targets write to a temp
  location first, so re-running the task is safe, and the surviving temp
  duckdb file lets the export be copied again.
- Include the last 20 stderr lines of the sidecar in the death error, so the
  real cause reaches telemetry instead of just the exit status.
- Report a cancelled query as a sidecar death when the process exited first;
  the cancel is then a consequence, not the cause.
- Bump flarco/g to v0.2.1.
- ClickHouse driver only speaks HTTP: build the URI from http_url or
  native port (8123/8443 by secure), pass user/password as ADBC
  options, and route BulkImportStream through ADBC when enabled.
- Rebuild the MySQL URI as a mysql:// URL so special characters in
  credentials survive; the old tcp() DSN form did not decode them.
- Target each driver's hierarchy in ingest options: a MySQL schema is
  the catalog, ClickHouse has no catalog level.
- Cover with unit tests and a typed load/merge replication matrix
  (suite 624-626) for both drivers.
- Use the source primary key for StarRocks Primary Key tables instead of
  _sling_row_id as hash key; fixes Error 1064 from inline PRIMARY KEY
- Place NOT NULL before AUTO_INCREMENT/DEFAULT; drop unsupported UNIQUE
  constraint and non-BIGINT AUTO_INCREMENT
- Templates: bitmap indexes, table comments, foreign_key_constraints
  table property, quoted default literals via default_value_map
- Tolerate empty Stream Load batches (empty_load_as_error) so
  incremental runs with no new rows succeed; emit column comments when
  schema migration descriptions are enabled
- Add DDL unit tests and CLI suite cases 627-630 (postgres -> starrocks
  schema migration, delete_missing soft, empty incremental)
- `sling conns set` no longer rejects literal secret values; ${VAR}
  refs are now optional, so test 527 expects the password to be stored.
- Add EnvFile.CheckFile and refuse set/unset writes when env.yaml does
  not fully parse, since a partial-parse rewrite would silently drop
  entries it failed to decode.
- GetLocalConns warns once per path when ignoring connections from an
  invalid env file instead of silently dropping them.
- MySQL BulkExportStream reads via ADBC when the lane gate opts in, falling back to the native driver if declined
- Skip the MaxDecimals=11 cap for ClickHouse when use_adbc is set; the ADBC driver writes typed decimals
- Strip http_port from the native ClickHouse conn string; the driver sends unknown keys as server settings, which get rejected and hang
- DuckDB staging tables use the Temporary ingest option; driver v1.5+ cannot find the temp table with catalog/schema options
- Map DECIMAL32/64/256 through the arrow.DecimalType interface in ArrowSchemaToColumns
- env.yaml files edited with tabs or NBSP-style spaces (e.g. from
  macOS option-space) failed YAML parsing with cryptic errors; now
  indentation is respaced to 2/4/8-wide tab stops on load when the
  result validates, and odd spaces after a key colon become regular
  spaces.
- CheckFile additionally rejects mapping keys that start with a space
  character, an indentation mistake that still parses.
- Repair applies to all body-parsing paths (load, root node,
  connection lookup/parse) so stored Body and reloads stay consistent.
- Track indentation repairs via a new EnvFile.Repaired field, set when repairEnvYAML changes the loaded body
- Warn once per env file path for both invalid files and repaired indentation using the envFileWarned map, telling users the next sling conns set saves the fix
- Refuse repairs that change scalar values via valuesKept, and only rewrite an odd space after a plain key colon so passwords and quoted values keep their NBSPs
- Add TestRepairKeepsValues covering quoted and plain colon NBSP, leading NBSP values, and a tab inside a block scalar
- Add env.EnvFileEditor: edits splice only the lines of the target
  entry, so comments, quotes, blank lines, indentation and key order
  of untouched lines stay byte-for-byte; saves are atomic with a sha
  staleness guard, and unsafe splices (anchors in replace mode,
  multi-doc files) are refused without touching the file.
- Route conns set/unset and env writes through the editor; add
  RenameValidated, a Replace mode for full-entry writes, and template
  key order for new entries.
- Filter query history in SQL (connection, status, multi-term search)
  before count and paging via GetQueryHistoryFiltered.
- copy_format (arrow|csv) replaces copy_method / DUCKDB_USE_ARROW; arrow is
  now the default read/import format for DuckDB-class connections, falling
  back to CSV automatically when the arrow community extension can't load
- DuckDB reads report mid-stream failures instead of truncating silently:
  CSV output goes through COPY, the arrow proc reader waits at EOF, and
  early-close/cancel paths tear the process down cleanly
- Type fidelity: hugeint/ubigint map to decimal, uint64 -> decimal(20),
  timestamp_s/ms/ns recognized, +/-infinity sentinels decoded; ADBC DuckDB
  ingest widens ts[s|ms] to microseconds to dodge a driver crash
- Arrow lane engine moves behind hooks in the open build (task_run_arrow.go
  and fs_arrow.go removed); MySQL ADBC batch size raised to 100k
- StarRocks fail-safes transactions off when version detection fails (4.x
  rejects DDL in tx); replication selection errors when no stream matches
  and tolerates "(part-N)" stream suffixes
- Replace the separate Arrow IPC process: SELECTs COPY ARROWS to a FIFO in the same session (temp file on Windows), so temp tables and attached databases stay visible in arrow mode.
- Select cast columns by position (duckArrowSQL) instead of * REPLACE, removing the ambiguous-name fallback, and set preserve_insertion_order around ARROWS COPY so parallel scan batches don't interleave.
- Fall back to csv once per session with a warning when the arrow extension can't load (ImportFormat renamed to SessionFormat); binary columns skip the hex unhex when the session already uses arrow.
- Add TIME64 arrow writer support and retry Iceberg appends that fail with "context canceled" mid-upload (iceberg-go v0.6.0).
- Parallelize the arrow lane CLI suite with a locked, idempotent seed and per-case targets.
- Take all of a test's comma-separated groups atomically so multi-group tests cannot deadlock against each other
- Release the concurrency slot while waiting on groups and retry asynchronously instead of busy-waiting
- Track every attempt, including re-queued ones, with a WaitGroup so TestCLI waits for all tests before exiting
- Replace the cmap-based group tracker with a mutex-guarded map plus take/release helpers
- Stop the heartbeat ticker via a done channel and only log running tests when there are any
- Count cancellation in the WaitGroup so canceled tests that never run do not hang completion
…ts/build.sh

Suppresses duplicate library warnings emitted by the macOS linker when building sling.
The D1 connector ignored the `limit` query option, so callers
requesting a bounded preview (e.g. row-limit) still streamed the full
result set. Stop the iterator once the fetched row counter reaches the
requested limit.
Cloudflare D1 fails with error 7429 (429) or 7500 (500/504) when a
single response holds too many rows, and repeating the same query just
fails again.

- Stream select/with queries through a pager (default 5000 rows,
  overridable via the page_size prop) instead of one /raw request, so
  only one page is in memory at a time.
- On a too-large response, halve the page size down to 50 rows and
  retry the offset; non-select statements stay unpaged.
- Stop retrying rows-limit 429s in makeRequest, since they are not
  transient; retry only 502+ and unrelated 429s after reading the body.
- Expose apiURL on D1Conn so tests can point at a fake server; add
  TestD1StreamRowsPaged covering paging, exact multiples, page-size
  halving, and unpaged statements.
- Bump snowflakedb/gosnowflake from v1.17.1 to v1.19.1.
- Stop forcing AWS_EC2_METADATA_DISABLED=true in core/version.go and the S3
  connection defaults. The blanket opt-out broke EC2 instance-profile
  credentials; IMDS is now probed normally so S3 streams can authenticate
  via IAM roles, while the checksum defaults remain set to limit the
  noisy metadata warning path.
- Collect `assist` telemetry into one prop map and record a terminal outcome
  for every command: completed, error, no_tty, cancelled, agent_exit,
  not_launched.
- Add `SessionOptions.OnLaunch`, fired before the agent process starts, so
  launch, agent name, duration and exit code get reported; emit
  `assist_launch` in a goroutine so the event survives a closed terminal
  (SIGHUP).
- Detect terminals with `term.IsTerminal` instead of `ModeCharDevice`:
  `/dev/null` is a char device and agents often attach it as stdin. Setup
  forms now fail with `ErrNoTTY` (pointing at `--non-interactive`), and
  Esc/Ctrl+C maps to `ErrUserAborted`; sentinels are returned unwrapped
  because `g.Error` breaks `errors.Is`.
- Lock `env.TelMap` while `Track` copies it.
A cancel ran cleanup while bcp was still writing, so the _tmp table was
dropped under bcp and the run failed with "Invalid object name ..._tmp"
instead of reporting an interrupt.

- BcpImportFile now takes a ctx and uses exec.CommandContext, and returns
  ctx.Err() instead of retrying once the context is cancelled.
- On interrupt, wait up to 5s for the writer to stop before cleanup, then
  force it, and always mark the run as terminated.
- Add CLI test 792: SIGINT to sling during an MSSQL BCP load must be an
  interrupt and must leave no _tmp table behind.
The decline reason is already surfaced by the info-level
"falling back to row-based" line, so the per-stream debug message
duplicated it on every -d run and became the only signal the quiet
cases were tested against.

- Remove the g.Debug call in arrowLaneStub
- Update the arrow suite so quiet/decline cases assert the absence of
  the fallback line and of "arrow lane: enabled", instead of grepping
  the removed debug text
- Bump agent-wire to v0.4.0
@flarco
flarco merged commit 200dd5e into main Sep 27, 2026
1 check passed
@flarco
flarco deleted the v1.6.4 branch September 27, 2026 23:57
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants