Repository navigation
Sync jdk-8 with main: 3.4.3 - #1718
Conversation
## Description Release v3.0.7 ## Changes This release includes: ### Updated - Log timestamps now explicitly display timezone. - **[Breaking Change]** `PreparedStatement.setTimestamp(int, Timestamp, Calendar)` now properly applies Calendar timezone conversion using LocalDateTime pattern (inline with `getTimestamp`). Previously Calendar parameter was ineffective. - `DatabaseMetaData.getColumns()` with null catalog parameter now retrieves columns from all catalogs when using SQL Execution API, aligning the behaviour with thrift. - `DatabaseMetaData.getFunctions()` with null catalog parameter now retrieves columns from the current catalog when using SQL Execution API, aligning the behaviour with thrift. ### Fixed - Fix timeout exception handling to throw `SQLTimeoutException` instead of `DatabricksSQLException` when queries timeout. - Removes dangerous global timezone modification that caused race conditions. - Fixed `Statement.getLargeUpdateCount()` to return -1 instead of throwing Exception when there were no more results or result is not an update count. - CVE-2025-66566. Updated lz4-java dependency to 1.10.1. - Fix `INVALID_IDENTIFIER` error when using catalog/schema/table names for SQL Exec API with hyphens or special characters in metadata operations (`getSchemas()`, `getTables()`, `getColumns()`, etc.) and connection methods (`setCatalog()`, `setSchema()`). Per Databricks identifier rules, special characters are now properly enclosed in backticks. - Fix Auth_Scope handling inconsistency in Azure U2M OAuth. ## Testing Version bump and release notes have been updated across all relevant files. OVERRIDE_FREEZE=true Co-authored-by: Samikshya Chand <148681192+samikshya-db@users.noreply.github.com>
## Description - Checks in progress, freeze main till then. ## Testing <!-- Describe how the changes have been tested--> ## Additional Notes to the Reviewer NO_CHANGELOG=true
## Description Fixes multichunk test by only counting the unique urls requested that'll ensure that the test is not counting the retries ## Testing <!-- Describe how the changes have been tested--> Tested locally ## Additional Notes to the Reviewer <!-- Share any additional context or insights that may help the reviewer understand the changes better. This could include challenges faced, limitations, or compromises made during the development process. Also, mention any areas of the code that you would like the reviewer to focus on specifically. --> NO_CHANGELOG=true
## Description <!-- Provide a brief summary of the changes made and the issue they aim to address.--> Improve logging when jdbc is shaded ## Testing <!-- Describe how the changes have been tested--> Unit tests + manually in benchmarking ## Additional Notes to the Reviewer <!-- Share any additional context or insights that may help the reviewer understand the changes better. This could include challenges faced, limitations, or compromises made during the development process. Also, mention any areas of the code that you would like the reviewer to focus on specifically. --> Signed-off-by: Vikrant Puppala <vikrant.puppala@databricks.com>
## Description <!-- Provide a brief summary of the changes made and the issue they aim to address.--> Excluded circuit breaker test from SEA in the post merge workflow ## Testing <!-- Describe how the changes have been tested--> ## Additional Notes to the Reviewer <!-- Share any additional context or insights that may help the reviewer understand the changes better. This could include challenges faced, limitations, or compromises made during the development process. Also, mention any areas of the code that you would like the reviewer to focus on specifically. --> NO_CHANGELOG=true
Bumps org.apache.logging.log4j:log4j-core from 2.22.1 to 2.25.3. [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot merge` will merge this PR after your CI passes on it - `@dependabot squash and merge` will squash and merge this PR after your CI passes on it - `@dependabot cancel merge` will cancel a previously requested merge and block automerging - `@dependabot reopen` will reopen this PR if it is closed - `@dependabot close` will close this PR and stop Dependabot recreating it. You can achieve the same result by closing it manually - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) You can disable automated security fix PRs for this repo from the [Security Alerts page](https://github.com/databricks/databricks-jdbc/network/alerts). </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
## Description - We should be caching the tokens for custom oauth providers - This is to incorporate caching as mentioned in the[ internal code audit ](https://docs.google.com/document/d/1O6cmsqYw6JMYIzW6NrJR_RLZw6KQkzXNy0oKzRK7xAM/edit?tab=t.0) - There are 2 options to add tokenCache : one is re-using the persistent token cache being used in refresh flow, another is to extend cachedTokenSource (in-memory). The decision is made by benchmarking both : [internal doc](https://docs.google.com/document/d/1aO6befanIuO-OIJ4NZMx3SK3R7JXmwtV54LCpFo02EA/edit?tab=t.0). <img width="630" height="352" alt="Screenshot 2025-12-12 at 5 18 17 PM" src="https://github.com/user-attachments/assets/b394ab0f-c9bc-43b4-b00f-6beda7a1a9d2" /> ## Testing - added unit tests - Tested each of the flow end to end. ## Additional Notes to the Reviewer <!-- Share any additional context or insights that may help the reviewer understand the changes better. This could include challenges faced, limitations, or compromises made during the development process. Also, mention any areas of the code that you would like the reviewer to focus on specifically. --> --------- Signed-off-by: samikshya-chand_data <samikshya.chand@databricks.com>
…hen using Thrift protocol. (#1066) ## Description [Thrift protocol](https://github.com/databricks-eng/runtime/blob/master/sql/hive-thriftserver/if/TCLIService.thrift#L2186) has a orientation field with values FETCH_NEXT, FETCH_PRIOR or FETCH_FIRST. This field is always set to FETCH_NEXT resulting in incorrect refetch. To fetch from a particular chunk index the Thrift protocol requires the start row offset to be set. The chunk index and start row offset information is available from the expired links. Use the start row offset to fetch the links in the Thrift protocol. ## Testing This fix is tested with an integration test that validates that the correct links are fetched when fetching from a pair of chunk index and start row offset. There are also unit tests to validate correct client behaviour when unexpected responses are received from the server. ## Additional Notes to the Reviewer I also made some changes to the validation of the results. Commented within the PR. --------- Co-authored-by: tejassp-db <> Co-authored-by: Samikshya Chand <148681192+samikshya-db@users.noreply.github.com>
## Description <!-- Provide a brief summary of the changes made and the issue they aim to address.--> NO_CHANGELOG=true ## Testing <!-- Describe how the changes have been tested--> ## Additional Notes to the Reviewer <!-- Share any additional context or insights that may help the reviewer understand the changes better. This could include challenges faced, limitations, or compromises made during the development process. Also, mention any areas of the code that you would like the reviewer to focus on specifically. -->
## Description NO_CHANGELOG=true <!-- Provide a brief summary of the changes made and the issue they aim to address.--> When LogLevel.OFF was set, setupLogger() returned early without configuring the JUL logger. This caused Java's default logging behavior to kick in, resulting in deprecation warnings (e.g., ignoreTransactions warnings) being logged to console despite logging being disabled. Now properly initializes the logger with Level.OFF to suppress all output while using STDOUT to avoid file system access issues in restricted environments. Fixes #1158 ## Testing <!-- Describe how the changes have been tested--> Manual testing ## Additional Notes to the Reviewer <!-- Share any additional context or insights that may help the reviewer understand the changes better. This could include challenges faced, limitations, or compromises made during the development process. Also, mention any areas of the code that you would like the reviewer to focus on specifically. -->
Updated version v3.0.4 to be marked as deprecated and added a note to use v3.0.5 instead. Added additional details for geospatial data type support. ## Description <!-- Provide a brief summary of the changes made and the issue they aim to address.--> ## Testing <!-- Describe how the changes have been tested--> ## Additional Notes to the Reviewer <!-- Share any additional context or insights that may help the reviewer understand the changes better. This could include challenges faced, limitations, or compromises made during the development process. Also, mention any areas of the code that you would like the reviewer to focus on specifically. --> NO_CHANGELOG=true
## Description <!-- Provide a brief summary of the changes made and the issue they aim to address.--> NO_CHANGELOG=true Optimize setAutoCommit to avoid unnecessary server round-trips when the requested autoCommit value matches the cached session value. This optimization only applies when FetchAutoCommitFromServer is disabled (the default), ensuring we still respect server state when that mode is enabled. ## Testing <!-- Describe how the changes have been tested--> Unit tests ## Additional Notes to the Reviewer <!-- Share any additional context or insights that may help the reviewer understand the changes better. This could include challenges faced, limitations, or compromises made during the development process. Also, mention any areas of the code that you would like the reviewer to focus on specifically. -->
…1101) ## Description NO_CHANGELOG=true <!-- Provide a brief summary of the changes made and the issue they aim to address.--> Complete link futures for upfront-fetched chunks to prevent deadlock When chunk links are fetched upfront, the corresponding futures were never completed, causing threads to wait indefinitely. Now we complete these futures in the constructor for all pre-fetched chunks. ## Testing <!-- Describe how the changes have been tested--> - Unit tests - Manual testing ## Additional Notes to the Reviewer <!-- Share any additional context or insights that may help the reviewer understand the changes better. This could include challenges faced, limitations, or compromises made during the development process. Also, mention any areas of the code that you would like the reviewer to focus on specifically. -->
## Description <!-- Provide a brief summary of the changes made and the issue they aim to address.--> ## Testing <!-- Describe how the changes have been tested--> ## Additional Notes to the Reviewer <!-- Share any additional context or insights that may help the reviewer understand the changes better. This could include challenges faced, limitations, or compromises made during the development process. Also, mention any areas of the code that you would like the reviewer to focus on specifically. -->
## Description <!-- Provide a brief summary of the changes made and the issue they aim to address.--> Added query tags to telemetry ## Testing <!-- Describe how the changes have been tested--> Tested with real workspace in both cases: when query tags are present / not present. Behaviour is working as expected. ## Additional Notes to the Reviewer <!-- Share any additional context or insights that may help the reviewer understand the changes better. This could include challenges faced, limitations, or compromises made during the development process. Also, mention any areas of the code that you would like the reviewer to focus on specifically. --> NO_CHANGELOG=true Signed-off-by: Sreekanth Vadigi <sreekanth.vadigi@databricks.com>
## Problem Custom user agent from useragententry parameter wasn't included in connector service HTTP requests for feature flag retrieval. ## Root Cause Method execution order issue in UserAgentManager.setUserAgent() - custom user agent was set AFTER the connector service request was made. ## Solution Reordered the method to set custom user agent before calling getClientUserAgent() (which triggers feature flag fetch). ## Testing - Added testCustomUserAgentIncludedBeforeClientTypeEvaluation() test NO_CHANGELOG=true --------- Co-authored-by: Claude Sonnet 4.5 <noreply@anthropic.com>
…action preview (#1176) ## Summary - Changes default value of `IgnoreTransactions` parameter from `0` to `1`, making transactions disabled by default - Updates `supportsTransactions()` to respect the `IgnoreTransactions` flag, returning `false` when transactions are ignored (default) and `true` when explicitly enabled - Adds test case for when transactions are explicitly enabled via `IgnoreTransactions=0` ## Background The multi-statement transaction feature is currently in private preview for limited workspaces. When BI tools (Tableau, Power BI, DBeaver) detect transaction support via `supportsTransactions()`, they automatically use transaction methods, causing failures for customers not enrolled in the preview. This change prevents unexpected failures for non-preview customers while allowing preview participants to opt-in by explicitly setting `IgnoreTransactions=0` in their connection string. ## Migration Path - **Non-preview customers**: No action required - transactions are now disabled by default - **Preview participants**: Set `IgnoreTransactions=0` in connection string to enable transaction support - **GA migration**: When multi-statement transactions reach GA, flip the default back to `0` ## Test plan - [ ] Verify existing tests pass - [ ] Verify default connection returns `supportsTransactions() = false` - [ ] Verify connection with `IgnoreTransactions=0` returns `supportsTransactions() = true` - [ ] Verify transaction methods (`setAutoCommit`, `commit`, `rollback`) are no-ops by default 🤖 Generated with [Claude Code](https://claude.com/claude-code) Signed-off-by: Vikrant Puppala <vikrant.puppala@databricks.com> Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
## Description Bump version to 3.1.1
### Problem
When executing queries that return 0 rows (e.g., `WHERE 1=0`), complex
types (ARRAY, MAP, STRUCT) showed only generic type names instead of
detailed type information:
**Before:**
- `ARRAY` instead of `ARRAY<INT>`
- `MAP` instead of `MAP<STRING,STRING>`
- `STRUCT` instead of `STRUCT<field: TYPE>`
**After:**
- Detailed type information is correctly preserved for all row counts
### Root Cause
In `AbstractArrowResultChunk.java`, Arrow field metadata was only
extracted inside the `while(arrowStreamReader.loadNextBatch())` loop.
For queries with 0 rows, no batches are loaded, so the loop never
executes and metadata is never extracted.
**Code location:**
`/src/main/java/com/databricks/jdbc/api/impl/arrow/AbstractArrowResultChunk.java:338-359`
### Solution
Extract metadata from `VectorSchemaRoot` immediately after obtaining it,
**before** the `loadNextBatch()` loop.
The Arrow IPC format always sends the schema message first (before any
record batches), so field metadata is available even when there are 0
rows. `VectorSchemaRoot` contains field vectors with metadata regardless
of row count.
**Key changes:**
1. Moved metadata extraction from inside the while loop to before it
2. Added defensive null checks for `VectorSchemaRoot` and field vectors
3. Added debug logging to track metadata extraction
### Testing
#### Unit Test Coverage
- ✅ Added `testMetadataExtractionWithZeroRows()` to
`ArrowResultChunkTest`
- ✅ Verifies Arrow field metadata is extracted correctly with 0 rows
- ✅ Tests complex types: `ARRAY<INT>`, `MAP<STRING,STRING>`
- ✅ All 2,693 unit tests pass
#### Manual Verification
Tested with queries returning 0 rows:
```sql
SELECT array_col, map_col, struct_col
FROM table
WHERE 1=0
Result: Metadata now correctly shows detailed type information
Impact
- Scope: Both SQL Exec API and Thrift Server (shared code path)
- Risk: Low - backward compatible change, only affects metadata
extraction timing
- Benefits:
- Fixes schema discovery for WHERE 1=0 pattern
- Improves metadata availability for empty result sets
- Aligns with Arrow IPC specification behavior
Additional Context
- Arrow IPC specification guarantees schema is sent before record
batches
- VectorSchemaRoot.getFieldVectors() is available immediately after
ArrowStreamReader.getVectorSchemaRoot()
- No performance impact: metadata extraction is now done once upfront
instead of conditionally on first batch
---------
Co-authored-by: Claude Sonnet 4.5 <noreply@anthropic.com>
## Description
<!-- Provide a brief summary of the changes made and the issue they aim
to address.-->
This PR introduces lazy loading support for inline Arrow results to
improve memory efficiency when handling large result sets.
Previously, InlineChunkProvider would eagerly fetch all arrow batches
upfront when results had hasMoreRows = true, which could lead to memory
issues with large datasets. This change splits the handling into two
separate paths:
1. Lazy path (new): For Thrift-based inline Arrow results (when
ARROW_BASED_SET is returned), we now use LazyThriftInlineArrowResult
which fetches arrow batches on-demand as the client iterates through
rows. This is similar to how LazyThriftResult works for columnar data.
2. Remote path (existing): For URL-based Arrow results (URL_BASED_SET),
we continue using ArrowStreamResult with RemoteChunkProvider which
downloads chunks from cloud storage.
The InlineChunkProvider is now only used for SEA results with JSON_ARRAY
format and INLINE disposition (contain all data inline {no hasMoreRows
flag set}).
This will reduce memory consumption and improve performance when dealing
with large inline Arrow result sets similar to
#975.
## Testing
<!-- Describe how the changes have been tested-->
- Unit tests
- Integration tests
- Manual testing
## Additional Notes to the Reviewer
<!-- Share any additional context or insights that may help the reviewer
understand the changes better. This could include challenges faced,
limitations, or compromises made during the development process.
Also, mention any areas of the code that you would like the reviewer to
focus on specifically. -->
Bypassing an existing failure on CI/CD because of 3e4f21c
…istency (#1182) ## Summary This PR adds TIMESTAMP_NTZ normalization in the Thrift path to ensure consistent metadata behavior across both SEA and Thrift API paths. ## Background PR #1177 moved Arrow metadata extraction earlier in the processing pipeline, which exposed an inconsistency: the Thrift path started returning the correct "TIMESTAMP_NTZ" from server metadata, while the SEA path was already normalizing it to "TIMESTAMP" for backward compatibility. ## Changes - Added TIMESTAMP_NTZ → TIMESTAMP normalization in `DatabricksResultSetMetaData.java` Thrift constructor (lines 205-208) - This brings Thrift path behavior in line with existing SEA path normalization - Fixes test failure in `PreparedStatementIntegrationTests.testGetMetaData_NoResultSet` ## Testing - ✅ Local test run: `PreparedStatementIntegrationTests.testGetMetaData_NoResultSet` passes - ✅ Metadata now consistent before and after `executeQuery()` for TIMESTAMP_NTZ columns - ✅ Both SEA and Thrift paths return "TIMESTAMP" for TIMESTAMP_NTZ columns ## Related - Builds on PR #1177 (Fix Arrow field metadata not available for queries with 0 rows) - Fixes issue introduced by early metadata extraction in PR #1177 - Maintains backward compatibility with existing behavior 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Sonnet 4.5 <noreply@anthropic.com>
…oad parameter (#1183) ## Summary Add support for disabling CloudFetch via `EnableQueryResultDownload=0` connection parameter to use inline Arrow results instead. ## Changes - Add `isCloudFetchEnabled()` method to `IDatabricksConnectionContext` interface - Implement the method in `DatabricksConnectionContext` using existing `EnableQueryResultDownload` parameter - Update `DatabricksThriftServiceClient` to respect this setting when making execute requests - Add unit tests for the new functionality ## Usage To disable CloudFetch and use inline Arrow results: ``` jdbc:databricks://host:port/default;EnableQueryResultDownload=0;... ``` --------- Signed-off-by: Vikrant Puppala <vikrant.puppala@databricks.com>
…lemetry code audit comments- part1 (#1163) ## Description - 4 things that is improved with respect to telemetry : - Common object mapper across telemetry use-case (This is already thread safe and is expensive to create, i.e., good tor re-use) - Make`flushIntervalMillis` config same across both telemetry clients (un-auth and auth) - Clear connection param cache when connection is closed : this was a memory leak before - Rather than creating a scheduledExecutor for each telemetry client, we share it across a factory. ## Testing - unit tests ## Additional Notes to the Reviewer - When the LAST connection to a host is closed, all pending telemetry events for that host are flushed across all prior connections (since they all shared the same TelemetryClient). i.e., If you have 5 connections to `host-A`, closing connections 1-4 does nothing (just decrements refCount). Only when you close connection 5 (the last one) does the flush occur, sending all accumulated telemetry from all 5 connections. NO_CHANGELOG=true
## Description
This PR enhances geospatial datatype handling to include SRID (Spatial
Reference System Identifier) information in column type names and fixes
multiple issues related to complex datatype handling across different
result formats.
### Key Changes
1. **Geospatial Type Name Enhancement**
- Column type names now include SRID: `GEOMETRY(4326)` instead of
`GEOMETRY`
- Applies to both GEOMETRY and GEOGRAPHY types
- Preserves full type information in metadata for better type
identification
2. **SEA Inline Mode Complex Type Fix**
- Fixed issue where complex types (ARRAY, MAP, STRUCT) were not returned
as complex objects in SEA Inline mode (JSON array result format)
- Now properly converts to complex datatype objects when
`EnableComplexDatatypeSupport=true`
3. **Thrift CloudFetch Metadata Enhancement**
- Fixed error when extracting type details (e.g., `INT` from
`ARRAY<INT>`) in Thrift CloudFetch mode
- Enhanced `getColumnInfoFromTColumnDesc()` to use Arrow schema metadata
alongside `TColumnDesc`
- Arrow schema provides complete type information (e.g., `ARRAY<INT>`)
while `TColumnDesc` only contains base type (e.g., `ARRAY`)
4. **Arrow Metadata Extraction**
- Added `DatabricksThriftUtil.getArrowMetadata()` to deserialize Arrow
schema from `TGetResultSetMetadataResp`
- Fixed null arrow metadata issue in `DatabricksResultSet` constructor
for Thrift CloudFetch mode
## Testing
### Unit Tests
- All existing unit tests pass and additional tests are added for new
methods
### Integration Tests
- `GeospatialTests.java` - Comprehensive E2E integration test
- Tests geospatial types (GEOMETRY and GEOGRAPHY)
- Validates **24 configuration combinations**:
- Protocol: Thrift / SEA
- Serialization: Arrow / Inline
- CloudFetch: Enabled / Disabled (only with Arrow, as CloudFetch
requires Arrow)
- GeoSpatial Support: Enabled / Disabled
- Complex Type Support: Enabled / Disabled
- Validates metadata: column types, type names, class names
- Validates values: WKT representation, SRID
- Validates behavior when geospatial objects are enabled vs. disabled
(STRING fallback)
- **All 24 tests pass** ✅
## Additional Notes to the Reviewer
Other required details are mentioned in comments in the diff
---------
Signed-off-by: Sreekanth Vadigi <sreekanth.vadigi@databricks.com>
## Summary Implements proactive prefetching with a sliding window for both Thrift columnar and inline Arrow results, eliminating blocking at batch boundaries and improving throughput. ## Key Components ### New Streaming Infrastructure - **`ThriftStreamingProvider<T>`**: Generic type-safe streaming provider with background prefetch thread and configurable sliding window - **`StreamingBatch<T>`**: Type-safe batch container with lifecycle management and error handling - **`ThriftResponseProcessor<T>`**: Interface for pluggable response processors - `ColumnarResponseProcessor`: Processes Thrift columnar results - `InlineArrowResponseProcessor`: Processes inline Arrow results with schema caching ### Result Implementations - **`StreamingInlineArrowResult`**: High-throughput streaming implementation for inline Arrow results with background prefetching - **`StreamingColumnarResult`**: Streaming implementation for Thrift columnar results with prefetch ### Supporting Classes - **`ThriftBatchFetcher`** / **`ThriftBatchFetcherImpl`**: Abstraction for fetching batches from the Thrift server <img width="1792" height="1234" alt="streaming inline" src="https://github.com/user-attachments/assets/66ea9b83-a16b-42d5-9280-cb1fb81dadeb" /> ## Configuration | Parameter | Description | Default | |-----------|-------------|---------| | `EnableInlineStreaming` | Toggle streaming mode for inline results | `1` (enabled) | | `ThriftMaxBatchesInMemory` | Sliding window size (max batches kept in memory) | `3` | ## Key Features 1. **Background Prefetching**: Dedicated thread fetches batches ahead of consumption 2. **Sliding Window**: Configurable memory limit prevents unbounded memory growth 3. **Type Safety**: Generic `ThriftStreamingProvider<T>` eliminates unsafe casting 4. **Graceful Error Handling**: - Try-catch around resource cleanup to prevent cascading failures - Timeout on batch creation wait to prevent indefinite blocking 5. **Comprehensive Logging**: Debug/error logging for troubleshooting ## Testing - Updated `ExecutionResultFactoryTest` for new factory logic - Updated `DatabricksThriftServiceClientTest` for CloudFetch control - Existing integration tests cover streaming behavior ## Usage Streaming is enabled by default. To disable and use lazy loading instead: ``` jdbc:databricks://host:port/default;EnableInlineStreaming=0;... ``` To adjust the sliding window size: ``` jdbc:databricks://host:port/default;ThriftMaxBatchesInMemory=5;... ``` --------- Signed-off-by: Vikrant Puppala <vikrant.puppala@databricks.com>
…ents (#1186) ## Description Fixed `IndexOutOfBoundsException` that occurs when executing DDL statements (e.g., `CREATE DATABASE`) using the Thrift protocol. The bug manifests when there's a mismatch between the number of Thrift column descriptors and Arrow schema fields. ### Root Cause When executing DDL statements, the Databricks server behavior is: - **Thrift Protocol**: Returns column descriptors including a "Result" status column (1 column) - **Arrow Schema**: Returns an empty schema with 0 fields (no actual data) - **The Bug**: Code attempted to access `arrowMetadata[0]` without checking if the list was empty This mismatch caused `IndexOutOfBoundsException` when the driver tried to access arrow metadata at index 0 of an empty list. ### Debug Evidence **TColumnDesc (Thrift)**: ``` Column[0]: name: Result type: STRING_TYPE position: 1 Full TColumnDesc: TColumnDesc(columnName:Result, typeDesc:TTypeDesc(...), position:1, comment:) ``` **Arrow Schema**: ``` Arrow schema bytes length: 72 Deserialized Arrow schema, field count: 0 ← Empty! Arrow metadata list: size=0 ``` ### Changes Made Added bounds checking in two locations where arrow metadata is accessed: 1. **`ArrowUtil.java:247`** - Used by `StreamingInlineArrowResult` 2. **`DatabricksResultSetMetaData.java:195`** - Used for result set metadata construction **Before:** ```java String columnArrowMetadata = arrowMetadata != null ? arrowMetadata.get(columnIndex) : null; ``` **After:** ```java String columnArrowMetadata = arrowMetadata != null && columnIndex < arrowMetadata.size() ? arrowMetadata.get(columnIndex) : null; ``` ## Testing ### Manual Testing **Test Case**: Execute CREATE DATABASE statement ```java String sqlQuery = "CREATE DATABASE IF NOT EXISTS hive_metastore.test_db"; boolean hasResultSet = stmt.execute(sqlQuery); ``` **Before Fix**: `IndexOutOfBoundsException: Index 0 out of bounds for length 0` **After Fix**: Executes successfully, returns `hasResultSet=false` ## Additional Notes to the Reviewer NO_CHANGELOG=true Co-authored-by: Claude Sonnet 4.5 <noreply@anthropic.com>
#1181) ## Description <!-- Provide a brief summary of the changes made and the issue they aim to address.--> Modified the logic for enableMultipleCatalogSupport parameter to only return results when the catalog provided in metadata calls is null or is equal to the current catalog when the param is disabled. This matches the behaviour with existing driver. ## Testing <!-- Describe how the changes have been tested--> Tested locally ## Additional Notes to the Reviewer <!-- Share any additional context or insights that may help the reviewer understand the changes better. This could include challenges faced, limitations, or compromises made during the development process. Also, mention any areas of the code that you would like the reviewer to focus on specifically. -->
## Description Implements NonRowcountQueryPrefixes flag to match exiting JDBC driver behavior. This allows users to specify comma-separated query prefixes (like INSERT, UPDATE, DELETE) that should return result sets instead of row counts. Changes: - Added NON_ROWCOUNT_QUERY_PREFIXES parameter to DatabricksJdbcUrlParams - Added getNonRowcountQueryPrefixes() method to connection context interface and implementation - Updated shouldReturnResultSet() logic to check configured prefixes before SQL patterns - Added 11 comprehensive unit tests covering various scenarios - Updated NEXT_CHANGELOG.md with feature description Usage: NonRowcountQueryPrefixes=INSERT,UPDATE,DELETE,MERGE ## Testing Tests: All 75 tests pass (64 existing + 11 new) <!-- Provide a brief summary of the changes made and the issue they aim to address.--> <!-- Describe how the changes have been tested--> ## Additional Notes to the Reviewer <!-- Share any additional context or insights that may help the reviewer understand the changes better. This could include challenges faced, limitations, or compromises made during the development process. Also, mention any areas of the code that you would like the reviewer to focus on specifically. --> --------- Co-authored-by: Claude Sonnet 4.5 <noreply@anthropic.com>
… new resultset (#1187) Added support for getClientInfoProperties and getTypeInfo to return a new resultset and not return the same resultset matching the JDBC spec ## Description <!-- Provide a brief summary of the changes made and the issue they aim to address.--> Added support for getClientInfoProperties and getTypeInfo to return a new resultset and not return the same resultset matching the JDBC spec ## Testing <!-- Describe how the changes have been tested--> Added tests ## Additional Notes to the Reviewer <!-- Share any additional context or insights that may help the reviewer understand the changes better. This could include challenges faced, limitations, or compromises made during the development process. Also, mention any areas of the code that you would like the reviewer to focus on specifically. --> Fixes: #1178
## 🥞 Stacked PR Use this [link](https://github.com/databricks/databricks-jdbc/pull/1620/files) to review incremental changes. - [stack/native-batch-legacy-seam](#1620) [[Files changed](https://github.com/databricks/databricks-jdbc/pull/1620/files)] --------- ## Description Extract the existing PreparedStatement batch implementation into a dedicated legacy executor. This preserves current rewrite, interpolation, chunking, fallback, and error behavior while leaving `PreparedStatementBatchExecutor` as the coordination layer for future native batching. ## Testing - Focused batch regression suites: 85 passed - Full `jdbc-core` suite: 3,602 passed, 88 skipped - Live serverless warehouse validation: - Individual parameter-set execution - Parameterized multi-row rewrite - Interpolated multi-row rewrite with chunking - Verified update counts and inserted rows ## Additional Notes to the Reviewer This is a behavior-preserving refactor. It does not add native batching, connection properties, or transport changes. NO_CHANGELOG=true Signed-off-by: Sreekanth Vadigi <sreekanth.vadigi@databricks.com>
## Description Telemetry-visible JDBC errors need two things to remain useful downstream: a stable `DatabricksDriverErrorCode` and an explicit driver/server/user classification. Missing either can leave the telemetry dashboard and automation taxonomy out of sync. This PR adds lightweight repository guardrails for future error changes: - `CLAUDE.md` tells contributors to reuse or add an enum-backed code, verify the emitted name and numeric code in tests, and record the classification. - The pull-request template asks authors to confirm those steps or request maintainer help when they cannot access the classification system. The classification source is maintainer-owned and is not available to public-repository CI. The checks therefore remain review-based instead of adding CI that could validate only a PR checkbox, not the underlying classification. NO_CHANGELOG=true ## Testing Documentation and review-routing changes only; this PR does not change driver build or runtime behavior. - Verified the exact error-code enum path referenced by the guidance. - Verified the pull-request checklist covers both error-code and classification updates. - `git diff --check` passes. ## Telemetry Errors - [x] Not applicable — this PR does not add or change a telemetry-visible error. - [ ] Applicable — the error uses `DatabricksDriverErrorCode` where appropriate, and any new code is uniquely numbered and tested. - [ ] Applicable — its driver/server/user classification is linked, or maintainer help is requested because the author cannot access the classification. ## Additional Notes to the Reviewer This is intentionally a lightweight contributor and reviewer guardrail. It does not duplicate the maintainer-owned taxonomy in the public repository or automatically classify existing errors. --------- Signed-off-by: Prathamesh Baviskar <prathamesh.baviskar@databricks.com>
## Description - Allow a connection parameter to be supplied in both the JDBC URL and `Properties` without failing on a duplicate map key. - Preserve JDBC URL precedence when duplicate values differ. - Add regression coverage for identical and conflicting values with case-insensitive parameter names. Fixes #1648 ## Testing - `DatabricksConnectionContextTest`: 147 tests passed. - Live warehouse repro confirmed that duplicate parameters connect successfully and the JDBC URL value takes precedence. - `isaac review --uncommitted`: 0 final findings. ## Telemetry Errors - [x] Not applicable — this PR does not add or change a telemetry-visible error. - [ ] Applicable — the error uses `DatabricksDriverErrorCode` where appropriate, and any new code is uniquely numbered and tested. - [ ] Applicable — its driver/server/user classification is linked, or maintainer help is requested because the author cannot access the classification. ## Additional Notes to the Reviewer The precedence behavior is documented in the helper method and the changelog. --------- Signed-off-by: Sreekanth Vadigi <sreekanth.vadigi@databricks.com>
## 🥞 Stacked PR Use this [link](https://github.com/databricks/databricks-jdbc/pull/1621/files/520eb7f425afb66804cb4749bfb100680dc844dd..473a4fba3e88e9293d3740bb2e7a24d9c9c188c0) to review incremental changes. - [stack/native-batch-legacy-seam](#1620) [[Files changed](https://github.com/databricks/databricks-jdbc/pull/1620/files)] - [**stack/native-batch-foundation**](#1621) [[Files changed](https://github.com/databricks/databricks-jdbc/pull/1621/files/520eb7f425afb66804cb4749bfb100680dc844dd..473a4fba3e88e9293d3740bb2e7a24d9c9c188c0)] ← _this PR_ --------- ## Description Add the dormant foundation for native parameter batching. - Introduce `EnableNativeBatching`, disabled by default. - Add an immutable, ordered parameter-set model with zero-based wire ordinals. - Snapshot mutable parameter values when creating a parameter set. - Preserve empty, sparse, and incomplete sets for backend validation. This PR does not change batch execution or send native requests. ## Testing - Added tests for connection-property behavior. - Added tests for ordering, ordinals, sparse/empty sets, nulls, and mutable-value snapshots. - PreparedStatement batch regression suites passed. - Full `jdbc-core` suite: 3,611 passed, 88 skipped. ## Additional Notes to the Reviewer No telemetry field is included because that requires the corresponding backend telemetry proto change. NO_CHANGELOG=true --------- Signed-off-by: Sreekanth Vadigi <sreekanth.vadigi@databricks.com>
## Description Fixes #1630. - Validates unconditionally required connection parameters. - Returns a `DatabricksValidationException` with `INPUT_VALIDATION_ERROR` when a required parameter is missing or blank, preventing the `NullPointerException`. - Handles null JDBC URLs safely during URL validation. ## Testing - Added parser coverage for null URLs and missing, empty, and blank required parameters. - Added a public `Driver.connect()` regression test for the issue reproduction. - Verified the existing valid base-URL flow where `httpPath` is supplied through `Properties`. ## Additional Notes to the Reviewer `httpPath` is currently the only parameter required unconditionally to construct a connection context. Authentication requirements remain validated according to their existing mode-specific rules. --------- Signed-off-by: Cathleen Yan <cathleen.yan@databricks.com>
## Summary - upgrade Apache HttpClient to 5.6.3, which officially resolves both `httpcore5` and `httpcore5-h2` to patched version 5.4.3 - update Jackson, lz4-java, and shaded Netty dependencies to clear the remaining current OSV findings - document the user-visible dependency updates in `NEXT_CHANGELOG.md` Closes #1584. ## Test plan - [x] `mvn clean package -Dmaven.test.skip=true -Ddependency-check.skip=true -B` - [x] Maven dependency tree resolves `httpclient5:5.6.3`, `httpcore5:5.4.3`, and `httpcore5-h2:5.4.3` - [x] verified the uber JAR embeds those versions plus Jackson 2.18.9, lz4-java 1.11.1, and Netty 4.2.15.Final - [x] GitHub Security Scan - [x] local PR integration tests - [x] external JDBC integration tests ## Telemetry Errors - [x] Not applicable — this change does not add or change a telemetry-visible error. --------- Signed-off-by: Madhavendra Rathore <madhavendra.rathore@databricks.com> Signed-off-by: Sreekanth Vadigi <sreekanth.vadigi@databricks.com> Co-authored-by: Madhavendra Rathore <madhavendra.rathore@databricks.com>
## Summary
Adds opt-in `EnableThriftNativeMetadata` support for metadata operations
executed through SEA. The server can return Thrift-shaped metadata rows
through the Statement Execution API, while the driver preserves the same
filtering, normalization, error behavior, and JDBC metadata exposed by
the Thrift client.
Supported operations are catalogs, schemas, tables, columns, functions,
primary keys, and cross references. Procedures and procedure columns
continue to use the existing SEA path.
## Request and result flow
```text
DatabaseMetaData request
→ send X-Databricks-Metadata-Operation-Type
→ when enabled and supported, request Thrift-native metadata
→ inspect ResultManifest.is_native_metadata_result
false / absent → existing SEA SHOW-result processing
true → copy native rows and rebuild them with the existing
Thrift normalization and JDBC metadata builders
```
The manifest flag is authoritative: neither the request header nor a
native-looking schema changes how a result is processed. This keeps the
feature opt-in and preserves legacy behavior when the server does not
return a native result.
## Requests requiring special handling
- `getTables`: the native path sends catalog and types to runtime, then
reapplies the JDBC filters to returned rows. This is necessary because
runtime treats catalog as a pattern, may return temporary views outside
the requested catalog, and does not consistently handle empty or exact
table-type filters. SEA native results reuse this processing to preserve
existing JDBC-over-Thrift behavior.
| Path | Catalog behavior | `types` behavior |
| --- | --- | --- |
| Native GetTables | Applies the direct-Thrift exact filter using the
original JDBC catalog | `null` accepts every runtime-returned type, an
empty array returns no rows, and a non-empty array is matched exactly |
| SHOW TABLES | SQL is scoped to the resolved catalog | `null` uses the
driver's supported default types because SHOW has no JDBC `types`
argument |
- `getFunctions`: the runtime paths have different catalog semantics:
| Path | Search behavior | Runtime `FUNCTION_CAT` | Driver handling |
| --- | --- | --- | --- |
| Native GetFunctions | Ignores the requested catalog and searches the
session catalog | Always `""` | Replaces it with the original
`DatabaseMetaData.getFunctions` catalog, including `null` |
| SHOW FUNCTIONS | Searches the requested/resolved catalog | Returns the
function identifier's catalog | Leaves it unchanged |
The driver therefore saves the original JDBC catalog before resolving a
catalog for SQL construction and passes that original value only when
rebuilding a manifest-confirmed native result. This preserves Thrift
compatibility, but it is a column-label correction rather than catalog
filtering: native rows can come from catalog A and be labeled as catalog
B.
- `getCrossReference`: the SQL query narrows only the foreign-key side,
so the returned native rows are additionally filtered by the requested
parent catalog, schema, and table.
- Native metadata failures use Thrift-compatible propagation and timeout
codes instead of the legacy SHOW-query compatibility fallbacks.
## SQLState correction
Key-based metadata validation now reports the applicable SQLState
(`42000` or `08000`) and keeps `EXECUTE_STATEMENT_FAILED` (`1003`) as
the driver error code. Previously these errors were constructed with
`DatabricksDriverErrorCode.INVALID_STATE`, which incorrectly exposed
`INVALID_STATE` through `SQLException.getSQLState()`.
Tests: `mvn spotless:check`; focused core metadata suites (331 tests).
---------
Signed-off-by: Vu Anh Phung <vu.phung@databricks.com>
…elements. (#1659) Null elements in strongly typed nested arrays are valid elements and must be returned as such. This fixes the instance of checking by adding a null-check on both nested branches (arrays of arrays and arrays of maps). No AI used. This closes #1658. ## Testing End to end tested in `ComplexTypeQueryTests`, three new test methods. ## Telemetry Errors - [x] Not applicable — this PR does not add or change a telemetry-visible error. - [ ] Applicable — the error uses `DatabricksDriverErrorCode` where appropriate, and any new code is uniquely numbered and tested. - [ ] Applicable — its driver/server/user classification is linked, or maintainer help is requested because the author cannot access the classification. ## Additional Notes to the Reviewer Contributed as part of my work at Neo4j on our Databricks integration. No licence or usage restriction / source restriction on my contribution, also no AI used. Do whatever you want with the code. --------- Signed-off-by: Michael Simons <michael@simons.ac> Co-authored-by: Sreekanth <sreekanth.vadigi@databricks.com>
## Description Sets `UseBoundedSeaApi` and `EnableThriftNativeMetadata` defaults to `1`. When unset, both features require the existing server-side `databricks.partnerplatform.clientConfigsFeatureFlags.enableSqlExecForJdbc` flag on SQL warehouses; explicit JDBC settings continue to take precedence. ## Testing - `mvn test -pl jdbc-core -Dtest=DatabricksConnectionContextTest#testNativeMetadataViaSea*` (4 passed) - `mvn spotless:check` - Full `DatabricksConnectionContextTest` run: 157 passed; one existing Mockito test could not initialize Byte Buddy attachment on the local JDK. ## Telemetry Errors - [x] Not applicable — this PR does not add or change a telemetry-visible error. - [ ] Applicable — the error uses `DatabricksDriverErrorCode` where appropriate, and any new code is uniquely numbered and tested. - [ ] Applicable — its driver/server/user classification is linked, or maintainer help is requested because the author cannot access the classification. ## Additional Notes to the Reviewer The existing SQL Exec server flag controls the rollout only when the JDBC properties are not explicitly configured. --------- Signed-off-by: Vu Anh Phung <vu.phung@databricks.com>
## Description Fix process-wide JUL initialization when the first connection uses `LogLevel=OFF`. - `OFF` suppresses the shared parent logger without creating a handler or permanently completing initialization. - The first logging-enabled connection installs the single shared handler. - Once enabled, later connections—including `OFF` connections—do not reconfigure the logger. - Failed handler creation remains retryable. ## Testing - Unit tests cover `OFF -> TRACE`, `TRACE -> OFF`, concurrent initialization without duplicate handlers, and retry after failed handler creation. - `JulLoggerTest` and `LoggingUtilTest`: 30 tests passed. - Thin and uber JARs were manually verified: `OFF -> TRACE` creates one handler, while `TRACE -> OFF` retains the original handler and level. ## Additional Notes to the Reviewer This is the short-term fix and intentionally keeps handler creation inside `JulLogger.initLogger()`. --------- Signed-off-by: Cathleen Yan <cathleen.yan@databricks.com> Signed-off-by: Cathleen Yan <58714163+cathleeny@users.noreply.github.com>
## Description - Upgrade bundled Apache Thrift from 0.23.0 to 0.24.0. - Address CVE-2026-43871 reported by the repository security scan. - Document the dependency update in the next release changelog. ## Testing - mvn package -DskipTests -Ddependency-check.skip=true - mvn -pl jdbc-core dependency:tree -Dincludes=org.apache.thrift:libthrift -DskipTests -Ddependency-check.skip=true — resolves 0.24.0 - Broad unit-test run: 3,638 tests passed. One JDK no-open-only allocator test was selected under the generic --add-opens JVM configuration; its documented filtered invocation passes and is unrelated to Thrift. - git diff --check ## Telemetry Errors - [x] Not applicable — this PR does not add or change a telemetry-visible error. - [ ] Applicable — the error uses DatabricksDriverErrorCode where appropriate, and any new code is uniquely numbered and tested. - [ ] Applicable — its driver/server/user classification is linked, or maintainer help is requested because the author cannot access the classification. ## Additional Notes to the Reviewer No runtime code changes are included. Signed-off-by: Vikrant Puppala <vikrant.puppala@databricks.com>
## Description - Reorder the static `DatabaseMetaData.getTypeInfo()` catalogue so the `INTERVAL` row is emitted with the other `Types.VARCHAR` rows, before `BOOLEAN`, `DATE`, and `TIMESTAMP`. - Add regression coverage requiring returned `DATA_TYPE` values to be nondecreasing. - Document the user-visible metadata fix in `NEXT_CHANGELOG.md`. The catalogue is shared by the Thrift and SEA metadata clients, so this fixes both backends without changing any row contents or type mappings. This also allows databricks-driver-test#1413 to replace its known-failure tripwire with the strict JDBC ordering assertion. Fixes #1661. ## Testing - Focused metadata suites: 391 tests passed, 0 failed (`MetadataResultSetBuilderTest` and `DatabricksDatabaseMetaDataTest`). - Full `databricks-jdbc-core` suite: 3,640 tests passed, 0 failed, 88 skipped. - `mvn spotless:check`: passed across the full reactor. Signed-off-by: Cathleen Yan <58714163+cathleeny@users.noreply.github.com>
## What
Transparently auto-recover the Reyden onboarding case for JDBC: when a
**default-Thrift**
connection to a Real-Time (Reyden) SQL warehouse is rejected by the
gateway with SQLSTATE
`KP001` ("Lakehouse/RT is not supported for Thrift protocol"), the
driver re-opens the session
on its SEA path (`DatabricksSdkClient`) and remembers the warehouse so
later connections skip
Thrift — no customer configuration change. This mirrors the same feature
already in the Python
(#948), Go (#479), and Node (#523) drivers; design follows ADBC
#670.
## Changes
- **`ReydenWarehouseCache`** (new): process-wide, thread-safe cache
keyed by
`host_lowercased|warehouse_id`, ~6h TTL, lazy eviction on read plus an
opportunistic sweep on
write.
- **`DatabricksSession.open`**:
- **Pre-check** — on the default Thrift path, if the warehouse is
already known-Reyden, open
SEA directly and skip the Thrift round-trip.
- **Reactive catch** — a `KP001` on the Thrift `OpenSession` marks the
cache and (default path
only) retries once on SEA. On a double failure both error chains are
preserved (SEA error as
cause, original Thrift `KP001` as a suppressed exception).
- Detection (`getSQLState() == "KP001"`) is scoped to the `OpenSession`
`createSession` call
only — there is no shared status checker that maps `KP001` elsewhere.
## Guardrail
Recovery engages only on the **default** Thrift path
(`isDefaultThriftPath()` = `UseThriftClient`
unset). An explicit `UseThriftClient=1` is honored: `KP001` is surfaced
to the caller, never
auto-switched to SEA.
### Known edge case (documented, not fixed here)
`isDefaultThriftPath()` gates on `UseThriftClient` alone. A user who
sets a Thrift-forcing
metadata param (`UseQueryForMetadata=0` or
`TreatMetadataCatalogNameAsPattern=1`) on a Reyden
warehouse will therefore be auto-recovered to SEA, silently dropping the
native-Thrift-metadata
behavior. This is judged acceptable — the connection would otherwise
hard-fail with `KP001`, and
a Reyden warehouse cannot serve Thrift metadata at all — but flagging it
for reviewer input.
## Testing
- **Unit** (`ReydenThriftAutoRecoveryTest`, 8 tests): reactive fallback
success, cache pre-check,
explicit-Thrift guardrail (2 cases incl. cached-Reyden), double-failure
error chaining,
non-`KP001` propagation, and cache host/warehouse isolation.
`DatabricksSessionTest` (21)
passes unchanged.
- **End-to-end** against a prod Reyden warehouse:
- Forced Thrift (`UseThriftClient=1`) returns the real
`TStatus(ERROR_STATUS, sqlState:KP001)`
and the guardrail surfaces it (no fallback).
- Reactive path: connect #1 recovers `KP001` → SEA (10 rows, cache
marked); connect #2
pre-checks (skips Thrift), opens SEA directly (10 rows).
## Notes
- No new `DatabricksDriverErrorCode` is introduced; the double-failure
exception reuses the
existing `CONNECTION_ERROR` code, so no telemetry-taxonomy change is
required.
- `NEXT_CHANGELOG.md` updated under `### Added`.
This pull request and its description were written by Isaac.
---------
Signed-off-by: Rahul Singhal <rahul.singhal@databricks.com>
Co-authored-by: Isaac <no-reply@databricks.com>
## Description - Create SQL Exec API sessions with `execution_mode=FAST`. - Track the maximum server-provided `session_version` per session and include it in subsequent execute requests. - Update the tracker from create-session, the initial execute response, and synchronous polling within the originating execute call. - Keep detached async polling non-authoritative because statement results can be retrieved through another connection. - Document the Lakehouse Real-Time async session-state limitation. Create-session responses without a version remain supported. Execute requests omit `session_version` until the server supplies one. ## Testing - `mvn -q spotless:check` - `mvn -q -pl jdbc-core -Dtest=DatabricksSessionTest,DatabricksSdkClientTest,SessionVersionJsonTest test` (87 tests, with the Byte Buddy agent required by the local FIPS host) ## Telemetry Errors - [x] Not applicable — this PR does not add or change a telemetry-visible error. - [ ] Applicable — the error uses `DatabricksDriverErrorCode` where appropriate, and any new code is uniquely numbered and tested. - [ ] Applicable — its driver/server/user classification is linked, or maintainer help is requested because the author cannot access the classification. ## Additional Notes to the Reviewer The public async API supports polling a statement from another connection, while statement-result responses do not identify the originating session. The client therefore updates session state only from responses attributable to the originating session. Signed-off-by: Aakash Saravanan <aakash.saravanan@databricks.com>
## Description Bump the Databricks JDBC driver to 3.4.3 and roll the pending changelog into the release notes. ## Testing - `mvn clean package -DskipTests` - JDBC core suite: 3,659 tests passed, including all version assertions - Verified the uber JAR manifest and compiled `DriverUtil` report 3.4.3 The aggregate `mvn test` reached two unrelated thin-packaging harness errors: the live query lacks local workspace credentials, and the packaging assertion loads unshaded reactor classes instead of the shaded JAR. ## Telemetry Errors - [x] Not applicable — this PR does not add or change a telemetry-visible error. - [ ] Applicable — the error uses `DatabricksDriverErrorCode` where appropriate, and any new code is uniquely numbered and tested. - [ ] Applicable — its driver/server/user classification is linked, or maintainer help is requested because the author cannot access the classification. ## Additional Notes to the Reviewer Please review and merge after CI and release signoff; the release automation will not merge this PR. --- <!-- GITHUB_MCP_FOOTER: This attribution is automatically appended by GitHub MCP. --> _This PR was created with [GitHub MCP](http://go/mcps)._ Signed-off-by: Vu Anh Phung <vu.phung@databricks.com>
## Summary Automated remediation for findings from the weekly OSS driver security scan. Updates: - org.bouncycastle:bcprov-jdk18on@1.84 -> 1.85 The repository's Security Scan check is the authoritative validation. This PR is draft until that check and the normal driver CI pass. NO_CHANGELOG=true Source: https://github.com/databricks/databricks-driver-test/actions/runs/35779611394 Signed-off-by: peco-engineer-bot[bot] <287056288+peco-engineer-bot[bot]@users.noreply.github.com> Co-authored-by: peco-engineer-bot[bot] <287056288+peco-engineer-bot[bot]@users.noreply.github.com>
## Description Create an ES Incident when a GitHub issue is opened, then post the Jira link back to the issue. The workflow uses the OSS JDBC component and avoids duplicate tickets on reruns. NO_CHANGELOG=true ## Testing - Parsed the workflow YAML locally - Compiled the embedded Python - Ran `git diff --check` ## Telemetry Errors - [x] Not applicable — this PR does not add or change a telemetry-visible error. - [ ] Applicable — the error uses `DatabricksDriverErrorCode` where appropriate, and any new code is uniquely numbered and tested. - [ ] Applicable — its driver/server/user classification is linked, or maintainer help is requested because the author cannot access the classification. ## Additional Notes to the Reviewer Requires the `JIRA_API_TOKEN` repository secret. Jira assignee is intentionally omitted. --- <!-- GITHUB_MCP_FOOTER: This attribution is automatically appended by GitHub MCP. --> _This PR was created with [GitHub MCP](http://go/mcps)._ Signed-off-by: Vu Anh Phung <vu.phung@databricks.com>
## Description Run the Jira issue automation on the Databricks protected runner so Atlassian accepts its source IP. NO_CHANGELOG=true ## Testing - Parsed the workflow YAML locally - Ran git diff --check ## Telemetry Errors - [x] Not applicable — this PR does not add or change a telemetry-visible error. - [ ] Applicable — the error uses DatabricksDriverErrorCode where appropriate, and any new code is uniquely numbered and tested. - [ ] Applicable — its driver/server/user classification is linked, or maintainer help is requested because the author cannot access the classification. ## Additional Notes to the Reviewer Fixes the Jira 403 from GitHub Actions run 36050943676 caused by the Atlassian IP allowlist. Signed-off-by: Vu Anh Phung <vu.phung@databricks.com>
#1656) Bumps [org.apache.httpcomponents.client5:httpclient5](https://github.com/apache/httpcomponents-client) from 5.6.3 to 5.6.4. <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/apache/httpcomponents-client/blob/rel/v5.6.4/RELEASE_NOTES.txt">org.apache.httpcomponents.client5:httpclient5's changelog</a>.</em></p> <blockquote> <h2>Release 5.6.4</h2> <p>This maintenance release fixes SSL parameter application in the async TLS upgrade strategy.</p> <h2>Change Log</h2> <ul> <li> <p>BearerScheme to reject control characters in bearer token. Contributed by Javid Khan <!-- raw HTML omitted --></p> </li> <li> <p>Corrects application of SSL parameters in the async TLS upgrade method. Contributed by Oleg Kalnichevski <!-- raw HTML omitted --></p> </li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/apache/httpcomponents-client/commit/36508cead5c4388d04d20aa7efad5a68595631a0"><code>36508ce</code></a> HttpClient 5.6.4 release</li> <li><a href="https://github.com/apache/httpcomponents-client/commit/59b3d2ed71b206530e286fa86548e7833b4815eb"><code>59b3d2e</code></a> Updated release notes for HttpClient 5.6.4 release</li> <li><a href="https://github.com/apache/httpcomponents-client/commit/c0af75978b64d66542dd1e1a4bde723a96e4f77b"><code>c0af759</code></a> reject control characters in bearer token in BearerScheme</li> <li><a href="https://github.com/apache/httpcomponents-client/commit/2422b6c89aaacb8d25823b7535479a6fb5a09d0d"><code>2422b6c</code></a> Corrects application of SSL parameters in the async TLS upgrade method</li> <li><a href="https://github.com/apache/httpcomponents-client/commit/66452ea55d9477d721b3c2c1d8c42b3a934a9408"><code>66452ea</code></a> Upgraded HttpClient version to 5.6.4-SNAPSHOT</li> <li>See full diff in <a href="https://github.com/apache/httpcomponents-client/compare/rel/v5.6.3...rel/v5.6.4">compare view</a></li> </ul> </details> <br /> --------- Signed-off-by: dependabot[bot] <support@github.com> Signed-off-by: Vu Anh Phung <vu.phung@databricks.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: vuanhphung <vu.phung@databricks.com>
## Description Fix the M2M private-key integration test's stale WireMock session stub. SQL Exec API session requests now include `execution_mode=FAST`, so the exact request matcher must expect it. ## Testing Validated the fixture JSON and exact expected request body; ran `git diff --check`. The credential-dependent integration test was not run locally. ## Telemetry Errors - [x] Not applicable — this PR does not add or change a telemetry-visible error. NO_CHANGELOG=true --------- Signed-off-by: Vu Anh Phung <vu.phung@databricks.com>
## Description Fixes #1708. For SEA requests with a valid `UserAgentEntry`, place it and `Java/SQLExecHttpClient` immediately after `os/...`, including after a Thrift-to-SEA switch. The wrapper is attached only to `DatabricksSdkClient`'s API client, so OAuth, OIDC, and DBFS traffic use the shared transport unchanged. ## Testing `mvn test -pl jdbc-core -am -Dtest=DatabricksSdkClientUserAgentTest,UserAgentOrderingTest,UserAgentManagerTest,DatabricksSdkClientTest -Dsurefire.failIfNoSpecifiedTests=false` (77 passed). ## Telemetry Errors - [x] Not applicable — this PR does not add or change a telemetry-visible error. --------- Signed-off-by: Vu Anh Phung <vu.phung@databricks.com>
## Summary Automated remediation for findings from the weekly OSS driver security scan. Updates: - com.fasterxml.jackson.core:jackson-core@2.18.9 -> 2.18.11 The repository's Security Scan check is the authoritative validation. This PR is draft until that check and the normal driver CI pass. NO_CHANGELOG=true Source: https://github.com/databricks/databricks-driver-test/actions/runs/37187456062 Signed-off-by: peco-engineer-bot[bot] <287056288+peco-engineer-bot[bot]@users.noreply.github.com> Co-authored-by: peco-engineer-bot[bot] <287056288+peco-engineer-bot[bot]@users.noreply.github.com>
## Description Honor `Auth_Scope` for Databricks OAuth client-secret M2M so scoped service-principal secrets can authenticate. Client-secret M2M, JWT-assertion M2M, and U2M now parse one space-separated property value consistently. Blank values use `all-apis` for client-secret M2M and `sql offline_access` for U2M; JWT-assertion M2M omits the scope parameter. Fixes #1706. ## Testing `mvn -q test -pl jdbc-core -Dtest=ClientConfiguratorTest,DatabricksConnectionContextTest` (190 passed). `mvn -q test -pl jdbc-core -Dtest=PrivateKeyClientCredentialProviderTest,JwtPrivateKeyClientCredentialsTest,ClientConfiguratorTest` (61 passed). ## Telemetry Errors - [x] Not applicable — this PR does not add or change a telemetry-visible error. ## Additional Notes to the Reviewer A live connection with a `sql`-scoped secret was not exercised. --- <!-- GITHUB_MCP_FOOTER: This attribution is automatically appended by GitHub MCP. --> _This PR was created with [GitHub MCP](http://go/mcps)._ --------- Signed-off-by: Vu Anh Phung <vu.phung@databricks.com>
Merge main (a64ae0e, 3.4.3) into jdk-8, using the last sync point (3ef8603, synced via #1275) as the merge base, and apply the JDK 8 transformations from .claude/commands/sync-jdk8-branch.md. - Keep JDK 8 pins (Arrow 13.0.0, Mockito 4.11.0, nimbus-jose-jwt 9.47), source/target 1.8, and no spotless. - Remove OWASP dependency-check-maven (12.x needs Java 11 and breaks `mvn install` on JDK 8); the cyclonedx SBOM used by OSV is kept. - Drop fakeservice/e2e tests and all WireMock resources, Arrow patch sources/tests, and JPMS flags/profiles. - Replace Java 9+ APIs introduced on main (List/Set/Map.of, List.copyOf, String.isBlank/strip, Objects.requireNonNullElse, Optional.isEmpty, Files.writeString, JDBC 4.3 supportsSharding/beginRequest/endRequest). - coverageReport.yml: JDK 8, 80% threshold. - Remove jdk-8-only Java sources that are not on main and unreferenced. Co-authored-by: Isaac <no-reply@databricks.com> Signed-off-by: Cathleen Yan <58714163+cathleeny@users.noreply.github.com>
There was a problem hiding this comment.
Verdict: 1 Low
Sync PR bringing jdk-8 up to main@3.4.3 plus JDK 8 adaptations. Reviewability caveat: the diff is truncated (1487 hunks omitted) — the core JDK 8 source changes (Java 9+ API replacements, JDBC 4.3 method removal, Arrow patch deletion, 23 removed sources) are NOT visible and were not reviewed; this review covers only the CI/workflow/config files shown. Visible portion is consistent with the documented sync approach (coverageReport.yml correctly switches to JDK 8 / 80% threshold / spotless.skip; release workflows gated if: false). One low-severity pinned-action version-comment inconsistency noted inline. Minor nit (not blocking): actions/upload-artifact appears at two different v4 SHAs both commented # v4 across release*.yml vs securityScan.yml/runJdbcComparator.yml — harmless (both are v4.x) but worth unifying.
| steps: | ||
| - name: Generate GitHub App token (this repo) | ||
| id: app-token | ||
| uses: actions/create-github-app-token@d72941d797fd3113feb6b93fd0dec494b13a2547 # v1.12.0 |
There was a problem hiding this comment.
🔵 Low — The same pinned commit SHA d72941d797fd3113feb6b93fd0dec494b13a2547 for actions/create-github-app-token is annotated inconsistently across the synced workflows: # v1.12.0 here (and in trigger-integration-tests.yml) versus # v2.1.4 in bot-prelude/action.yml. A single immutable SHA cannot be both versions — one comment is wrong. The SHA pin itself is correct and binding, so this is cosmetic, but the mislabeled comment defeats the purpose of the version annotation and will mislead anyone later bumping the pin. Recommend normalizing all occurrences to the actual tag for this SHA.
Description
Syncs
jdk-8withmainata64ae0e0(3.4.3) and applies the JDK 8 changes from.claude/commands/sync-jdk8-branch.md. I followed the updated version of that skill frommain(#1298), which this merge also brings intojdk-8. Replaces #1717.mainsince the last sync point,3ef86030(Fix DatabaseMetaData.getColumns() returning type name in COLUMN_DEF for columns with no default #1270, synced via Sync jdk-8 with main: 3.2.2-SNAPSHOT #1275). That covers releases 3.3.1, 3.3.2, 3.3.3, 3.4.2 and 3.4.3.jdk-8/mainis still Fixed multichunk test #1150 (23a0bbdc). A plaingit merge origin/mainwould re-conflict changes that were already synced. I merged with3ef86030as the base (using a temporary graft), so only changes since the last sync were merged. This is a real merge commit: its parents are thejdk-8tip and themaintip.JDK 8 changes
Root
pom.xml1.8. The JDK 8 pins stay: Arrow13.0.0, Mockito4.11.0, nimbus-jose-jwt9.47.dependency-check-mavenis removed (plugin,pluginManagemententry and version property). Version 12.x needs Java 11, so it breaksmvn installon JDK 8. Onmain, OSV replaced OWASP as the security gate (Unify weekly + per-PR security scanning into a single workflow #1460).main: the dependency bumps, the CycloneDX aggregate-SBOM plugin (works on JDK 8), and thedependencyManagementCVE overrides for commons-lang3 and gson.jdbc-core/pom.xml--add-opens.jdk17-NioNotOpenandjdk21-NioNotOpenprofiles, the Arrow-patch JaCoCo exclusions, and the!Jvm17PlusAndArrowToNioReflectionDisabledgroup filter are removed.Workflows
coverageReport.yml: JDK 8, 80% threshold, andmvn -pl jdbc-core clean test -Dspotless.skip=true jacoco:report.main. That fixes a duplicated "Compile" step the merge produced inprCheckJDK8.yml.Java 9+ APIs in code added on
mainsince the last syncList.of,Set.of,Map.ofandList.copyOfbecame GuavaImmutableList/ImmutableSet/ImmutableMapin main code, andArrays.asListorCollections.empty*in tests.String.isBlank()/strip()becametrim().isEmpty()/trim().Objects.requireNonNullElsebecame a null-check ternary.Optional.isEmpty()became!isPresent().Files.writeStringbecameFiles.write.supportsSharding()again and dropped thebeginRequest/endRequesttest asserts.jdk-8rewrites and applying main's logic changes on top.Removed
IntegrationTestUtilandDatabricksDriverExamples(which uses it).sqlexecapi,thriftserverapiandcloudfetchapi, plus leftoversqlgatewayapi,cloudfetchsqlgatewayapi,cloudfetchthriftserverapi,jwttokenendpoint,__filesand the*fakeservicetest.propertiesfiles. Only the removed fakeservice tests used them.src/test/resources/arrow.Cleanup beyond the skill
I removed 23 Java sources that exist only on
jdk-8: they were never onmain, or were deleted from it long ago. They came back in the 2025-09 "post merge" commitde3efe43. Nothing references them:api/IDatabricksConnectionContextandapi/IDatabricksSession, stale duplicates of theapi.internalinterfacesChunkDownloadCallbackVolumeOperationProcessorDirectOAuthEndpointResolverErrorCodes,ErrorTypesDeviceInfoLogUtilTDBSql*,TExpressionInfo,TSQLVariable, etc.)Non-code leftovers (
release-notes/v1.0.x,.vscode/settings.json,runBenchmarks.yml) are unchanged.Testing
All runs were local on OpenJDK 1.8.0:
mvn clean install -DskipTests -Dspotless.skip=true(theprCheckJDK8"Compile" step): BUILD SUCCESS for all 6 modules.mvn -pl jdbc-core test -Dspotless.skip=true(theprCheckJDK8unit-test step): 3,569 tests, 0 failures, 0 errors, 0 skipped.javaponcom.databricks.client.jdbc.Driverreports 52).META-INF/versionsentries excluded).TestThinPackaging/TestUberPackaging: the offline packaging checks pass.executeLargeQueryneeds live workspace credentials, so it wasn't run.Telemetry Errors
main.DatabricksDriverErrorCodewhere appropriate, and anynew code is uniquely numbered and tested.
requested because the author cannot access the classification.
Additional Notes to the Reviewer
main(git diff origin/main...on this branch): the poms,coverageReport.yml, and the Java 8 rewrites listed above. Everything else ismain's code as already reviewed.jdk-8. I formatted the lines I touched with google-java-format 1.18.1, the versionmainuses.Synced commits (173)
NULLelements. (Fix NPE on materialising nested arrays or arrays of maps withNULLelements. #1659)NO_CHANGELOG=true
This pull request and its description were written by Isaac.