Raw API
For use cases which don’t require transformation between ClickHouse data and native or third party data types and structures, the ClickHouse Connect client provides methods for direct usage of the ClickHouse connection.Client raw_query method
The Client.raw_query method allows direct usage of the ClickHouse HTTP query interface using the client connection. The return value is an unprocessed bytes object. It offers a convenient wrapper with parameter binding, error handling, retries, and settings management using a minimal interface:
It is the caller’s responsibility to handle the resulting
bytes object. Note that the Client.query_arrow is just a thin wrapper around this method using the ClickHouse Arrow output format.
Client raw_stream method
The synchronous Client.raw_stream method has the same API as raw_query, but returns an io.IOBase stream of byte chunks. Close the stream when processing finishes. AsyncClient.raw_stream is awaited and returns an async StreamContext for use with async with and async for.
Client raw_insert method
The Client.raw_insert method allows direct inserts of bytes objects or bytes object generators using the client connection. Because it does no processing of the insert payload, it is highly performant. The method provides options to specify settings and insert format:
It is the caller’s responsibility to ensure that the
insert_block is in the specified format and uses the specified compression method. ClickHouse Connect uses these raw inserts for file uploads and PyArrow Tables, delegating parsing to the ClickHouse server.
Saving query results as files
You can stream files directly from ClickHouse to the local file system using theraw_stream method. For example, if you’d like to save the results of a query to a CSV file, you could use the following code snippet:
output.csv file with the following content:
Multithreaded, multiprocess, and async/event driven use cases
ClickHouse Connect works well in multithreaded, multiprocess, and event-loop-driven/asynchronous applications. All query and insert processing occurs within a single thread, so operations are generally thread-safe. (Parallel processing of some operations at a low level is a possible future enhancement to overcome the performance penalty of a single thread, but even in that case thread safety will be maintained.) Because each query or insert executed maintains state in its ownQueryContext or InsertContext object, respectively, these helper objects aren’t thread-safe, and they shouldn’t be shared between multiple processing streams. See the additional discussion about context objects in the QueryContexts and InsertContexts sections.
Additionally, in an application that has two or more queries and/or inserts “in flight” at the same time, there are two further considerations to keep in mind. The first is the ClickHouse “session” associated with the query/insert, and the second is the HTTP connection pool used by ClickHouse Connect Client instances.
AsyncClient
ClickHouse Connect provides a native aiohttp-based client for asyncio applications. Install the optional dependency before using it:get_async_client to create and initialize a client. I/O methods such as query, command, and insert are coroutines:
get_async_client disables automatic session IDs by default so concurrent coroutines can share a client. Pass an explicit session_id or autogenerate_session_id=True only when you need session state and will avoid concurrent queries in that session.
Managing ClickHouse session IDs
Each ClickHouse query occurs within the context of a ClickHouse “session”. Sessions are currently used for two purposes:- To associate specific ClickHouse settings with multiple queries (see the user settings). The ClickHouse
SETcommand is used to change the settings for the scope of a user session. - To track temporary tables.
Client uses a generated session ID. SET statements and temporary tables therefore persist across requests from that client. The async factory does not generate a session ID by default. ClickHouse doesn’t allow concurrent queries in the same session, and the client raises a ProgrammingError if this is attempted, so use one of the following patterns:
- Create a separate
Clientinstance for each thread/process/event handler that needs session isolation. This preserves per-client session state (temporary tables andSETvalues). - Use a unique
session_idfor each query via thesettingsargument when callingquery,command, orinsert, if you don’t require shared session state. - Disable sessions on a shared client by setting
autogenerate_session_id=Falsebefore creating the client (or pass it directly toget_client).
autogenerate_session_id=False directly to get_client(...).
In this case ClickHouse Connect doesn’t send a session_id; the server doesn’t treat separate requests as belonging to the same session. Temporary tables and session-level settings won’t persist across requests.
Customizing the HTTP connection pool
ClickHouse Connect usesurllib3 connection pools to handle the underlying HTTP connection to the server. By default, all client instances share the same connection pool, which is sufficient for the majority of use cases. This default pool maintains up to 8 HTTP Keep Alive connections to each ClickHouse server used by the application.
For large multi-threaded applications, separate connection pools may be appropriate. Customized connection pools can be provided as the pool_mgr keyword argument to the main clickhouse_connect.get_client function:
urllib3 PoolManager documentation.
The async client owns an aiohttp pool rather than using urllib3. Configure it through connector_limit, connector_limit_per_host, and keepalive_timeout on get_async_client. Calling await async_client.close_connections() rotates the pool without interrupting in-flight requests.