Setup
Let’s first create a virtual environment:ipython to run the commands in the rest of the guide, which you can launch by running:
Querying a JSON file in S3
Let’s now have a look at how to query a JSON file that’s stored in an S3 bucket. The YouTube dislikes dataset contains more than 4 billion rows of dislikes on YouTube videos up to 2021. We’re going to work with one of the JSON files from that dataset. Import chdb:Configuring the output format
The default output format isCSV, but we can change that via the output_format parameter.
chDB supports the ClickHouse data formats, as well as some of its own, including DataFrame, which returns a Pandas DataFrame:
Creating a table from JSON file
Next, let’s have a look at how to create a table in chDB. We need to use a different API to do that, so let’s first import that:dislikes table based on the schema from the JSON file, using the CREATE...EMPTY AS technique.
We’ll use the schema_inference_make_columns_nullable setting so that column types aren’t all made Nullable.
DESCRIBE clause to inspect the schema:
CREATE...AS technique.
Let’s create a different table using that technique: