# Connecting pyspark to CrateDB inside Jupyter Notebooks

**URL:** https://community.cratedb.com/t/connecting-pyspark-to-cratedb-inside-jupyter-notebooks/1351
**Category:** CrateDB
**Created:** [February 1, 2023, 5:49pm UTC](https://community.cratedb.com/t/connecting-pyspark-to-cratedb-inside-jupyter-notebooks/1351 "2023-02-01T17:49:50Z")
**Posts on this page:** 6
**Page:** 1

<div class="post-metadata">

### Author: ![suchrandomstuff](https://sea2.discourse-cdn.com/flex020/user_avatar/community.cratedb.com/suchrandomstuff/32/802_2.png) [@suchrandomstuff](https://community.cratedb.com/u/suchrandomstuff)
#### Post date: [February 1, 2023, 5:49pm UTC](https://community.cratedb.com/t/connecting-pyspark-to-cratedb-inside-jupyter-notebooks/1351/1 "2023-02-01T17:49:50Z")

</div>

I am trying to retrieve data from my CrateDB database using pyspark in a jupyter notebook.

Here is my code:

> from pyspark.sql import SparkSession  
> import crate  
> import os
> 
> spark = SparkSession.builder.appName(“ConnectToCrateDB”).getOrCreate()  
> os.environ[‘PYSPARK\_SUBMIT\_ARGS’] = ‘–packages io.crate:crate-jdbc-standalone:2.6.0 pyspark-shell’
> 
> df = spark.read   
> .format(“jdbc”)   
> .option(“url”, “crate://address:4200/”)   
> .option(“dbtable”, “tablename”)   
> .option(“user”, “crate”)   
> .load()

However, it keeps giving me the following error:

> Py4JJavaError: An error occurred while calling o107.load.  
> : java.sql.SQLException: No suitable driver

Can someone please help me with the setup?

---

<div class="post-metadata">

### Author: ![hammerhead](https://sea2.discourse-cdn.com/flex020/user_avatar/community.cratedb.com/hammerhead/32/270_2.png) [@hammerhead](https://community.cratedb.com/u/hammerhead)
#### Post date: [February 1, 2023, 8:51pm UTC](https://community.cratedb.com/t/connecting-pyspark-to-cratedb-inside-jupyter-notebooks/1351/2 "2023-02-01T20:51:29Z")

</div>

Hi @suchrandomstuff,

can you please try adding an `option` call setting the driver class (`.option("driver", "io.crate.client.jdbc.CrateDriver")`?

You can also use a standard PostgreSQL JDBC driver instead, we have a Spark-based example of how to connect in this post:

> [@Introduction to Azure Databricks with CrateDB](https://community.cratedb.com/t/introduction-to-using-azure-databricks-with-cratedb/764):
>
> This is a quick intro into getting started with [Azure Databricks](https://azure.microsoft.com/en-us/services/databricks/) and [CrateDB](https://crate.io/). Setup Azure Databricks Add a new Databricks service to your Azure Subscription Once this is done use “Launch Workspace” After you are signed into Azure Databricks use the common task “New Cluster” to start a cluster for your Spark jobs execution Install the [pgjdbc](https://jdbc.postgresql.org/) library (as of time of publishing org.postgresql:postgresql:42.2.23) from Maven for your cluster [azure-databricks-server-install-library]

---

<div class="post-metadata">

### Author: ![suchrandomstuff](https://sea2.discourse-cdn.com/flex020/user_avatar/community.cratedb.com/suchrandomstuff/32/802_2.png) [@suchrandomstuff](https://community.cratedb.com/u/suchrandomstuff)
#### Post date: [February 1, 2023, 9:00pm UTC](https://community.cratedb.com/t/connecting-pyspark-to-cratedb-inside-jupyter-notebooks/1351/3 "2023-02-01T21:00:30Z")

</div>

> [@hammerhead](#):
>
> .option(“driver”, “io.crate.client.jdbc.CrateDriver”)

Hi, thanks a lot for your response.

I get the following error after adding that option:

> Py4JJavaError: An error occurred while calling o141.load.  
> : java.lang.ClassNotFoundException: io.crate.client.jdbc.CrateDriver

I am not sure how or where to add the driver within a jupyter notebook.

---

<div class="post-metadata">

### Author: ![hammerhead](https://sea2.discourse-cdn.com/flex020/user_avatar/community.cratedb.com/hammerhead/32/270_2.png) [@hammerhead](https://community.cratedb.com/u/hammerhead)
#### Post date: [February 1, 2023, 9:04pm UTC](https://community.cratedb.com/t/connecting-pyspark-to-cratedb-inside-jupyter-notebooks/1351/4 "2023-02-01T21:04:30Z")

</div>

Haven’t tested it, but maybe the solution suggested here in the reply works?

> <https://stackoverflow.com/questions/51772350/how-to-specify-driver-class-path-when-using-pyspark-within-a-jupyter-notebook/51986646#51986646>

---

<div class="post-metadata">

### Author: ![suchrandomstuff](https://sea2.discourse-cdn.com/flex020/user_avatar/community.cratedb.com/suchrandomstuff/32/802_2.png) [@suchrandomstuff](https://community.cratedb.com/u/suchrandomstuff)
#### Post date: [February 1, 2023, 9:24pm UTC](https://community.cratedb.com/t/connecting-pyspark-to-cratedb-inside-jupyter-notebooks/1351/5 "2023-02-01T21:24:06Z")

</div>

This doesn’t work either. Starting to seem impossible, lol. Been on it for 2 days now.

---

<div class="post-metadata">

### Author: ![amotl](https://sea2.discourse-cdn.com/flex020/user_avatar/community.cratedb.com/amotl/32/617_2.png) [@amotl](https://community.cratedb.com/u/amotl)
#### Post date: [February 1, 2023, 11:09pm UTC](https://community.cratedb.com/t/connecting-pyspark-to-cratedb-inside-jupyter-notebooks/1351/6 "2023-02-01T23:09:49Z")

</div>

Dear @suchrandomstuff,

thank you for writing in, and for evaluating CrateDB in the context of Jupyter Notebooks and Spark, which also sparks [sic!] my interest. I can look into further details of this topic next week.

In general, to second @hammerhead, it is recommended to use the vanilla PostgreSQL JDBC driver \[1\] with CrateDB. Also in general, when aiming to connect to the PostgreSQL-compatible interface of CrateDB, addressing it on port 4200 is probably wrong, because this is the standard port of its HTTP interface.

It will probably not improve anything on your error, because it looks like the application is not even connecting to CrateDB, but croaks when loading the driver already. Still, I wanted to make you aware of the details I’ve spotted within your original post.

Please let us know about the outcome when using the vanilla driver, where the correct driver class name is `org.postgresql.Driver`.

With kind regards,  
Andreas.

* * *

1. [https://jdbc.postgresql.org/](https://jdbc.postgresql.org/)
