Webchem provides offline access to three databases: ChEMBL, FooDB, and the EU Pesticides database. This means you download them once, and then you can query them locally without an internet connection. Reliable and lightning fast!
webchem uses the hoardr package for
managing database files. By default, hoardr will save into
a users .cache/R/webchem directory. We can set a custom
directory for databases using wc_cache$cache_path_set().
This can be useful when we work on a long running project and we want to
store database files in the project directory. The current cache path
can be retrieved using wc_cache$cache_path_get().
Databases can be downloaded using db_download_chembl(),
db_download_foodb() or db_download_eup(). Note
that ChEMBL maintainers regularly issue new releases. We can pin a
ChEMBL release for our project using
chembl_check_db_version() or download the latest version.
At the time of writing the other two databases only had one version so
their downloaders do not have a version argument yet.
ChEMBL’s main query function is chembl_query(). It will
use the webservice by default but mode = "offline" will use
the offline database instead. It is intended to be a drop-in replacement
so you can expect the structure of the offline response to be at least
very similar to the webservice response. Fields that are missing from
the offline response are listed in a warning message.
res <- chembl_query("CHEMBL2", resource = "drug", mode = "offline", version = "37")
names(res$CHEMBL2)The other two databases are implemented in offline mode only.
eup_list_entries() and foodb_list_compounds()
list available substances in the database, eup_convert()
and foodb_convert() convert between identifiers,
eup_query() and foodb_query() query the
database for information on a substance.
Look at compounds in foodb:
Convert folic acid to some other supported ID:
Let’s see some synonyms:
If we want low level access to the database, we can use
db_connect() to establish the connection and then use
e.g. dplyr to interact with the database.
db_connect() is essentially a wrapper around
DBI::dbConnect() which automatically resolves the database
path.
Let’s connect to FooDB
List tables:
Let’s look at the first few rows and columns of the “Compound” table: