📓 Download this page as a Jupyter notebook

Introduction

This series of pages shows you how UniProt’s RDF data model works and how to query it with SPARQL.

Authors: SIB Swiss-Prot group

Page Comment
00 Introduction this page
01 Basic information accession, mnemonic, dates, …
02 Protein names protein names
03 Replicon & genes replicon and gene names
04 Taxonomy organism and taxonomy
05 Sequence & isoforms sequences, isoforms, canonical sequence, processing, fragments, mass spectrometry
06 Domains & topology protein domains, zinc fingers, coiled-coils, transmembrane regions, membrane topology
07 Natural variants sequence variants, their position, evidence and literature references
08 Disease disease involvement annotations, linked disease resources, cross-references to OMIM
09 Cross References Cross References and links to other databases
10 Evidence & citations Why data was added to UniProt and how much you can trust it
11 Transporters transporter proteins, Rhea transport annotations, lipid transport, tissue-specific expression
12 Metabolism & Rhea How catalytic activity is represented and linked to Rhea
13 GO terms & keywords GO and UniProt keywords
14 Chemistry ligands, catalytic activity, cofactors, similarity search

UniProt RDF

Documentation

Documentation about the data model is available here. UniProt uses standard, community-supported vocabularies (Dublin Core, SKOS, etc.) where possible, extended by the UniProt core vocabulary.

Distribution

UniProt SPARQL endpoint

The UniProt SPARQL endpoint sparql.uniprot.org is free to use. It is updated in sync with the www.uniprot.org and FTP releases.

SPARQL is a W3C-standardized query language for the Semantic Web. If you know SQL, it will look familiar, and you can do similar kinds of queries with it. SPARQL also lets you combine data from a variety of SPARQL endpoints, providing a low-cost alternative to building your own data warehouse - you can combine UniProt data from sparql.uniprot.org with data from other SPARQL endpoints (Rhea, Bgee, OMA, OrthoDB, neXtProt, etc.).

You can also fetch a single UniProtKB entry directly in RDF/XML or Turtle format, without going through the SPARQL endpoint at all, e.g. P0A877.rdf or P0A877.ttl.

How these pages work

Every runnable example on this site is a small Turtle snippet (a tiny, self-contained excerpt of what a real UniProt entry looks like in RDF) paired with a SPARQL query. Click Run query and the query runs immediately, in your browser, against that snippet - powered by Comunica, a SPARQL engine written in JavaScript. Nothing is sent to a server. Both the data and the query are editable, so feel free to change either one and run it again. Click Visualize as graph on the data box to see it drawn out as a graph of resources and relationships. Try it all below.

Example data (Turtle) — edit it, then re-run any query below
base <http://purl.uniprot.org/uniprot/>
prefix up: <http://purl.uniprot.org/core/>
prefix rdf: <http://www.w3.org/1999/02/22-rdf-syntax-ns#>

<P0A877> a up:Protein ;
  up:mnemonic "TRPE_ECOLI" ;
  up:reviewed true .
PREFIX up: <http://purl.uniprot.org/core/>

SELECT ?protein
WHERE {
  ?protein a up:Protein .
}

The examples on the following pages build on the exact same idea, using slightly larger fixtures that mirror the shape of real UniProtKB entries. Once you’re comfortable with a pattern, try it for real against the full dataset at sparql.uniprot.org.