Introduction
This series of pages shows you how UniProt’s RDF data model works and how to query it with SPARQL.
Authors: SIB Swiss-Prot group
| Page | Comment |
|---|---|
| 00 Introduction | this page |
| 01 Basic information | accession, mnemonic, dates, … |
| 02 Protein names | protein names |
| 03 Replicon & genes | replicon and gene names |
| 04 Taxonomy | organism and taxonomy |
| 05 Sequence & isoforms | sequences, isoforms, canonical sequence, processing, fragments, mass spectrometry |
| 06 Domains & topology | protein domains, zinc fingers, coiled-coils, transmembrane regions, membrane topology |
| 07 Natural variants | sequence variants, their position, evidence and literature references |
| 08 Disease | disease involvement annotations, linked disease resources, cross-references to OMIM |
| 09 Cross References | Cross References and links to other databases |
| 10 Evidence & citations | Why data was added to UniProt and how much you can trust it |
| 11 Transporters | transporter proteins, Rhea transport annotations, lipid transport, tissue-specific expression |
| 12 Metabolism & Rhea | How catalytic activity is represented and linked to Rhea |
| 13 GO terms & keywords | GO and UniProt keywords |
| 14 Chemistry | ligands, catalytic activity, cofactors, similarity search |
UniProt RDF
Documentation
Documentation about the data model is available here. UniProt uses standard, community-supported vocabularies (Dublin Core, SKOS, etc.) where possible, extended by the UniProt core vocabulary.
Distribution
UniProt SPARQL endpoint
The UniProt SPARQL endpoint sparql.uniprot.org is free to use. It is updated in sync with the www.uniprot.org and FTP releases.
SPARQL is a W3C-standardized query language for the Semantic Web. If you know SQL, it will look familiar, and you can do similar kinds of queries with it. SPARQL also lets you combine data from a variety of SPARQL endpoints, providing a low-cost alternative to building your own data warehouse - you can combine UniProt data from sparql.uniprot.org with data from other SPARQL endpoints (Rhea, Bgee, OMA, OrthoDB, neXtProt, etc.).
You can also fetch a single UniProtKB entry directly in RDF/XML or Turtle format, without going through the SPARQL endpoint at all, e.g. P0A877.rdf or P0A877.ttl.
How these pages work
Every runnable example on this site is a small Turtle snippet (a tiny, self-contained excerpt of what a real UniProt entry looks like in RDF) paired with a SPARQL query. Click Run query and the query runs immediately, in your browser, against that snippet - powered by Comunica, a SPARQL engine written in JavaScript. Nothing is sent to a server. Both the data and the query are editable, so feel free to change either one and run it again. Click Visualize as graph on the data box to see it drawn out as a graph of resources and relationships. Try it all below.
Example data (Turtle) — edit it, then re-run any query below
base <http://purl.uniprot.org/uniprot/>
prefix up: <http://purl.uniprot.org/core/>
prefix rdf: <http://www.w3.org/1999/02/22-rdf-syntax-ns#>
<P0A877> a up:Protein ;
up:mnemonic "TRPE_ECOLI" ;
up:reviewed true .PREFIX up: <http://purl.uniprot.org/core/>
SELECT ?protein
WHERE {
?protein a up:Protein .
}The examples on the following pages build on the exact same idea, using slightly larger fixtures that mirror the shape of real UniProtKB entries. Once you’re comfortable with a pattern, try it for real against the full dataset at sparql.uniprot.org.