SPARQL training
This site teaches you how to query SIB (Swiss Institute of Bioinformatics) resources with SPARQL. It also looks at RDF and some tools you can use, so it could be used as general introduction to RDF & SPARQL, however these pages will have a life science orientation.
Every example query on these pages runs directly in your browser, using the Comunica SPARQL engine loaded as JavaScript. Each example ships with a small, self-contained snippet of Turtle data - both the data and the query are editable, so you can click Run query, see real results immediately, tweak either one, and run it again. No account, no server, no installation. Every example dataset can also be drawn out as a graph with the Visualize as graph button. You can’t break anything on our servers so feel free to experiment here to your ❤️ delight.
New to SPARQL?
- RDF and linked data - start here if you’ve never seen a “triple” before: what RDF is, why it exists, and the “I ❤️ ELIXIR” example built up step by step
- RDF file formats - the same triple written four ways (Turtle, N-Triples, RDF/XML, JSON-LD), so you recognise them as the same data
- SPARQL basics - the classic introductory tutorial (a small people-and-pets dataset) covering triple patterns, property paths,
OPTIONAL/FILTER, aggregation, and federated queries - schema.org & Bioschemas - how the same RDF ideas show up as structured data embedded in ordinary web pages, and how Bioschemas applies that to life-science resources
UniProt: SPARQL and RDF
The Universal Protein Resource (UniProt) is a comprehensive resource for protein sequence and annotation data, available as RDF and queryable with SPARQL at sparql.uniprot.org.
- Introduction - what UniProt RDF and SPARQL are, and how the examples on this site work
- Basic information - accession, entry name, status, dates and versions
- Protein names - recommended, alternative and EC names
- Replicon & genes - gene names and the replicon (chromosome, plasmid, organelle) a gene sits on
- Taxonomy - organisms, taxonomic ranks, hierarchy and host organisms
- Sequence & isoforms - sequences, isoforms, canonical sequence selection, processing (initiator methionine, chains, signal peptides), fragments and mass spectrometry measurements
- Domains & topology - protein domains, zinc fingers, coiled-coils, transmembrane regions and membrane topology
- Natural variants - sequence variants, their position, evidence and longest descriptions
- Disease - disease involvement annotations, linked disease resources, and cross-references to OMIM
- Cross-references - links to PDB, UniRef, UniParc and other external databases, including federated queries
- Evidence & citations - how annotations are backed by evidence tags, protein existence levels, and citation scope
- Transporters - transporter proteins via Rhea’s transport annotations, lipid transport, and tissue-specific expression
- Metabolism & Rhea - catalytic activity, EC classification, and pathway cross-references, queried from the UniProt side
- GO terms & keywords - classifying proteins with Gene Ontology terms and UniProt keywords
- Chemistry - ligands, cofactors, PTMs, catalytic activity, and a reference example of an IDSM/Sachem chemical similarity search
Rhea
Rhea is an expert-curated resource of biochemical reactions, cross-referenced with UniProt, ChEBI, and other resources, queryable at sparql.rhea-db.org/sparql.
- Metabolism tutorial - a hands-on walk through querying metabolism data across Rhea, UniProt, ChEBI and more
- Citations & cross-references - how Rhea reactions cite PubMed literature and cross-reference KEGG, MetaCyc and other reaction databases
Tools & ecosystem
- Tools & ecosystem - a live-generated list of SPARQL/RDF/semantic-web tools, queried straight from Wikidata
Publishing your own data
- Tips and tricks: publishing your own RDF - a practical checklist for turning your own life science dataset into RDF: identifiers, reusing vocabulary, content negotiation, SHACL shapes and common pitfalls
Source
The material is developed in the open at github.com/sib-swiss/sparql-training. We used AI (specifically Claude by Anthropic) to build this site, but only to convert prior existing material into this page, adding the visualizations and comunica tools.