219_human_gene_count_vs_isoform_count

rq turtle/ttl

Compare the number of reviewed human UniProtKB entries (one canonical accession per gene) to the number of distinct sequence resources those entries carry (the canonical sequence plus every annotated splice isoform), showing directly why searching human in UniProt returns far more sequences than there are protein-coding genes — the full picture also includes unreviewed (TrEMBL) entries on top of this. This is also due to there being more than one 'genome' source (proteome).

Use at

PREFIX up: <http://purl.uniprot.org/core/>
PREFIX taxon: <http://purl.uniprot.org/taxonomy/>

SELECT
  (COUNT(DISTINCT ?gene) AS ?reviewedGenes)
  (COUNT(DISTINCT ?isoformSequence) AS ?distinctSequences)
WHERE {
  ?protein a up:Protein ;
    up:organism taxon:9606 ;
    up:reviewed true ;
    up:sequence ?isoformSequence .
  OPTIONAL { ?protein up:encodedBy ?gene }
}
graph TD
classDef projected fill:lightgreen;
classDef literal fill:orange;
classDef iri fill:yellow;
  v5("?distinctSequences")
  v3("?gene"):::projected 
  v2("?isoformSequence"):::projected 
  v1("?protein")
  v4("?reviewedGenes")
  c6(["true^^xsd:boolean"]):::literal 
  c2(["up:Protein"]):::iri 
  c4(["taxon:9606"]):::iri 
  v1 --"a"-->  c2
  v1 --"up:organism"-->  c4
  v1 --"up:reviewed"-->  c6
  v1 --"up:sequence"-->  v2
  subgraph optional0["(optional)"]
  style optional0 fill:#bbf,stroke-dasharray: 5 5;
    v1 -."up:encodedBy".->  v3
  end
  bind2[/"count(?gene)"/]
  v3 --o bind2
  bind2 --as--o v4
  bind3[/"count(?isoformSequence)"/]
  v2 --o bind3
  bind3 --as--o v5