> For the complete documentation index, see [llms.txt](https://nks.gitbook.io/rna-seq/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://nks.gitbook.io/rna-seq/6.-mapping-alignment-and-quantification/salmon.md).

# Salmon

How did it end up to be named after a fish I have no idea but I learnt this in Japan - happy coincident

There are many commonly used mapping workflow available, and [benchmark](https://genomebiology.biomedcentral.com/articles/10.1186/s13059-016-0940-1) paper are a good starting point in selecting suitable variant.  It is painful but true that one has to try many before settling to the most suitable, but considering time, learning curve and the foreseeable continuous support to the related packages, this manual would focus on the [Salmon](https://combine-lab.github.io/salmon/getting_started/) method.  &#x20;

Salmon has 2 quantification modes.  Practically, without going into technical details, in the first mode Salmon maps the fragment (raw reads stored inside fastq) to an indexed reference genome ([quasi-mapping](https://www.biostars.org/p/331160/)) and count the hit, then move on to the next.  In the second mode a SAM/BAM alignment files were provided to Salmon and Salmon will produce the quantification from the alignment result.  One does not need to index reference genome by Salmon before running the quantification for the second mode. &#x20;

![ENST means that this is a gene annotation from Ensembl database and it indicates a transcript instead of a gene, which will be ENSG](https://hbctraining.github.io/DGE_workshop_salmon/img/quant_screenshot.png)

### 1.  Selecting a reference genome

The most common reference genome database are Ensembl, Refseq (NCBI), and UCSC.  I worked exclusively with genome curated by Ensembl so let's start from there.  Google "[Ensembl FTP](http://asia.ensembl.org/info/data/ftp/index.html/)" and you should safely land on the server within first 3 hits.  The FASTA file of cDNA of Human is what we are after.  &#x20;

![Human, cDNA, FASTA](https://4217215440-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2Fo6SLVna6kyWn4aaQaIFl%2Fuploads%2FbaACJ18yswld0gH2Db7B%2Fpic4.png?alt=media\&token=7a1e17f0-9eb0-4612-bd11-a529ae443d80)

We want `Homo_sapiens.GRCh38.cdna.all.fa.gz` and please click to have it downloaded.  This is the reference genome.

{% hint style="info" %}
The difference between DNA and cDNA, with the former containing all the sequence of a genome and the latter carrying only the coding RNA.  In case of mRNA sequencing, selecting cDNA as the reference genome is more sensible since we should have no non-coding RNA in our library.  This will significantly speed up the mapping and quantification.&#x20;
{% endhint %}

### 2.  Install and activate Salmon

Before installing Salmon, we need to install [conda](https://docs.conda.io/en/latest/miniconda.html#linux-installers) first to provide the python environment for Salmon.&#x20;

`sudo sh miniconda.sh`

```
$ conda config --add channels conda-forge
$ conda config --add channels bioconda
$ conda create -n salmon salmon
```

The last line means to create an <mark style="color:purple;">environment called salmon</mark> and install <mark style="color:orange;">package called salmon</mark> inside <mark style="color:purple;">the environment</mark>.  So every time when you want to fire up Salmon -&#x20;

```
conda activate salmon
```

The indication that you are in `conda` environment is the attachment of bracketed <mark style="color:purple;">environment name</mark> in front of your user name in the terminal, like this

{% hint style="success" %}
(salmon) <mark style="color:green;">**user\@computer**</mark> :
{% endhint %}

### 3.  Index reference genome and quantify

Then use this line to index the reference genome, `GRCh38.cDNA.fa.gz`, for mapping and quantification and store the indexed files inside `cDNA_index` -&#x20;

`salmon index -t GRCh38.cDNA.fa.gz -i cDNA_index`

{% hint style="info" %}
By now one should know that the single character after a hyphen (-), `-t` and `-i` in this case, is the parameter/argument that being passed to the command at front for additional condition/options. &#x20;
{% endhint %}

Then we can map and quantify the fastq file using Salmon.  Our example is the sequence data from a single-end library so we should use

```
salmon quant -i /home/user/FUS/cDNA_index -l A \
	 -r SRR19241828.fastq \
         -p 7 --validateMappings --gcBias -o /home/user/FUS/quants/SRR19241828
```

{% hint style="info" %}
Refer to [here](https://combine-lab.github.io/salmon/getting_started/#quantifying-samples) for the meaning of the parameter
{% endhint %}

{% hint style="danger" %}
Replace `-r` with `-1` and `-2` for paired-end read library to specific the paired .fastq file.  `-r` parameter is for single-end library. &#x20;
{% endhint %}

![This is what salmon would give you](https://4217215440-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2Fo6SLVna6kyWn4aaQaIFl%2Fuploads%2FHuWZ9L9KzugTkeI3WwpG%2Fpic5.png?alt=media\&token=a7ab045f-745b-4d06-901e-40b575db7d5b)

### 4.  Batch process

The way to loop through the whole folder and process all files in one go - First of all create the below `.sh` file.  You can do that with a `.txt` in the GUI and then `save as .sh`.  One can surely do that within terminal using their favorite word processors such as `nano`. &#x20;

```
#!/bin/bash
while read line
do
   samp=`basename ${line}`
echo "Processing sample ${samp}"
salmon quant -i /home/user/STB/FUS/cDNA_index -l A \
	 -r ${samp}.fastq \
         -p 7 --validateMappings --gcBias -o /home/user/FUS/quants/${samp}
done < file.txt
```

the `file.txt` at the end of line 9 means to input this file for the `while` loop to `read`, that means the <mark style="color:orange;">name of the fastq file</mark> that the `while loop` is reading in are from the `file.txt`, which is essentially the `SRR_Acc_List.txt` that we generated [before](/rna-seq/5.-fastq-and-quality-control/getting-fastq-files-from-online-database.md#demonstration-with-fus-neuron-rna-seq-data). &#x20;

Run the file by `bash file.sh`, and quit conda by `conda deactivate`
