Tuesday, November 13, 2012


Merge files together

Merging involves putting together the information in two (or sometimes more) files into one output file. In some scripts, a line from file1 and a line from file2 are joined together into just one line of output. In others, you would get two lines in the output. See the documentation for details.
To use a script, cut and paste the code from the light green or blue box into a terminal window, change the bold, red text as needed, and hit Enter.
See More Information for notes on using these tools.

Merge files without changing individual lines

In this set of tools, lines are copied unchanged from the input files into the output file. The scripts in this section differ based on which lines from each file are used in the output file.
Use these tools when you want to merge the same kinds of data about different things. For example, two files in the same format which contain annotations about different sets of genes. The output will contain annotations about both sets genes.

Take lines in file 1 plus lines in file 2, removing duplicates (merge_files_union)

All lines appearing in either or both input files will be printed. The lines will be printed in the order they appear in the first file, followed by the order of lines in the second file.
Note: Even if a line appears more than once in a file, or appears in both files, it will be printed only once. (Having the same first column in tabular data is not the same as being a duplicate.)
Input file(s)First file
Input file(s)Second file
Output file
perl -e ' $count=0; while (<>) { if (! ($save{$_}++)) { print $_; $count++; } } warn "\n\nRead $. lines.\nTook union and removed duplicates, yielding $count lines.\n" ' file1.txt file2.txt > union.txt
Example: Given two gene lists, get all genes found in either list. Run the above script on files file1.txt and file2.txt to get a file called union.txt:
First file
(file1.txt)
 ap23
 ap23
 CG2500
 cxb7
Second file
(file2.txt)
 cxb7
 CG12345
 CG2500
Output file
(union.txt)
 ap23
 CG2500
 CG12345
 cxb7
Screen Output
 Read 7 lines.
 Took union and removed duplicates, yielding 4 lines.

Take any line that appears in file 1 and also appears in file 2 (merge_files_intersection)

Any line that appears in the first file and also appears in the second file will be printed. The lines will be printed in the order they appear in the first file.
Note: Even if a line appears more than once, it will be printed only once. (Having the same first column in tabular data is not the same as being a duplicate line.)
Input file(s)First file
Input file(s)Second file
Output file
perl -e ' ($file1, $file2) = @ARGV; open F2, $file2 or die $!; while (<F2>) { $h2{$_}++ }; open F1, $file1 or die; $total=$.; $printed=0; while (<F1>) { $total++; if ($h2{$_}) { print $_; $h2{$_} = ""; $printed++; } } warn "\n\nRead $total lines.\nTook intersection and then removed duplicates, yielding $printed lines.\n" ' file1.txt file2.txt > intersection.txt
Example: Given two gene lists, get only genes that are found in both lists. Run the above script on files file1.txt and file2.txt to get a file called intersection.txt:
First file
(file1.txt)
 ap23
 ap23
 CG2500
 cxb7
Second file
(file2.txt)
 cxb7
 CG12345
 CG2500
Output file
(intersection.txt)
 CG2500
 cxb7
Screen Output
 Read 7 lines.
 Took intersection and then removed duplicates, yielding 2 lines.

Take lines in file 1 or file 2, but not both (merge_files_lines_not_in_both)

All lines that appear in first file or in the second file, but not in both, will be printed. The lines from the first file will be printed (in the order they appear) followed by the lines in the second file.
Note: this tool will use more memory and run more slowly if the second file is very large.
Input file(s)First file
Input file(s)Second file
Output file
perl -e ' ($file1, $file2) = @ARGV; $printed = 0; $total = 0; open F2, $file2 or die "$file2: $!\n"; @lines2 = <F2>; $total += $.; foreach (@lines2) { $h2{$_}++; } open F1, $file1 or die "$file1: $!\n"; while (<F1>) { if (exists $h2{$_}) { $h2{$_} = 0; } else { print $_; $printed++; } } $total += $.; foreach (@lines2) { if ($h2{$_}) { print $_; $printed ++; } } warn "\nRead $total lines.\nPrinted $printed lines found in $file1 or $file2, but not both\n\n"; ' xor1 xor2 > not_shared
Example: Given two gene lists, get genes that are NOT shared between the lists. Run the above script on files xor1 and xor2 to get a file called not_shared:
First file
(xor1)
 ap23
 ap23
 CG2500
 CG2500
 cxb7
Second file
(xor2)
 cxb7
 CG12345
 CG12345
 CG2500
Output file
(not_shared)
 ap23
 ap23
 CG12345
 CG12345
Screen Output
 Read 9 lines.
 Printed 4 lines found in xor1 or xor2, but not both

Take lines in file 1 but NOT in file 2 (merge_files_lines_only_in_first)

All lines that appear in first file but not in the second file will be printed. (Make sure to give the files in the correct order!) The lines will be printed in the order they appear in the (first) file.
Input file(s)Take lines in this file
Input file(s)Exclude lines in this file
Output file
perl -e ' ($file1, $file2) = @ARGV; $printed = 0; open F2, $file2; while (<F2>) { $h2{$_}++ }; $count2 = $.; open F1, $file1; while (<F1>) { if (! $h2{$_}) { print $_; $printed++; } } $count1 = $.; warn "\nRead $count1 lines from $file1 and $count2 lines from $file2.\nPrinted $printed lines found in $file1 but not in $file2\n\n" ' yes_list no_list > only_yes
Example: Given two gene lists, get genes that are only in the first list. Run the above script on files yes_list and no_list to get a file called only_yes:
First file
(yes_list)
 ap23
 ap23
 CG2500
 cxb7
Second file
(no_list)
 cxb7
 CG12345
 CG2500
Output file
(only_yes)
 ap23
 ap23
Screen Output
 Read 4 lines from yes_list and 3 lines from no_list.
 Printed 2 lines found in yes_list but not in no_list
If we give no_list as the first argument and yes_list as the second, then the result will contain (only) CG12345.

Merge lines from separate files to create single lines

In this set of tools, a line from file1 and a line from file2 are merged together to create just one line of output. The tools in this section differ based on which lines are merged together in the output file.
Use these tools when you want to merge different kinds of data about the same things. For example, one file contains disease associations about a certain set of genes, the other contains mouse orthologs for the same genes. The output file will contain disease associations and orthologs for the genes.

Join two tables based on columns sharing a value (merge_lines_based_on_shared_column)

Join tables in tab-separated files file1 and file2. For all lines where the mth column in file 1 equals the nth column in file2, print the line from file1, a tab, and the line from file2. This operation is similar to a SQL join.
If a value appears more than once in the file, then the line containing it will appear that many times in the output file. A value appearing three times merging with a value appearing twice will yield six output lines. The lines will be in the order that they appear in the first file.
$col1Column in first file
$col2Column in second file
Input file(s)First file
Input file(s)Second file
Output file
perl -e ' $col1=1; $col2=0; ($f1,$f2)=@ARGV; open(F2,$f2); while (<F2>) { s/\r?\n//; @F=split /\t/, $_; $line2{$F[$col2]} .= "$_\n" }; $count2 = $.; open(F1,$f1); while (<F1>) { s/\r?\n//; @F=split /\t/, $_; $x = $line2{$F[$col1]}; if ($x) { $num_changes = ($x =~ s/^/$_\t/gm); print $x; $merged += $num_changes } } warn "\nJoining $f1 column $col1 with $f2 column $col2\n$f1: $. lines\n$f2: $count2 lines\nMerged file: $merged lines\n"; ' ortho.tab human_func.tab > fly_func.tab
Example: Given a table of potential fly-human orthologs ortho.tab and a table of human genes' functions human_func.tab, join the tables to create a file fly_func.tab with potential functions of fly genes.
First file (ortho.tab)Second file (human_func.tab)
 Fly1   Hum11
 Fly3   Hum7
 Fly7   Hum36
 Hum7   light-sensing   22
 Hum11  overeating      4
 Hum11  oversleeping    32
 Hum17  blue eyes       X
Output file (fly_func.tab)Screen Output
 Fly1   Hum11   Hum11   overeating      4
 Fly1   Hum11   Hum11   oversleeping    32
 Fly3   Hum7    Hum7    light-sensing   22
 Joining ortho.tab column 1 with human_func.tab column 0
 ortho.tab: 3 lines
 human_func.tab: 4 lines
 Merged file: 3 lines
Example: Given a list of gene names and a tab-separated table of annotations, take only the lines where the fourth column has names from the list. Simply treat the list as a table with only one column. That is:
 perl -e '$col1=0; $col2=3; ...' gene_names.list all_annot.tab > some_annot.tab

Join two tables side by side (merge_lines_side_by_side)

Join a line from file1 and a line from file2 into a single line in the output file. Print a line from one file, a separator, and a line from the next file. (The default separator is a tab, \t.) Given two tables, this will print corresponding lines from each table next to each other, effectively joining the tables side by side. The tool will print a warning if the files have a different number of lines.
This tool can be useful if you remove a couple columns from a file, manipulate those columns with other Scriptome tools, and then want to put the changed columns back into the original file.
$separatorSeparator between lines from file1 and file2
Input file(s)First file
Input file(s)Second file
Output file
perl -e ' $separator="\t"; ($file1, $file2) = @ARGV; open (F1, $file1) or die; open (F2, $file2) or die; while (<F1>) { if (eof(F2)) { warn "WARNING: File $file2 ended early\n"; last } $line2 = <F2>; s/\r?\n//; print "$_$separator$line2" } if (! eof(F2)) { warn "WARNING: File $file1 ended early\n"; } warn "\nMerged $. lines side by side with separator $separator\nMerged files $file1 and $file2 side by side\n\n" ' annotations.tab ids.tab > annotations_and_ids.tab
Example: Your original file (abbreviated BLAST results) has long identifiers like "gi|33504569|ref|NP_878312.1|" in the second column. You want just the gi number (e.g., 33504569) and the Refseq identifier (NP_878312.1) to be in separate columns. You used other Scriptome tools to pull out the identifiers, but now they're in a separate file. Running the above script will take the original file,annotations.tab. To each line, it will add a tab (\t) and then the two columns from the identifier file, ids.tab, to yield the combined table, annotation_and_ids.tab.
First file
(annotations.tab)
 gi|33504569|ref|NP_878312.1|   Agene 6 123456
 gi|33504561|ref|NP_878311.1|   Bgene   1000    X       234567
 gi|42476237|ref|NP_571328.2|   Cgene   2500    Y       987654
Second file
(ids.tab)
 33504569       NP_878312.1
 33504561       NP_878311.1
 42476237       NP_571328.2
Output file
(annotations_and_ids.tab)
 gi|33504569|ref|NP_878312.1|   Agene 6 123456  33504569        NP_878312.1
 gi|33504561|ref|NP_878311.1|   Bgene   1000    X       234567  33504561        NP_878311.1
 gi|42476237|ref|NP_571328.2|   Cgene   2500    Y       987654  42476237        NP_571328.2
Screen Output)
 Merged 3 lines side by side with separator $separator
 Merged files annotations.tab and ids.tab side by side

Sort lines in a file

To use a script, cut and paste the code from the light green or blue box into a terminal window, change the bold, red text as needed, and hit Enter.
See More Information for notes on using these tools.

Sort a file alphabetically

See General Sorting Notes for details on sorting order.

Sort alphabetically, in ASCENDING order (sort_alpha_asc)

Sort the lines in the file in ascending alphabetical order.
Input file(s)
Output file
perl -e ' while(<>) { push @lines, $_; } warn "\nSorted $. lines in ascending alphabetical order\n\n"; print sort @lines ' unsorted > sorted_asc
Example: Sort a gene list. Run the above script on a file called unsorted to get a file called sorted:
Original file (unsorted)Output file (sorted_asc)Screen Output
 ap23   7
 ap23   30
 CG2500 3
 cxb7   2
 CG12345        9
 CG2500 3
 CG12345        9
 CG2500 3
 CG2500 3
 ap23   30
 ap23   7
 cxb7   2
  
 Sorted 6 lines in ascending alphabetical order

Sort alphabetically, in DESCENDING order (sort_alpha_desc)

Sort the lines in the file in descending alphabetical order.
Input file(s)
Output file
perl -e ' while(<>) { push @lines, $_; } warn "\nSorted $. lines in descending alphabetical order\n\n"; print sort { $b cmp $a } @lines ' unsorted > sorted_desc
Example: Sort a gene list. Run the above script on a file called unsorted to get a file called sorted_desc:
Original file (unsorted)Output file (sorted_desc)Screen Output
 ap23   7
 ap23   30
 CG2500 3
 cxb7   2
 CG12345        9
 CG2500 3
 cxb7   2
 ap23   7
 ap23   30
 CG2500 3
 CG2500 3
 CG12345        9
  
 Sorted 6 lines in descending alphabetical order

Sort alphabetically, ignoring case

Z and z will count as the same letter in these sorts. See General Sorting Notes for more details on sorting order.

Sort alphabetically, ignoring case, in ASCENDING order (sort_nocase_asc)

Sort the lines in the file in ascending alphabetical order, counting z and Z as the same.
Input file(s)
Output file
perl -e ' while(<>) { push @lines, $_; } warn "\nSorted $. lines in ascending alphabetical order, ignoring case\n\n"; print sort { lc($a) cmp lc($b) } @lines ' unsorted > sorted_nc_asc
Example: Sort a gene list. Run the above script on a file called unsorted to get a file called sorted_nc_asc:
Original file (unsorted)Output file (sorted_nc_asc)Screen Output
 ap23   7
 ap23   30
 CG2500 3
 cxb7   2
 CG12345        9
 CG2500 3
 ap23   30
 ap23   7
 CG12345        9
 CG2500 3
 CG2500 3
 cxb7   2
  
 Sorted 6 lines in ascending alphabetical order, ignoring case

Sort alphabetically, ignoring case, in DESCENDING order (sort_nocase_desc)

Sort the lines in the file in descending alphabetical order, counting z and Z as the same.
Input file(s)
Output file
perl -e ' while(<>) { push @lines, $_; } warn "\nSorted $. lines in descending alphabetical order, ignoring case\n\n"; print sort { lc($b) cmp lc($a) } @lines ' unsorted > sorted_nc_desc
Example: Sort a gene list. Run the above script on a file called unsorted to get a file called sorted_nc_desc:
Original file (unsorted)Output file (sorted_nc_desc)Screen Output
 ap23   7
 ap23   30
 CG2500 3
 cxb7   2
 CG12345        9
 CG2500 3
 cxb7   2
 CG2500 3
 CG2500 3
 CG12345        9
 ap23   7
 ap23   30
  
 Sorted 6 lines in descending alphabetical order, ignoring case

Sort numerically

See General Sorting Notes for more details on sorting order.

Sort numerically, in ASCENDING order (sort_num_asc)

Sort a simple list numerically.
Input file(s)
Output file
perl -e ' while(<>) { push @lines, $_; } warn "\nSorted $. lines in ascending numerical order\n\n"; print sort { $a <=> $b } @lines ' unsorted > sorted_num_asc
Example: Sort a list of numbers. Run the above script on a file called unsorted to get a file called sorted_num_asc:
Original file (unsorted)Output file (sorted_num_asc)Screen Output
 3
 12
 78.5
 -10
 1e5
 1e-3
 -10
 1e-3
 3
 12
 78.5
 1e5
  
 Sorted 6 lines in ascending numerical order

Sort numerically, in DESCENDING order (sort_num_desc)

Input file(s)
Output file
perl -e ' while(<>) { push @lines, $_; } warn "\nSorted $. lines in ascending numerical order\n\n"; print sort { $b <=> $a } @lines ' unsorted > sorted_num_desc
Example: Sort a list of numbers. Run the above script on a file called unsorted to get a file called sorted_num_desc:
Original file (unsorted)Output file (sorted_num_desc)Screen Output
 3
 12
 78.5
 -10
 1e5
 1e-3
 1e5
 78.5
 12
 3
 1e-3
 -10
 
 Sorted 6 lines in ascending numerical order

Sort tabular data, based on the values in a given column

Sort numerically (ascending order) based on a given column (sort_on_column_num)

Sort lines in a tab-separated file in ascending order, based on the numerical value in a given column.
$columnNumerical column to sort by
Input file(s)
Output file
perl -e ' $column=1; while(<>) { s/\r?\n//; @F=split /\t/, $_; push @sort_col, $F[$column]; push @lines, "$_\n"; } warn "\nSorted $. lines in ascending order, based on numerical values in column $column\n\n"; print @lines[sort { $sort_col[$a] <=> $sort_col[$b] } 0..$#sort_col] ' unsorted > sorted_col
Example: Sort a gene list based on the score in the first column. Run the above script on a file called unsorted to get a file called sorted_col:
Original file (unsorted)Output file (sorted_col)Screen Output
 ap23   7
 ap23   30
 CG2500 3
 cxb7   2
 CG12345        9
 CG2500 3
 cxb7   2
 CG2500 3
 CG2500 3
 ap23   7
 CG12345        9
 ap23   30
  
 Sorted 6 lines in ascending order, based on numerical values in column 1

Sort numerically (descending order) based on a given column (sort_on_column_num_desc)

Sort lines in a tab-separated file in descending order, based on the numerical values in a given column.
$columnNumerical column to sort by
Input file(s)
Output file
perl -e ' $column=1; while(<>) { s/\r?\n//; @F=split /\t/, $_; push @sort_col, $F[$column]; push @lines, "$_\n"; } warn "\nSorted $. lines in descending order, based on numerical values in column 1\n\n"; print @lines[sort { $sort_col[$b] <=> $sort_col[$a] } 0..$#sort_col] ' unsorted > sorted_col
Example: Sort a gene list based on the score in the second, tab-separated column. Run the above script on a file called unsorted to get a file called sorted_col:
Original file (unsorted)Output file (sorted_col)Screen Output
 ap23   7
 ap23   30
 CG2500 3
 cxb7   2
 CG12345        9
 CG2500 3
 ap23   30
 CG12345        9
 ap23   7
 CG2500 3
 CG2500 3
 cxb7   2
  
 Sorted 6 lines in descending order, based on numerical values in column 1

Sort alphabetically (ascending order) based on a given column (sort_on_column_alpha)

Sort lines in a tab-separated file in ascending order, based on the text strings in a given column.
$columnText column to sort by
Input file(s)
Output file
perl -e ' $column=1; while(<>) { s/\r?\n//; @F=split /\t/, $_; push @sort_col, $F[$column]; push @lines, "$_\n"; } warn "\nSorted $. lines in ascending order, based on text strings in column 1\n\n"; print @lines[sort { $sort_col[$a] cmp $sort_col[$b] } 0..$#sort_col] ' unsorted > sorted_col
Example: Sort a gene list based on the gene name in the second, tab-separated column. Run the above script on a file called unsorted to get a file called sorted_col:
Original file (unsorted)Output file (sorted_col)Screen Output
 1      ap23    7
 2      ap23    30
 3      CG2500  3
 4      cxb7    2
 5      CG12345 9
 6      CG2500  3
 5      CG12345 9
 3      CG2500  3
 6      CG2500  3
 1      ap23    7
 2      ap23    30
 4      cxb7    2
  
 Sorted 6 lines in ascending order, based on text strings in column 1

Sort alphabetically (descending order) based on a given column (sort_on_column_alpha_desc)

Sort lines in a tab-separated file in descending order, based on the text strings in a given column.
$columnText column to sort by
Input file(s)
Output file
perl -e ' $column=1; ; while(<>) { s/\r?\n//; @F=split /\t/, $_; push @sort_col, $F[$column]; push @lines, "$_\n"; } warn "\nSorted $. lines in descending order, based on text strings in column 1\n\n"; print @lines[sort { $sort_col[$b] cmp $sort_col[$a] } 0..$#sort_col] ' unsorted > sorted_col
Example: Sort a gene list based on the gene name in the second, tab-separated column. Run the above script on a file called unsorted to get a file called sorted_col:
Original file (unsorted)Output file (sorted_col)Screen Output
 1      ap23    7
 2      ap23    30
 3      CG2500  3
 4      cxb7    2
 5      CG12345 9
 6      CG2500  3
 4      cxb7    2
 1      ap23    7
 2      ap23    30
 3      CG2500  3
 6      CG2500  3
 5      CG12345 9

Fetch files or sequences

The tools in this section get things, like files or sequences - e.g., from the Web.
To use a script, cut and paste the code from the light green or blue box into a terminal window, change the bold, red text as needed, and hit Enter.
See More Information for notes on using these tools.

Fetch a sequence from a popular Internet database

Fetch a sequence from a popular Internet database (fetch_sequence_web)

Gets a sequence with a given id from a given database. (The database must be one of: swiss, genbank, genpept, embl, refseq.) The format of the fetched sequence (fasta by default) can be embl, fasta, gcg, genbank, swiss, or a whole bunch of other formats: see
The Bioperl SeqIO HOWTO
for details.
This script requires Bioperl to be installed (on whichever machine the script runs on). Many biology computers will have it installed. If the script breaks because it "can't locate Bio/Perl.pm", you can download Bioperl from bioperl.org.
$databaseDatabase name
$idIdentifier
$formatFormat to write sequence in
Output file
perl -MBio::Perl -e ' $database="embl"; $id="AI129902"; $format="fasta"; $sequence = get_sequence($database, $id); write_sequence(">-", $format, $sequence); warn "Wrote $database sequence $id in $format format\n"; ' > seq.fasta
Example: Get the ROA1_HUMAN sequence from Swiss-Prot in FASTA format, and put it in seq.fasta by running the above script.
Output file (seq.fasta)Screen Output
 >AI129902; qc41b07.x1 Soares_pregnant_uterus_NbHPU Homo sapiens cDNA [etc.]
 CTCCGCGCCAACTCCCCCCACCCCCCCCCCACACCCC
 Wrote embl sequence AI129902 in fasta format

Get a file from the Web

Fetch a file from the web (fetch_file_web)

Given an http or ftp address, get a file and store it in a given filename. This assumes you have an Internet connection, the file exists, etc. If something breaks, it should print an error message.
$web_fileWeb address
$storeName of file to save to
perl -MLWP::Simple -e ' $web_file="ftp://ftp.ncbi.nih.gov/genbank/GB_Release_Number"; $store="GB.txt"; if (is_success(getstore($web_file, $store))) { warn "Downloaded $web_file into $store\n"; } else { warn "Error downloading $web_file\n" } '
Example: Run the above script to download the current GenBank release number to a file GB.txt. The resulting file will have in it one line, giving the release number (as of this writing, 151).
Example 2: Download the NCBI home page by setting $web_file to "http://ncbi.nih.gov" and $store to "ncbi.html".
  Perl -e


Calculate simple statistics

The tools in this section calculate very simple statistics about the data (often tabular data) in given input files.
To use a script, cut and paste the code from the light green or blue box into a terminal window, change the bold, red text as needed, and hit Enter.
See More Information for notes on using these tools.

Calculate statistics about a column

Calculate sum of the nth column of tabular data (calc_col_sum)

(This tool should not be confused with calc_row_sum, which calculates the sum of all columns for each row.)
$colColumn to sum
Input file(s)
perl -e ' $col=1; while(<>) { s/\r?\n//; @F=split /\t/, $_; $sum += $F[$col]; } warn "\nSum of column $col for $. lines\n\n"; print "$sum\n" ' file.tab
Example: Sum second column of file.tab by running the above script.
Input file (file.tab)Screen output
 Fly    7
 Human  14
 Worm   28
 Yeast  35
 Sum of column 1 for 4 lines
 
 84

Calculate length of a given column on each line (calc_col_length)

For a given column of a tab-separated file, calculate the length of the text in that column. For each line, add a column at the end of the line with the length of the chosen column.
$colColumn to measure length of
Input file(s)
Output file
perl -e ' $col=2; while (<>) { s/\r?\n//; @F = split /\t/, $_; $len = length($F[$col]); print "$_\t$len\n" } warn "\nAdded column with length of column $col for $. lines.\n\n"; ' seqs.tab > seqs_length.tab
Example: Take a FASTA file seqs.tab that we've converted to tabular format using change_fasta_to_tab. Calculate the length of column 2 (the third column), which has the sequence in it, by running the above script. Create a new, 4-column, file seqs_length.tab.
Input file
(seqs.tab)
 SEQ1   First seq       ACTGACTG
 SEQ2   Second seq      ACTG
 SEQ3   Third seq       
 SEQ4   Third seq       ACTGACTGACTG
Output file
(seqs_length.tab)
 SEQ1   First seq       ACTGACTG        8
 SEQ2   Second seq      ACTG    4
 SEQ3   Third seq               0
 SEQ4   Third seq       ACTGACTGACTG    12
Screen Output
 Added column with length of column 2 for 4 lines.

Calculate statistics about each row/line

Insert line numbers (calc_line_numbers)

For each line in a file, print the line number followed by a separator (by default, a tab, represented by \t), and then the rest of the line.
$separatorWhat to print between line number and rest of line - \t means tab
Input file(s)
Output file
perl -e ' $separator="\t"; while (<>) { print "$.$separator$_" } warn "\nInserted line numbers for $. lines, with separator $separator.\n\n" ' gene_list.txt > numbered_gene_list.txt
Example: Add line numbers to a list of genes gene_list.txt to generate a numbered list numbered_gene_list.txt by running the above script.
Input file
(gene_list.txt)
 Hsp90  Heat shock protein
 apo1   apoptosis-related protein
 glu7   Glucose metabolism
Output file
(numbered_gene_list.txt)
 1      Hsp90   Heat shock protein
 2      apo1    apoptosis-related protein
 3      glu7    Glucose metabolism
Screen Output
 Inserted line numbers for 3 lines, with separator 


Calculate sum of two or more columns for each row (calc_row_sum)

@colsWhich column(s) to add
Input file(s)
Output file
perl -e ' @cols=(1, 2, 3); while(<>) { s/\r?\n//; @F=split /\t/, $_; $sum = 0; foreach $col (@cols) { $sum += $F[$col] }; print "$_\t$sum\n"; } warn "\nSum of columns @cols for each line ($. lines)\n\n" ' in.tab > out.tab

Calculate how many times each value appears in a given column (calc_repeats_for_each_value_in_col)

For a given column of a tab-separated file, count how many times each value appears in that column. Each line of the output will have a value, a tab, and the number of times it appears. Values will be printed in the order of their first appearance.
$colColumn to count repeats in
Input file(s)
Output file
perl -e ' $col=1; while (<>) { s/\r?\n//; @F = split /\t/, $_; $val = $F[$col]; if (! exists $count{$val}) { push @order, $val } $count{$val}++; } foreach $val (@order) { print "$val\t$count{$val}\n" } warn "\nPrinted number of occurrences for ", scalar(@order), " values in $. lines.\n\n"; ' gene_go.txt > go_repeats.txt
Example: Given a list of genes with associated GO terms gene_go.txt, make a new file go_repeats.txt showing how many times each GO term is found by running the above script. This could be used to find whether a certain biological process is over-represented in a list of genes, for example. (Note: changing the column to 0 would calculate how many GO terms each gene was associated with.)
Input file
(gene_go.txt)
 Hsp90  GO:00171
 apo1   GO:00012
 apo1   GO:00233
 apo1   GO:01234
 glu7   GO:00012
 glu7   GO:56785
Output file
(go_repeats.txt)
 GO:00171       1
 GO:00012       2
 GO:00233       1
 GO:01234       1
 GO:56785       1
Screen Output
 Printed number of occurrences for 5 values in 6 lines.

Calculate sum of values in a column for each value in another column (calc_sum_of_col_for_groups_of_lines)

Find sets of rows that have the same value in column m. Then get the sum of column n for those sets of rows. Each line of the output will have a value, a tab, and the sum for that value. Values will be printed in the order of their first appearance.
$value_colColumn to determine grouping of lines
$sum_colColumn to sum for sets of similar lines
Input file(s)
Output file
perl -e ' $value_col=0; $sum_col=2; while (<>) { s/\r?\n//; @F = split /\t/, $_; $val = $F[$value_col]; if (! exists $sum{$val}) { push @order, $val } $sum{$val} += $F[$sum_col]; } foreach $val (@order) { print "$val\t$sum{$val}\n" } warn "\nPrinted sum of column $sum_col for each value in column $value_col\nFound ", scalar(@order), " values in $. lines\n\n"; ' exon_length.tab > gene_length.tab
Example: Calculate the sum of the length of the exons for each gene.
Input file
(exon_length.tab)
 Hsp90  exon1   300
 Hsp90  exon2   100
 Hsp90  exon3   250
 apo1   exon1   100
 apo1   exon2   350
Output file
(gene_length.tab)
 Hsp90  650
 apo1   450
Screen Output
 Printed sum of column 2 for each value in column 0
 Found 2 values in 5 lines

Count lines or records in a file

Count lines or records in a file

Count lines in a file (calc_num_lines)

Simply give a count of the number of lines in a file. (The result is printed to an output file as well as the screen.)
Input file(s)
Output file
perl -e ' while (<>) { } print "Counted $. lines\n"; warn "\nCounted $. lines\n\n" ' gene_list.txt > gene_count.txt
Example: Count how many genes are in file gene_list.txt by running the above script.
Input file
(gene_list.txt)
 Hsp90  Heat shock protein
 apo1   apoptosis-related protein
 glu7   Glucose metabolism
Output file
(gene_count.txt)
 Counted 3 lines
Screen Output
 Counted 3 lines
UNIX/Mac users: also check out the wc command.

Count records in a FASTA file (calc_num_fasta_records)

Counts the number of records (and, for convenience, total sequence length) in a FASTA file. (The result is printed to an output file as well as the screen.)

Input file(s)
Output file


perl -e ' $count=0; $len=0; while(<>) { s/\r?\n//; if (/^>/) { $count++; } else { $len += length($_) } } print "Read $count FASTA records in $. lines. Total sequence length: $len\n"; warn "\nRead $count FASTA records in $. lines. Total sequence length: $len\n\n"; ' seqs.fasta > seqs_count.txt
Example: See how many sequences are in seqs.fasta.
Input file
(seqs.fasta)
 >CG123 A small sequence
 ACGTTGCA
 GTTACCAG
 >EG12
 ACCGGA
 >DG124  A smaller sequence
 GTTACCAG
Output file
(seqs_count.txt)
 Read 3 FASTA records in 7 lines. Total sequence length: 30
Screen Output
 Read 3 FASTA records in 7 lines. Total sequence length: 30

Calculate line numbers for each line