Class: Retriever::FetchFiles
Instance Attribute Summary
Attributes inherited from Fetch
Instance Method Summary collapse
-
#autodownload ⇒ Object
when autodownload option is true, this will automatically go through the fetched file URL collection and download each one.
-
#download_file(path) ⇒ Object
given valid url, downloads file to current directory in /rr-downloads/.
-
#initialize(url, options) ⇒ FetchFiles
constructor
recieves target url and RR options, returns an array of all unique files (based on given filetype) found on the site.
Methods inherited from Fetch
#asyncGetWave, #async_crawl_and_collect, #dump, #errlog, #good_response?, #lg, #write
Constructor Details
#initialize(url, options) ⇒ FetchFiles
recieves target url and RR options, returns an array of all unique files (based on given filetype) found on the site
3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 |
# File 'lib/retriever/fetchfiles.rb', line 3 def initialize(url,) #recieves target url and RR options, returns an array of all unique files (based on given filetype) found on the site super @data = [] page_one = Retriever::Page.new(@t.source,@t) @linkStack = page_one.parseInternalVisitable lg("URL Crawled: #{@t.target}") lg("#{@linkStack.size-1} new links found") tempFileCollection = page_one.parseFiles @data.concat(tempFileCollection) if tempFileCollection.size>0 lg("#{@data.size} new files found") errlog("Bad URL -- #{@t.target}") if !@linkStack @linkStack.delete(@t.target) if @linkStack.include?(@t.target) @linkStack = @linkStack.take(@maxPages) if (@linkStack.size+1 > @maxPages) self.async_crawl_and_collect() @data.sort_by! {|x| x.length} @data.uniq! end |
Instance Method Details
#autodownload ⇒ Object
when autodownload option is true, this will automatically go through the fetched file URL collection and download each one.
35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 |
# File 'lib/retriever/fetchfiles.rb', line 35 def autodownload() #when autodownload option is true, this will automatically go through the fetched file URL collection and download each one. lenny = @data.count puts "###################" puts "### Initiating Autodownload..." puts "###################" puts "#{lenny} - #{@file_ext}'s Located" puts "###################" if File::directory?("rr-downloads") Dir.chdir("rr-downloads") else puts "creating rr-downloads Directory" Dir.mkdir("rr-downloads") Dir.chdir("rr-downloads") end file_counter = 0 @data.each do |entry| begin self.download_file(entry) file_counter+=1 lg(" File [#{file_counter} of #{lenny}]") puts rescue StandardError => e puts "ERROR: failed to download - #{entry}" puts e. puts end end Dir.chdir("..") end |
#download_file(path) ⇒ Object
given valid url, downloads file to current directory in /rr-downloads/
24 25 26 27 28 29 30 31 32 33 34 |
# File 'lib/retriever/fetchfiles.rb', line 24 def download_file(path) #given valid url, downloads file to current directory in /rr-downloads/ arr = path.split('/') shortname = arr.pop puts "Initiating Download to: #{'/rr-downloads/' + shortname}" File.open(shortname, "wb") do |saved_file| open(path) do |read_file| saved_file.write(read_file.read) end end puts " SUCCESS: Download Complete" end |