Class: RssFeed::Feed::Namespace

Inherits:
Object
  • Object
show all
Defined in:
lib/rss_feed/feed/namespace.rb

Overview

The Namespace module provides utility methods for accessing and manipulating XML namespaces in RSS feeds.

Constant Summary collapse

NAMESPACES =

Mapping of namespace prefixes to their corresponding URIs.

{
  'itunes' => 'http://www.itunes.com/dtds/podcast-1.0.dtd',
  'dc' => 'http://purl.org/dc/elements/1.1/',
  'feedburner' => 'http://rssnamespace.org/feedburner/ext/1.0',
  'content' => 'http://purl.org/rss/1.0/modules/content/',
  'trackback' => 'http://example.com/trackback',
  'media' => 'http://search.yahoo.com/mrss/',
  'atom' => 'http://www.w3.org/2005/Atom',
  'xmlns' => 'http://www.w3.org/2005/Atom'
}.freeze

Class Method Summary collapse

Class Method Details

.access_tag(tag, doc, feed) ⇒ Hash

Accesses the specified XML tag within the document with proper namespace handling.

Parameters:

  • tag (String) —

    The XML tag to access.

  • doc (Nokogiri::XML::Document) —

    The XML document.

Returns:

  • (Hash) —

    The tag data including text, nested elements flag, nested attributes flag, and the document.



26
27
28
29
30
31
# File 'lib/rss_feed/feed/namespace.rb', line 26

def access_tag(tag, doc, feed)
  feed_tag = %w[atom feed].include?(feed.detect_feed_type) && namespace(tag).blank? ? "xmlns:#{tag}" : tag
  doc = doc.xpath(feed_tag, namespace(tag))
  nested_elements = nested_elements?(doc)
  { text: doc.to_s, nested_elements: nested_elements, nested_attributes: nested_attributes?(doc), docs: doc }
end

.namespace(tag) ⇒ Hash

Resolves the namespace for the given XML tag.

Parameters:

  • tag (String) —

    The XML tag.

Returns:

  • (Hash) —

    The namespace declaration.



37
38
39
40
# File 'lib/rss_feed/feed/namespace.rb', line 37

def namespace(tag)
  namespace_key = tag.split(':').first
  NAMESPACES[namespace_key].blank? ? nil : { namespace_key.to_s => NAMESPACES[namespace_key] }.compact
end

.nested_attributes?(node) ⇒ Boolean

Checks if the XML node has nested attributes.

Parameters:

  • node (Nokogiri::XML::NodeSet) —

    The XML node.

Returns:

  • (Boolean) —

    Whether the node has nested attributes.



70
71
72
73
74
# File 'lib/rss_feed/feed/namespace.rb', line 70

def nested_attributes?(node)
  return false if node.blank? || node.to_s == 'NaN'

  node.any? { |thumbnail| !thumbnail.attributes.empty? }
end

.nested_elements?(node) ⇒ Boolean

Checks if the XML node has nested elements.

Parameters:

  • node (Nokogiri::XML::NodeSet) —

    The XML node.

Returns:

  • (Boolean) —

    Whether the node has nested elements.



58
59
60
61
62
63
64
# File 'lib/rss_feed/feed/namespace.rb', line 58

def nested_elements?(node)
  return false if node.blank? || node.to_s == 'NaN'

  return true if node.children.any?(&:element?)

  false
end

.remove_html_tags(content) ⇒ String

Removes HTML tags from the given content.

Parameters:

  • content (String) —

    The content containing HTML tags.

Returns:

  • (String) —

    The content without HTML tags.



46
47
48
49
50
51
52
# File 'lib/rss_feed/feed/namespace.rb', line 46

def remove_html_tags(content)
  if %r{([^-_.!~*'()a-zA-Z\d;/?:@&=+$,\[\]]%)}.match?(content)
    CGI.unescape(content)
  else
    content
  end.gsub(/(<!\[CDATA\[|\]\]>)/, '').strip.gsub(/<[^>]+>/, '')
end