[ 
https://issues.apache.org/jira/browse/NUTCH-3210?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Lewis John McGibbney updated NUTCH-3210:
----------------------------------------
    Description: 
Add an Atlas IndexWriter plugin that, at IndexingJob time, upserts lineage into 
[Apache Atlas|http://atlas.apache.org/] via AtlasClientV2 REST.

*Motivation*

Atlas will own Nutch typedefs and a CrawlDb batch bridge (ATLAS-XXXX). 
Index-time lineage (crawl/segment to index DataSet) belongs in Nutch so it runs 
with the existing indexer plugin path.

*Scope*
 # New plugin src/plugin/indexer-atlas implementing IndexWriter
 # Configure via conf/index-writers.xml and plugin.includes
 # One nutch_index_process per IndexingJob (not per URL)
 # Process.inputs: nutch_crawl and/or nutch_segment (must already exist in 
Atlas)
 # Process.outputs: generic DataSet named by collection/index + clusterName
 # Optional updates to existing crawl–host relationship indexedCount / 
lastIndexedTime only when the host already exists
 # Auth: Atlas REST username/password or token; atlas.cluster.name shared with 
the Atlas bridge
 # Unit tests with a mocked Atlas client; docs for config keys

*Out of scope*
 # Reusing indexer-kafka (payload is NutchDocument JSON, not HookNotification)
 # AtlasHook / Kafka ATLAS_HOOK
 # Creating crawl, seedlist, domain, or host entities (Atlas CrawlDb bridge)
 # URL-level Atlas entities
 # Restoring NutchServer / JobManager metadata

Depends on: ATLAS-5399 (Nutch type model; process/index DataSet types must 
exist)

  was:
Add an Atlas IndexWriter plugin that, at IndexingJob time, upserts lineage into 
[Apache Atlas|http://atlas.apache.org] via AtlasClientV2 REST.

*Motivation*

Atlas will own Nutch typedefs and a CrawlDb batch bridge (ATLAS-XXXX). 
Index-time lineage (crawl/segment to index DataSet) belongs in Nutch so it runs 
with the existing indexer plugin path.

*Scope*
 # New plugin src/plugin/indexer-atlas implementing IndexWriter
 # Configure via conf/index-writers.xml and plugin.includes
 # One nutch_index_process per IndexingJob (not per URL)
 # Process.inputs: nutch_crawl and/or nutch_segment (must already exist in 
Atlas)
 # Process.outputs: generic DataSet named by collection/index + clusterName
 # Optional updates to existing crawl–host relationship indexedCount / 
lastIndexedTime only when the host already exists
 # Auth: Atlas REST username/password or token; atlas.cluster.name shared with 
the Atlas bridge
 # Unit tests with a mocked Atlas client; docs for config keys

*Out of scope*
 # Reusing indexer-kafka (payload is NutchDocument JSON, not HookNotification)
 # AtlasHook / Kafka ATLAS_HOOK
 # Creating crawl, seedlist, domain, or host entities (Atlas CrawlDb bridge)
 # URL-level Atlas entities
 # Restoring NutchServer / JobManager metadata

Depends on: ATLAS-XXXX (Nutch type model; process/index DataSet types must 
exist)


> Add Atlas IndexWriter to emit crawl/segment lineage to Apache Atlas
> -------------------------------------------------------------------
>
>                 Key: NUTCH-3210
>                 URL: https://issues.apache.org/jira/browse/NUTCH-3210
>             Project: Nutch
>          Issue Type: New Feature
>          Components: indexer, plugin
>    Affects Versions: 1.24
>            Reporter: Lewis John McGibbney
>            Assignee: Lewis John McGibbney
>            Priority: Major
>             Fix For: 1.24
>
>
> Add an Atlas IndexWriter plugin that, at IndexingJob time, upserts lineage 
> into [Apache Atlas|http://atlas.apache.org/] via AtlasClientV2 REST.
> *Motivation*
> Atlas will own Nutch typedefs and a CrawlDb batch bridge (ATLAS-XXXX). 
> Index-time lineage (crawl/segment to index DataSet) belongs in Nutch so it 
> runs with the existing indexer plugin path.
> *Scope*
>  # New plugin src/plugin/indexer-atlas implementing IndexWriter
>  # Configure via conf/index-writers.xml and plugin.includes
>  # One nutch_index_process per IndexingJob (not per URL)
>  # Process.inputs: nutch_crawl and/or nutch_segment (must already exist in 
> Atlas)
>  # Process.outputs: generic DataSet named by collection/index + clusterName
>  # Optional updates to existing crawl–host relationship indexedCount / 
> lastIndexedTime only when the host already exists
>  # Auth: Atlas REST username/password or token; atlas.cluster.name shared 
> with the Atlas bridge
>  # Unit tests with a mocked Atlas client; docs for config keys
> *Out of scope*
>  # Reusing indexer-kafka (payload is NutchDocument JSON, not HookNotification)
>  # AtlasHook / Kafka ATLAS_HOOK
>  # Creating crawl, seedlist, domain, or host entities (Atlas CrawlDb bridge)
>  # URL-level Atlas entities
>  # Restoring NutchServer / JobManager metadata
> Depends on: ATLAS-5399 (Nutch type model; process/index DataSet types must 
> exist)



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to