Giter Club home page Giter Club logo

go-dcp-elasticsearch's Introduction

Go Dcp Elasticsearch

Go Reference Go Report Card

Go implementation of the Elasticsearch Connect Couchbase.

Go Dcp Elasticsearch streams documents from Couchbase Database Change Protocol (DCP) and writes to Elasticsearch index in near real-time.

Features

  • Less resource usage and higher throughput(see Benchmarks).
  • Custom routing support(see Example).
  • Update multiple documents for a DCP event(see Example).
  • Handling different DCP events such as expiration, deletion and mutation(see Example).
  • Elasticsearch compression request body support.
  • Managing batch configurations such as maximum batch size, batch bytes, batch ticker durations.
  • Scale up and down by custom membership algorithms(Couchbase, KubernetesHa, Kubernetes StatefulSet or Static, see examples).
  • Easily manageable configurations.

Benchmarks

The benchmark was made with the 1,001,006 Couchbase document, because it is possible to more clearly observe the difference in the batch structure between the two packages. Default configurations for Java Elasticsearch Connect Couchbase used for both connectors.

Package Time to Process Events Elasticsearch Indexing Rate(/s) Average CPU Usage(Core) Average Memory Usage
Go Dcp Elasticsearch(Go 1.20) 50s go 0.486 408MB
Java Elasticsearch Connect Couchbase(JDK15) 80s go 0.31 1091MB

Example

Struct Config

func mapper(event couchbase.Event) []document.ESActionDocument {
	if event.IsMutated {
		e := document.NewIndexAction(event.Key, event.Value, nil)
		return []document.ESActionDocument{e}
	}
	e := document.NewDeleteAction(event.Key, nil)
	return []document.ESActionDocument{e}
}

func main() {
	connector, err := dcpelasticsearch.NewConnectorBuilder(config.Config{
		Elasticsearch: config.Elasticsearch{
			CollectionIndexMapping: map[string]string{
				"_default": "indexname",
			},
			Urls: []string{"http://localhost:9200"},
		},
		Dcp: dcpConfig.Dcp{
			Username:   "user",
			Password:   "password",
			BucketName: "dcp-test",
			Hosts:      []string{"localhost:8091"},
			Dcp: dcpConfig.ExternalDcp{
				Group: dcpConfig.DCPGroup{
					Name: "groupName",
					Membership: dcpConfig.DCPGroupMembership{
						Type: "static",
					},
				},
			},
			Metadata: dcpConfig.Metadata{
				Config: map[string]string{
					"bucket":     "checkpoint-bucket-name",
					"scope":      "_default",
					"collection": "_default",
				},
				Type: "couchbase",
			},
		},
	}).
		SetMapper(mapper).
		Build()
	if err != nil {
		panic(err)
	}

	defer connector.Close()
	connector.Start()
}

File Config

Default Mapper

Configuration

Dcp Configuration

Check out on go-dcp

Elasticsearch Specific Configuration

Variable Type Required Default Description
elasticsearch.collectionIndexMapping map[string]string yes Defines which Couchbase collection events will be written to which index
elasticsearch.urls []string yes Elasticsearch connection urls
elasticsearch.typeName string no _doc Defines Elasticsearch index type name
elasticsearch.batchSizeLimit int no 1000 Maximum message count for batch, if exceed flush will be triggered.
elasticsearch.batchTickerDuration time.Duration no 10s Batch is being flushed automatically at specific time intervals for long waiting messages in batch.
elasticsearch.batchByteSizeLimit int no 10485760 Maximum size(byte) for batch, if exceed flush will be triggered.
elasticsearch.maxConnsPerHost int no 512 Maximum number of connections per each host which may be established
elasticsearch.maxIdleConnDuration time.Duration no 10s Idle keep-alive connections are closed after this duration.
elasticsearch.compressionEnabled boolean no false Compression can be used if message size is large, CPU usage may be affected.
elasticsearch.concurrentRequest int no 1 Concurrent bulk request count
elasticsearch.disableDiscoverNodesOnStart boolean no false Disable discover nodes when initializing the client.
elasticsearch.discoverNodesInterval time.Duration no 5m Discover nodes periodically

Exposed metrics

Metric Name Description Labels Value Type
elasticsearch_connector_latency_ms Time to adding to the batch. N/A Gauge
elasticsearch_connector_bulk_request_process_latency_ms Time to process bulk request. N/A Gauge

You can also use all DCP-related metrics explained here. All DCP-related metrics are automatically injected. It means you don't need to do anything.

Contributing

Go Dcp Elasticsearch is always open for direct contributions. For more information please check our Contribution Guideline document.

License

Released under the MIT License.

go-dcp-elasticsearch's People

Contributors

erayarslan avatar mhmtszr avatar abdulsametileri avatar burhanelgun avatar ramazan avatar canerpatir avatar oguzyildirim avatar alihanyalcin avatar

Watchers

 avatar

Recommend Projects

  • React photo React

    A declarative, efficient, and flexible JavaScript library for building user interfaces.

  • Vue.js photo Vue.js

    ๐Ÿ–– Vue.js is a progressive, incrementally-adoptable JavaScript framework for building UI on the web.

  • Typescript photo Typescript

    TypeScript is a superset of JavaScript that compiles to clean JavaScript output.

  • TensorFlow photo TensorFlow

    An Open Source Machine Learning Framework for Everyone

  • Django photo Django

    The Web framework for perfectionists with deadlines.

  • D3 photo D3

    Bring data to life with SVG, Canvas and HTML. ๐Ÿ“Š๐Ÿ“ˆ๐ŸŽ‰

Recommend Topics

  • javascript

    JavaScript (JS) is a lightweight interpreted programming language with first-class functions.

  • web

    Some thing interesting about web. New door for the world.

  • server

    A server is a program made to process requests and deliver data to clients.

  • Machine learning

    Machine learning is a way of modeling and interpreting data that allows a piece of software to respond intelligently.

  • Game

    Some thing interesting about game, make everyone happy.

Recommend Org

  • Facebook photo Facebook

    We are working to build community through open source technology. NB: members must have two-factor auth.

  • Microsoft photo Microsoft

    Open source projects and samples from Microsoft.

  • Google photo Google

    Google โค๏ธ Open Source for everyone.

  • D3 photo D3

    Data-Driven Documents codes.