Posts

Showing posts with the label Elasticsearch

Elasticsearch Integration with Spring Boot: A Comprehensive Guide

  Elasticsearch Integration with Spring Boot: A Comprehensive Guide Elasticsearch is a powerful search and analytics engine that is commonly used for various applications like full-text search, logging, and analytics. Integrating Elasticsearch with a Spring Boot application can significantly enhance its search capabilities and provide real-time search functionalities. In this guide, we'll walk through a simple example of how to integrate Elasticsearch with a Spring Boot application. We will cover the setup, configuration, and basic operations such as indexing and searching data. 1. Setting Up the Project First, you'll need a Spring Boot application. You can either create a new Spring Boot project or use an existing one. For this example, we'll use Spring Initializr to generate a new project. Go to Spring Initializr . Choose your project metadata (Group, Artifact, Name, etc.). Add the following dependencies: Spring Web Spring Data Elasticsearch Click "Generate" to ...

Learn more about rdd in spark

RDD (Resilient Distributed Dataset) is called the elastic distributed data set. It is the most basic data abstraction in Spark. It represents an immutable, partitionable, and set of elements that can be calculated in parallel. RDD has the characteristics of a data flow model: automatic fault tolerance, location-aware scheduling, and scalability. RDD allows users to explicitly cache the working set in memory when executing multiple queries. Subsequent queries can reuse the working set, which greatly improves query speed. Today, let's talk briefly about RDD in Spark. The RDD API will be put into the next chapter and then detailed RDD Introduction RDD can be regarded as an object of Spark, which runs in memory itself. For example, reading a file is an RDD, calculating a file is an RDD, and the result set is also an RDD. Different shards, data dependencies, key- Value type map data can be regarded as RDD. (Note: from Baidu Encyclopedia), here, RDD will not say much, just talk abou...

Spring Data Elasticsearch MappingElasticsearchConverter

Spring Data Elasticsearch MappingElasticsearchConverter  The MappingElasticsearchConverter uses metadata to drive the mapping of objects to documents. The metadata is taken from the entity’s properties which can be annotated. The following annotations are available: @Document: Applied at the class level to indicate this class is a candidate for mapping to the database. The most important attributes are: indexName: the name of the index to store this entity in type: the mapping type. If not set, the lowercased simple name of the class is used. (deprecated since version 4.0) shards: the number of shards for the index. replicas: the number of replicas for the index. refreshIntervall: Refresh interval for the index. Used for index creation. Default value is "1s". indexStoreType: Index storage type for the index. Used for index creation. Default value is "fs". createIndex: Configuration whether to create an index on repository bootstrapping. Default...

Spring Data Elasticsearch @Field

Spring Data Elasticsearch @Field The @Field annotation now supports nearly all of the types that can be used in Elasticsearch. @Document(indexName = "person", type = "dummy") public class Person implements Persistable<Long> {     @Nullable @Id     private Long id;     @Nullable @Field(value = "last-name", type = FieldType.Text, fielddata = true)     private String lastName;      (1)     @Nullable @Field(name = "birth-date", type = FieldType.Date, format = DateFormat.basic_date)     private LocalDate birthDate;  (2)     @CreatedDate     @Nullable @Field(type = FieldType.Date, format = DateFormat.basic_date_time)     private Instant created;      (3)     // other properties, getter, setter } in Elasticsearch this field will be named last-name, this mapping is handled transparently a property for a date without time informat...

Elasticsearch - async search

Elasticsearch - async search Asynchronous search Asynchronous search makes long-running queries feasible and reliable. Async search allows users to run long-running queries in the background, track the query progress, and retrieve partial results as they become available. Async search enables users to more easily search vast amounts of data with no more pesky timeouts. Submit async search API Executes a search request asynchronously. It accepts the same parameters and request body as the search API. POST /sales*/_async_search?size=0 {     "sort" : [       { "date" : {"order" : "asc"} }     ],     "aggs" : {         "sale_date" : {              "date_histogram" : {                  "field" : "date",                  "calendar_interval": "1d"           ...

ELK tutorial

Image
What is ELK? ELK is a combination of 3 open source products: Elasticsearch Logstash Kibana All developed and maintained by Elastic . Elasticsearch is a NoSQL database based on the Lucene search engine. Logstash is a log pipeline tool that accepts data input, performs data conversion, and then outputs data. Kibana is an interface layer that works on top of Elasticsearch. In addition, the ELK stack also contains a series of log collector tools called Beats. The most common usage scenario of ELK is as a log system for Internet products. Of course, the ELK stack can also be used in other aspects, such as: business intelligence, big data analysis, etc. Why use ELK? The ELK stack is very popular, because it is powerful, open source and free. For smaller companies such as SaaS companies and startups, using ELK to build a log system is very cost-effective. Netflix, Facebook, Microsoft, LinkedIn and Cisco also use ELK to monitor logs. Why use a logging system? The log sys...

Spring Elasticsearch Repositories

Elasticsearch Repositories Query creation Generally the query creation mechanism for Elasticsearch works as described in Query methods. Here’s a short example of what a Elasticsearch query method translates into: Query creation from method names interface BookRepository extends Repository<Book, String> {   List<Book> findByNameAndPrice(String name, Integer price); } The method name above will be translated into the following Elasticsearch json query { "bool" :     { "must" :         [             { "field" : {"name" : "?"} },             { "field" : {"price" : "?"} }         ]     } } Using @Query Annotation Declare query at the method using the @Query annotation. interface BookRepository extends ElasticsearchRepository<Book, String> {     @Query("{\"bool\" : {\"must\" : {\"field\" : {\"name\" : \"?0\"}}}}...

Spring Elasticsearch Operations

Elasticsearch Operations Spring Data Elasticsearch uses two interfaces to define the operations that can be called against an Elasticsearch index. These are ElasticsearchOperations and ReactiveElasticsearchOperations. Whereas the first is used with the classic synchronous implementations, the second one uses reactive infrastructure. The default implementations of the interfaces offer: Read/Write mapping support for domain types. A rich query and criteria api. Resource management and Exception translation. ElasticsearchTemplate The ElasticsearchTemplate is an implementation of the ElasticsearchOperations interface using the Transport Client. @Configuration public class TransportClientConfig extends ElasticsearchConfigurationSupport {   @Bean   public Client elasticsearchClient() throws UnknownHostException {                      Settings settings = Settings.builder().put("cluster.name", "elasticsearch")....

Spring Elasticsearch Object Mapping

Elasticsearch Object Mapping Spring Data Elasticsearch allows to choose between two mapping implementations abstracted via the EntityMapper interface: Jackson Object Mapping Meta Model Object Mapping Jackson Object Mapping The Jackson2 based approach (used by default) utilizes a customized ObjectMapper instance with spring data specific modules. Extensions to the actual mapping need to be customized via Jackson annotations like @JsonInclude. @Configuration public class Config extends AbstractElasticsearchConfiguration {    @Override   public RestHighLevelClient elasticsearchClient() {     return RestClients.create(ClientConfiguration.create("localhost:9200")).rest();   } } AbstractElasticsearchConfiguration already defines a Jackson2 based entityMapper via ElasticsearchConfigurationSupport. CustomConversions, @ReadingConverter & @WritingConverter cannot be applied when using the Jackson based EntityMapper. Setting the name of a...

Spring Elasticsearch Clients

Spring Elasticsearch Clients Spring data Elasticsearch operates upon an Elasticsearch client that is connected to a single Elasticsearch node or a cluster. Although the Elasticsearch Client can be used to work with the cluster, applications using Spring Data Elasticsearch normally use the higher level abstractions of Elasticsearch Operations and Elasticsearch Repositories. Transport Client static class Config {   @Bean   Client client() {   Settings settings = Settings.builder()     .put("cluster.name", "elasticsearch")          .build();   TransportClient client = new PreBuiltTransportClient(settings);     client.addTransportAddress(new TransportAddress(InetAddress.getByName("127.0.0.1")       , 9300));                                    return client;   } } // ... IndexRequest request = ...