È«ÎļìË÷ÒýÇæSolrϵÁСª¡ªÕûºÏÖÐÎÄ·Ö´Ê×é¼þIKAnalyzer
IK AnalyzerÊÇÒ»¿î½áºÏÁ˴ʵäºÍÎÄ·¨·ÖÎöËã·¨µÄÖÐÎÄ·Ö´Ê×é¼þ£¬»ùÓÚ×Ö·û´®Æ¥Å䣬֧³ÖÓû§´ÊµäÀ©Õ¹¶¨Ò壬֧³ÖϸÁ£¶ÈºÍÖÇÄÜÇз֣¬±ÈÈ磺
ÕÅÈý˵µÄȷʵÔÚÀí
ÖÇÄִܷʵĽá¹ûÊÇ£º
ÕÅÈý | ˵µÄ | ȷʵ | ÔÚÀí
×îϸÁ£¶È·Ö´Ê½á¹û£º
ÕÅÈý | Èý | ˵µÄ | µÄÈ· | µÄ | ȷʵ | ʵÔÚ | ÔÚÀí
ÕûºÏIK Analyzer±Èmmseg4jÒª¼òµ¥ºÜ¶à£¬ÏÂÔØ½âѹËõIKAnalyzer2012FF_u1.jar·Åµ½Ä¿Â¼£ºE:\solr-4.8.0\example\solr-webapp\webapp\WEB-INF\lib£¬ÐÞ¸ÄÅäÖÃÎļþschema.xml£¬Ìí¼Ó´úÂ룺
<field name="content" type="text_ik" indexed="true" stored="true"/> <fieldType name="text_ik" class="solr.TextField"> <analyzer type="index" isMaxWordLength="false" class="org.wltea.analyzer.lucene.IKAnalyzer"/> <analyzer type="query" isMaxWordLength="true" class="org.wltea.analyzer.lucene.IKAnalyzer"/> </fieldType> |
²éѯ²ÉÓÃIK×Ô¼ºµÄ×î´ó·Ö´Ê·¨,Ë÷ÒýÔò²ÉÓÃËüµÄϸÁ£¶È·Ö´Ê·¨
´Ëʱ¾ÍËãÅäÖÃÍê³ÉÁË£¬ÖØÆô·þÎñ£ºjava -jar start.jar£¬À´¿´¿´IKAnalyzerµÄ·Ö´ÊЧ¹ûÔõôÑù£¬´ò¿ªSolr¹ÜÀí½çÃæ£¬µã»÷×ó²àµÄAnalysisÒ³Ãæ

ĬÈÏ·Ö´ÊÆ÷½øÐÐ×îϸÁ£¶ÈÇз֡£IKAnalyzerÖ§³Öͨ¹ýÅäÖÃIKAnalyzer.cfg.xml
ÎļþÀ´À©³äÄúµÄÓëÓдʵäÒÔ¼°Í£Ö¹´Êµä£¨¹ýÂ˴ʵ䣩£¬Ö»Ðè°ÑIKAnalyzer.cfg.xmlÎļþ·ÅÈëclassĿ¼ÏÂÃæ£¬Ö¸¶¨×Ô¼ºµÄ´Êµämydic.dics
<?xml version="1.0" encoding="UTF-8"?> <!DOCTYPE properties SYSTEM "http://java.sun.com/dtd/properties.dtd"> <properties> <comment>IK Analyzer À©Õ¹ÅäÖÃ</comment> <!--Óû§¿ÉÒÔÔÚÕâÀïÅäÖÃ×Ô¼ºµÄÀ©Õ¹×Öµä --> <entry key="ext_dict">/mydict.dic; /com/mycompany/dic/mydict2.dic;</entry> <!--Óû§¿ÉÒÔÔÚÕâÀïÅäÖÃ×Ô¼ºµÄÀ©Õ¹Í£Ö¹´Ê×Öµä--> <entry key="ext_stopwords">/ext_stopword.dic</entry> </properties> |
ÊÂʵÉÏÇ°ÃæµÄFieldTypeÅäÖÃÆäʵ´æÔÚÎÊÌ⣬¸ù¾ÝĿǰ×îеÄIK°æ±¾IK Analyzer 2012FF_hf1.zip£¬Ë÷ÒýʱʹÓÃ×îϸÁ£¶È·Ö´Ê£¬²éѯʱ×î´ó·Ö´Ê£¨ÖÇÄÜ·Ö´Ê£©Êµ¼ÊÉÏÊDz»ÉúЧµÄ¡£
¾Ý×÷Õßlinliangyi˵£¬ÔÚ2012FF_hf1Õâ¸ö°æ±¾ÖÐÒѾÐÞ¸´£¬¾²âÊÔ»¹ÊÇûÓã¬ÏêÇéÇë¿´´ËÌù¡£
½â¾ö°ì·¨£ºÖØÐÂʵÏÖIKAnalyzerSolrFactory
package org.wltea.analyzer.lucene; import java.io.Reader; import java.util.Map; import org.apache.lucene.analysis.Tokenizer; import org.apache.lucene.analysis.util.TokenizerFactory; //lucene:4.8֮ǰµÄ°æ±¾ //import org.apache.lucene.util.AttributeSource.AttributeFactory; //lucene:4.9 import org.apache.lucene.util.AttributeFactory; public class IKAnalyzerSolrFactory extends TokenizerFactory{ private boolean useSmart; public boolean useSmart() { return useSmart; } public void setUseSmart(boolean useSmart) { this.useSmart = useSmart; } public IKAnalyzerSolrFactory(Map<String,String> args) { super(args); assureMatchVersion(); this.setUseSmart(args.get("useSmart").toString().equals("true")); } @Override public Tokenizer create(AttributeFactory factory, Reader input) { Tokenizer _IKTokenizer = new IKTokenizer(input , this.useSmart); return _IKTokenizer; } } |
ÖØÐ±àÒëºó¸üÐÂjarÎļþ£¬¸üÐÂschema.xmlÎļþ£º
<fieldType name="text_ik" class="solr.TextField" > <analyzer type="index"> <tokenizer class="org.wltea.analyzer.lucene.IKAnalyzerSolrFactory" useSmart="false"/> </analyzer> <analyzer type="query"> <tokenizer class="org.wltea.analyzer.lucene.IKAnalyzerSolrFactory" useSmart="true"/> </analy |
È«ÎļìË÷ÒýÇæSolrϵÁСª¡ªÕûºÏMySQL¡¢MongoDB
MySQL
¿½±´mysql-connector-java-5.1.25-bin.jarµ½E:\solr-4.8.0\example\solr-webapp\webapp\WEB-INF\libĿ¼ÏÂÃæ
ÅäÖÃE:\solr-4.8.0\example\solr\collection1\conf\solrconfig.xml
<requestHandler name="/dataimport" class="org.apache.solr.handler.dataimport.DataImportHandler"> <lst name="defaults"> <str name="config">data-config.xml</str> </lst> </requestHandler> |
µ¼ÈëÒÀÀµ¿âÎļþ£º
<lib dir="../../../dist/" regex="solr-dataimporthandler-\d.*\.jar"/> |
¼ÓÔÚ
<lib dir="../../../dist/" regex="solr-cell-\d.*\.jar" /> |
Ç°Ãæ¡£
´´½¨E:\solr-4.8.0\example\solr\collection1\conf\data-config.xml£¬Ö¸¶¨MySQLÊý¾Ý¿âµØÖ·£¬Óû§Ãû¡¢ÃÜÂëÒÔ¼°½¨Á¢Ë÷ÒýµÄÊý¾Ý±í
<?xml version="1.0" encoding="UTF-8" ?> <dataConfig> <dataSource type="JdbcDataSource" driver="com.mysql.jdbc.Driver" url="jdbc:mysql://localhost:3306/django_blog" user="root" password=""/> <document name="blog"> <entity name="blog_blog" pk="id" query="select id,title,content from blog_blog" deltaImportQuery="select id,title,content from blog_blog where ID='${dataimporter.delta.id}'" deltaQuery="select id from blog_blog where add_time > '${dataimporter.last_index_time}'" deletedPkQuery="select id from blog_blog where id=0"> <field column="id" name="id" /> <field column="title" name="title" /> <field column="content" name="content"/> </entity> </document> </dataConfig> |
query ÓÃÓÚ³õ´Îµ¼Èëµ½Ë÷ÒýµÄsqlÓï¾ä¡£
¿¼Âǵ½Êý¾Ý±íÖеÄÊý¾ÝÁ¿·Ç³£´ó£¬±ÈÈçǧÍò¼¶£¬²»¿ÉÄÜÒ»´ÎË÷ÒýÍ꣬Òò´ËÐèÒª·ÖÅú´ÎÍê³É£¬ÄÇô²éѯÓï¾äqueryÒªÉèÖÃÁ½¸ö²ÎÊý£º${dataimporter.request.length}
${dataimporter.request.offset}
query=¡±select id,title,content from blog_blog limit ${dataimporter.request.length} offset
${dataimporter.request.offset}¡± |
ÇëÇó£ºhttp://localhost:8983/solr/collection2/dataimport?command=full-import&commit=true&clean=false&offset=0&length=10000
deltaImportQuery ¸ù¾ÝIDÈ¡µÃÐèÒª½øÈëµÄË÷ÒýµÄµ¥ÌõÊý¾Ý¡£
deltaQuery ÓÃÓÚÔöÁ¿Ë÷ÒýµÄsqlÓï¾ä£¬ÓÃÓÚÈ¡µÃÐèÒªÔöÁ¿Ë÷ÒýµÄID¡£
deletedPkQuery ÓÃÓÚÈ¡³öÐèÒª´ÓË÷ÒýÖÐɾ³ýÎĵµµÄµÄID
ΪÊý¾Ý¿â±í×ֶν¨Á¢Óò£¨field£©£¬±à¼E:\solr-4.8.0\example\solr\collection1\conf\schema.xml:
<!-- mysql --> <field name="id" type="string" indexed="true" stored="true" required="true" /> <field name="title" type="text_cn" indexed="true" stored="true" termVectors="true" termPositions="true" termOffsets="true"/> <field name="content" type="text_cn" indexed="true" stored="true" termVectors="true" termPositions="true" termOffsets="true"/> <!-- mysql --> |
Mongodb
°²×°mongo-connector£¬×îºÃʹÓÃÊÖ¶¯°²×°·½Ê½£º
<code>git clone https://github.com/10gen-labs/mongo-connector.git
cd mongo-connector #°²×°Ç°ÐÞ¸Ämongo_connector/constants.pyµÄ±äÁ¿£ºÉèÖÃDEFAULT_COMMIT_INTERVAL
= 0 python setup.py install </code>
ĬÈÏÊDz»»á×Ô¶¯Ìá½»ÁË£¬ÕâÀïÉèÖóÉ×Ô¶¯Ìá½»£¬·ñÔòmongodbÊý¾Ý¿â¸üУ¬Ë÷ÒýÕâ±ßû·¨Í¬Ê±¸üУ¬»òÕßÔÚÃüÁîÐÐÖпÉÒÔÖ¸¶¨ÊÇ·ñ×Ô¶¯Ìá½»£¬²»¹ýÎÒÏÖÔÚ»¹Ã»·¢ÏÖ¡£
ÅäÖÃschema.xml£¬°ÑmongodbÖÐÐèÒª¼ÓÉÏË÷ÒýµÄ×Ö¶ÎÅäÖõ½schema.xmlÎļþÖУº
<?xml version="1.0" encoding="UTF-8" ?> <schema name="example" version="1.5"> <field name="_version_" type="long" indexed="true" stored="true"/> <field name="_id" type="string" indexed="true" stored="true" required="true" multiValued="false" /> <field name="body" type="string" indexed="true" stored="true"/> <field name="title" type="string" indexed="true" stored="true" multiValued="true"/> <field name="text" type="text_general" indexed="true" stored="false" multiValued="true"/> <uniqueKey>_id</uniqueKey> <defaultSearchField>title</defaultSearchField> <solrQueryParser defaultOperator="OR"/> <fieldType name="string" class="solr.StrField" sortMissingLast="true" /> <fieldType name="long" class="solr.TrieLongField" precisionStep="0" positionIncrementGap="0"/> <fieldType name="text_general" class="solr.TextField" positionIncrementGap="100"> <analyzer type="index"> <tokenizer class="solr.StandardTokenizerFactory"/> <filter class="solr.StopFilterFactory" ignoreCase="true" words="stopwords.txt" /> <filter class="solr.LowerCaseFilterFactory"/> </analyzer> <analyzer type="query"> <tokenizer class="solr.StandardTokenizerFactory"/> <filter class="solr.StopFilterFactory" ignoreCase="true" words="stopwords.txt" /> <filter class="solr.SynonymFilterFactory" synonyms="synonyms.txt" ignoreCase="true" expand="true"/> <filter class="solr.LowerCaseFilterFactory"/> </analyzer> </fieldType> </schema> |
Æô¶¯Mongod£º
<code>mongod --replSet myDevReplSet --smallfiles </code> |
³õʼ»¯:rs.initiate()
Æô¶¯mongo-connector:
<code>E:\Users\liuzhijun\workspace\mongo-connector\mongo_connector
\doc_managers>mongo-connector -m localhost:27017 -t
http://localhost:8983/solr/collection2 -n s_soccer.person -u id -d ./solr_doc_manager.py </code> |
-m£ºmongod·þÎñ
-t£ºsolr·þÎñ
-n£ºmongodbÃüÃû¿Õ¼ä£¬¼àÌýdatabase.collection£¬¶à¸öÃüÃû¿Õ¼ä¶ººÅ·Ö¸ô
-u£ºuniquekey
-d£º´¦ÀíÎĵµµÄmanagerÎļþ
×¢Ò⣺mongodbͨ³£Ê¹ÓÃ_id×÷Ϊuniquekey£¬¶øSolrmoreʹÓÃid×÷Ϊuniquekey£¬Èç¹û²»×ö´¦Àí£¬Ë÷ÒýÎļþʱ½«»áʧ°Ü£¬ÓÐÁ½ÖÖ·½Ê½À´´¦ÀíÕâ¸öÎÊÌ⣺
Ö¸¶¨²ÎÊý--unique-key=idµ½mongo-connector£¬Mongo
Connector ¾Í¿ÉÒÔ·Òë°Ñ_idת»»µ½id¡£
°Ñschema.xmlÎļþÖеÄ:
<code><uniqueKey>id<uniqueKey> </code> |
Ìæ»»³É
<code><uniqueKey>_id</uniqueKey> </code> |
ͬʱ»¹Òª¶¨ÒåÒ»¸ö_idµÄ×ֶΣº
<code><field name="_id" type="string" indexed="true" stored="true" /> </code> |
Æô¶¯Ê±Èç¹û±¨´í£º
<code>2014-06-18 12:30:36,648 - ERROR - OplogThread: Last entry no longer in
oplog cannot recover! Collection(Database(MongoClient('localhost', 27017), u'local'), u'oplog.rs') </code> |
Çå¿ÕE:\Users\liuzhijun\workspace\mongo-connector\mongo_connector\doc_managers\config.txtÖеÄÄÚÈÝ£¬ÐèҪɾ³ýË÷ÒýĿ¼ÏµÄÎļþÖØÐÂÆô¶¯
²âÊÔ
mongodbÖеÄÊý¾Ý±ä»¯¶¼»áͬ²½µ½solrÖÐÈ¥¡£
|