Kadyrm үчүн колдонуучунун салымдары
Навигацияга өтүү
Издөөгө өтүү
28 Август (Баш оона) 2026
- 10:4810:48, 28 Август (Баш оона) 2026 айырма тарыхы +12 Колдонуучу:KETC Parser for WikiMediaExport →AI text processing tools overview учурдагы
- 10:4610:46, 28 Август (Баш оона) 2026 айырма тарыхы +35 Калып:NestedList No edit summary учурдагы
- 10:4410:44, 28 Август (Баш оона) 2026 айырма тарыхы +79 Колдонуучу:KETC Parser for WikiMediaExport →AI text processing tools overview
- 10:4110:41, 28 Август (Баш оона) 2026 айырма тарыхы −126 Калып:NestedList No edit summary
- 10:2910:29, 28 Август (Баш оона) 2026 айырма тарыхы −13 Калып:NestedList No edit summary
- 10:2410:24, 28 Август (Баш оона) 2026 айырма тарыхы +227 Ж Калып:NestedList Created page with "<ul> <li>{{{item1}}} {{#if:{{{subitem1|}}}| <ul> <li>{{{subitem1}}}</li> {{#if:{{{subitem2|}}}|<li>{{{subitem2}}}</li>}} </ul> |}} </li> {{#if:{{{item2|}}}|<li>{{{item2}}}</li>|}} </ul>"
- 10:0410:04, 28 Август (Баш оона) 2026 айырма тарыхы +7 Колдонуучу:KETC Parser for WikiMediaExport →UP: which AI tool would you recommend to process a formatted text in docx or odt format
- 09:2909:29, 28 Август (Баш оона) 2026 айырма тарыхы +37 Колдонуучу:KETC Parser for WikiMediaExport →Classification for Parsing
- 09:2609:26, 28 Август (Баш оона) 2026 айырма тарыхы +64 Колдонуучу:KETC Parser for WikiMediaExport /* UP: there are 3 types of p/strong elements: 1) containing all capital unicode chars like ФЫАЫФА. this is actually an article title; 2) first letter capital like Авааф фыаа. This is a section header inside an article; 3) all letters small like выфафыа. This is just an inline bold text having no structural load. As a long term goal I need to classify all of them in my html document. As a specific quetion: can I use regex pattern inside XPath expression to match all capitals and first letter...
- 09:2209:22, 28 Август (Баш оона) 2026 айырма тарыхы +33 Колдонуучу:KETC Parser for WikiMediaExport →UP: still removes other imgs. #problematic;
- 09:2209:22, 28 Август (Баш оона) 2026 айырма тарыхы +15 Колдонуучу:KETC Parser for WikiMediaExport →UP: still removes other imgs
- 09:1809:18, 28 Август (Баш оона) 2026 айырма тарыхы +2 Колдонуучу:KETC Parser for WikiMediaExport →Gamechanger: Mammoth lib
- 09:0109:01, 28 Август (Баш оона) 2026 айырма тарыхы +56 Колдонуучу:KETC Parser for WikiMediaExport →Using Python (mammoth)
- 08:5908:59, 28 Август (Баш оона) 2026 айырма тарыхы +13 Колдонуучу:KETC Parser for WikiMediaExport →Using pandoc (The Gold Standard)
- 08:5808:58, 28 Август (Баш оона) 2026 айырма тарыхы −10 Колдонуучу:KETC Parser for WikiMediaExport →markup vs app/api/ftext considerations, and markup sanitizers coming to the stage
- 08:5708:57, 28 Август (Баш оона) 2026 айырма тарыхы +21 Колдонуучу:KETC Parser for WikiMediaExport →MD considerations, and markup sanitizers coming to the stage
- 08:5508:55, 28 Август (Баш оона) 2026 айырма тарыхы +81 Колдонуучу:KETC Parser for WikiMediaExport →Recommendation
- 08:5208:52, 28 Август (Баш оона) 2026 айырма тарыхы +1 Колдонуучу:KETC Parser for WikiMediaExport →UP: are u sure beautiful soup or mammoth can handle html like this:
- 08:5108:51, 28 Август (Баш оона) 2026 айырма тарыхы +3 Колдонуучу:KETC Parser for WikiMediaExport →3. The "Hybrid" AI-Assisted Shortcut
- 08:5008:50, 28 Август (Баш оона) 2026 айырма тарыхы +60 Колдонуучу:KETC Parser for WikiMediaExport →UP: will mammoth produce me a sanitized html of the docx file keeping formatting intact?
- 08:3908:39, 28 Август (Баш оона) 2026 айырма тарыхы +10 Колдонуучу:KETC Parser for WikiMediaExport →MD considerations, and mammoth coming to the stage
- 08:3308:33, 28 Август (Баш оона) 2026 айырма тарыхы +89 Колдонуучу:KETC Parser for WikiMediaExport →UP: So, processing docx files would be cumbersome for Google AI Studio. How to make transition to plain text then without loosing formatting data?
- 08:2908:29, 28 Август (Баш оона) 2026 айырма тарыхы +79 Колдонуучу:KETC Parser for WikiMediaExport →UP: Can it add tags inside docx file , I.e modify it?
- 08:2208:22, 28 Август (Баш оона) 2026 айырма тарыхы +54 Колдонуучу:KETC Parser for WikiMediaExport →UP: Would it be appropriate to use Google AI Studio for my project
- 06:3506:35, 28 Август (Баш оона) 2026 айырма тарыхы +12 Колдонуучу:KETC Parser for WikiMediaExport →MediaWiki Export Schema
- 06:3306:33, 28 Август (Баш оона) 2026 айырма тарыхы +122 Колдонуучу:KETC Parser for WikiMediaExport →UP: Yes
- 06:3206:32, 28 Август (Баш оона) 2026 айырма тарыхы +51 Колдонуучу:KETC Parser for WikiMediaExport →UP: So, build me a holistic recipe to deal with such encyclopedic formatted text files each containing a series of a series of articles in Natural language
- 06:3006:30, 28 Август (Баш оона) 2026 айырма тарыхы +70 Колдонуучу:KETC Parser for WikiMediaExport →UP: Actually the final xml structure is Mediawiki Export xml file which I get from simple conversion from StarDict xml. So the initial text can be directly annotated to that xml schema. The main tag is mediawiki. Then come a series of pages entries. In my case there is only one initial revision, as far as i remember in text tag. Bold and italic fonts remain in html. Section titles embraced in double equations.
- 06:2806:28, 28 Август (Баш оона) 2026 айырма тарыхы −2 Колдонуучу:KETC Parser for WikiMediaExport →StarDict Schema
- 06:2806:28, 28 Август (Баш оона) 2026 айырма тарыхы +3 Колдонуучу:KETC Parser for WikiMediaExport →Suggested Workflow for Large-Scale Annotation
- 06:2606:26, 28 Август (Баш оона) 2026 айырма тарыхы +55 Колдонуучу:KETC Parser for WikiMediaExport →UP: I'm using StarDict xml schema
- 06:2106:21, 28 Август (Баш оона) 2026 айырма тарыхы +7 857 Колдонуучу:KETC Parser for WikiMediaExport →Key Things to Know
- 06:1706:17, 28 Август (Баш оона) 2026 айырма тарыхы −7 823 Колдонуучу:KETC Parser for WikiMediaExport →Appendix: Garbage Prompts
- 06:1506:15, 28 Август (Баш оона) 2026 айырма тарыхы +61 Колдонуучу:KETC Parser for WikiMediaExport →UP: Not email but xml format
- 06:1206:12, 28 Август (Баш оона) 2026 айырма тарыхы +27 Колдонуучу:KETC Parser for WikiMediaExport →AI text processing tools overview
- 06:0206:02, 28 Август (Баш оона) 2026 айырма тарыхы +38 Колдонуучу:KETC Parser for WikiMediaExport →UP: which AI tool would you recommend to process a formatted text in docx or odt format
- 06:0006:00, 28 Август (Баш оона) 2026 айырма тарыхы +7 Колдонуучу:KETC Parser for WikiMediaExport →UP: I want xslt transform my html file. There is a div.imageCaption>p and neighboring (right before or after) p>img. My task is to move the img inside the div.imageCaption
- 05:5705:57, 28 Август (Баш оона) 2026 айырма тарыхы −10 Колдонуучу:KETC Parser for WikiMediaExport →UP: which AI tool would you recommend to process a formatted text in docx format or Oddity format Белги: Көрсөтмөлүү оңдоо
- 05:5305:53, 28 Август (Баш оона) 2026 айырма тарыхы +1 Колдонуучу:KETC Parser for WikiMediaExport /* UP: I have updated my dir tree based on your recommendation: root ── 4. Processor [parsing] ├── 1. strong text classifier │ ├── input │ │ └── 2-50 [V5][styles]_[ImgCapt]_clean.html │ ├── output │ │ ├── 2-13 v2 [vol5][imgC]_[ImgCapt]_transformed.html │ │ └── 2-50 v3 [styles][ImgCapt][cleanHtml][strongClassified].html │ └── src │ ├── strong_text_classifier.py │ └── strong_text_classifier_v2.py ├── 2. inline_title_anomaly_f...
- 05:5105:51, 28 Август (Баш оона) 2026 айырма тарыхы +1 Колдонуучу:KETC Parser for WikiMediaExport /* UP: are u sure beautiful soup or mammoth can handle html like this: ol{margin:0;padding:0}table td,table th{padding:0}.c5{color:#000000;font-weight:400;text-decoration:none;vertical-align:baseline;font-size:20pt;font-family:"Arial";font-style:normal}.c4{color:#000000;font-weight:400;text-decoration:none;vertical-align:sub;font-size:11pt;font-family:"Arial";font-style:italic}.c6{color:#000000;font-weight:700;text-decoration:none;vertical-align:baseline;font-size:11pt;font-family:"Arial";fon...
- 05:5005:50, 28 Август (Баш оона) 2026 айырма тарыхы +435 338 Колдонуучу:KETC Parser for WikiMediaExport No edit summary Белги: Көрсөтмөлүү оңдоо
- 05:0405:04, 28 Август (Баш оона) 2026 айырма тарыхы +1 943 Ж Колдонуучу:KETC Parser for WikiMediaExport Created page with "==Category Appending== def tokenize_pages(tree, output_image_dir, page_tag="page", title_tag="title", text_wrapper_tag="text"): body = tree.find("body") if body is None: return tree children = list(body) new_body_elements = [] current_page = None current_revision = None current_text = None current_title_str = "" # Prompt the user ONCE for the whole document before looping through pages global_category = p..."
27 Август (Баш оона) 2026
- 17:1617:16, 27 Август (Баш оона) 2026 айырма тарыхы 0 м ЛЕМ 1 версия учурдагы
- 17:1617:16, 27 Август (Баш оона) 2026 айырма тарыхы 0 м ЛЕЛЕВКИН 1 версия учурдагы
- 17:1617:16, 27 Август (Баш оона) 2026 айырма тарыхы 0 м ЛЕКЦИЯ 1 версия учурдагы
- 17:1617:16, 27 Август (Баш оона) 2026 айырма тарыхы 0 м ЛЕКСИКОН 1 версия учурдагы
- 17:1617:16, 27 Август (Баш оона) 2026 айырма тарыхы 0 м ЛЕКСИКОЛОГИЯ 1 версия учурдагы
- 17:1617:16, 27 Август (Баш оона) 2026 айырма тарыхы 0 м ЛЕКСИКОГ-РАФИЯ 1 версия учурдагы
- 17:1617:16, 27 Август (Баш оона) 2026 айырма тарыхы 0 м ЛЕКСИКАЛЫК МААНИ 1 версия учурдагы
- 17:1617:16, 27 Август (Баш оона) 2026 айырма тарыхы 0 м ЛЕКСИКАЛАШУУ 1 версия учурдагы