{"id":87696,"date":"2026-08-03T13:30:30","date_gmt":"2026-08-03T10:30:30","guid":{"rendered":"https:\/\/twodots.gr\/?p=87696"},"modified":"2026-08-03T13:30:31","modified_gmt":"2026-08-03T10:30:31","slug":"omegause-officeval-ai-agents-office-quality","status":"publish","type":"post","link":"https:\/\/twodots.gr\/en\/omegause-officeval-ai-agents-office-poiotita\/","title":{"rendered":"OmegaUse\u2011OfficeVal: Why AI Agents Are Fast but Not Yet Ready for the Office"},"content":{"rendered":"<div class=\"td-article-lede\">\n<p>An AI agent can generate a document in just a few minutes and at minimal cost. However, this does not mean that the document is accurate, usable, or ready to be delivered to a client. OmegaUse\u2011OfficeVal, a new open benchmark for time-consuming office tasks, aims to measure precisely this difference: not whether the model made a few correct clicks, but whether the final Word, Excel, PowerPoint, or PDF file is truly worthwhile for the person who requested it.<\/p>\n<p>For professionals, e-commerce teams, and marketers, the message is practical. Speed and low inference cost are important, but they do not, on their own, prove productivity. If a deliverable requires extensive testing, formatting corrections, or rework, part of the theoretical savings is offset by human labor costs. That is why adopting AI as a \u2019first-line solution\u00ab requires clear role boundaries and acceptance criteria, not just low execution costs.<\/p>\n<\/div>\n<div class=\"td-article-note\">\n<p><strong>Short answer:<\/strong> OmegaUse\u2011OfficeVal shows that the evaluated AI agents are cheaper and generally faster than the specific human baseline, but they lag behind in terms of the quality of the final Office artifact. For a business, the correct unit of measurement is not clicks or tokens; it is the file that opens, remains editable, meets requirements, and does not impose hidden repair time on the employee.<\/p>\n<\/div>\n<div class=\"td-article-toc\">\n<div class=\"td-toc-title\">Contents<\/div>\n<ul>\n<li><a href=\"#ti-metra-omegause-officeval\">What exactly does OmegaUse\u2011OfficeVal measure?<\/a><\/li>\n<li><a href=\"#apo-1715-protaseis-se-100-ergasies\">From 1,715 actual sentences in 100 graded assignments<\/a><\/li>\n<li><a href=\"#office-stin-praxi-polytropiko\">In practice, Office is multimodal<\/a><\/li>\n<li><a href=\"#dyo-oikonomika-simata\">Two occupational codes instead of a general occupational title<\/a><\/li>\n<li><a href=\"#christikotita-os-pyli\">Why Usability Acts as a Gateway<\/a><\/li>\n<li><a href=\"#apotelesmata-oikonomia-ochi-poiotita\">The results: the economy has not yet become a source of quality<\/a><\/li>\n<li><a href=\"#meso-score-oikonomiki-axia\">An average score and a financially significant project are not the same thing<\/a><\/li>\n<li><a href=\"#pou-emfanizontai-plireis-apotychies\">Where Do Total Failures Occur?<\/a><\/li>\n<li><a href=\"#ti-allazei-etairiko-ai-workflow\">What's Changing in the Design of a Corporate AI Workflow<\/a><\/li>\n<li><a href=\"#exi-elegchoi-office-agent-paragogi\">Six checks before an Office agent goes live<\/a><\/li>\n<li><a href=\"#sosto-symperasma-epicheiriseis-marketers\">The Right Conclusion for Businesses and Marketers<\/a><\/li>\n<\/ul>\n<\/div>\n<h2 id=\"ti-metra-omegause-officeval\">What exactly does OmegaUse\u2011OfficeVal measure?<\/h2>\n<p>The benchmark includes 100 long-running tasks, based on requests from professionals and adapted to ensure privacy. Each task provides the agent with a high-level instruction and one or more input files. The goal is to produce a final deliverable, such as a processed document, spreadsheet, presentation, or PDF.<\/p>\n<p>The choice of the final artifact as the unit of evaluation is critical. The agent can work with a GUI, scripts, APIs, or a hybrid approach. It is not graded for following a specific path, but for delivering something that is correct, accessible, editable, and useful. This is a much closer simulation of a real-world task assignment.<\/p>\n<div class=\"td-comparison\">\n<p class=\"td-comparison-title\">Four Criteria for Evaluating an Office Agent<\/p>\n<div class=\"td-comparison-cards td-comparison-cards--horizontal\">\n<div class=\"td-comparison-grid td-comparison-grid--two\">\n<div class=\"td-platform-card\">\n<h3>Success of Actions<\/h3>\n<p>It tracks whether the agent followed steps, made clicks, or completed a specified path. It is useful for debugging, but not sufficient for accepting a deliverable.<\/p>\n<div class=\"td-badge-row\"><span class=\"td-badge\">Trajectory<\/span><span class=\"td-badge\">Process<\/span><\/div>\n<\/div>\n<div class=\"td-platform-card td-platform-card--navy\">\n<h3>Artifact quality<\/h3>\n<p>It checks whether the DOCX, XLSX, PPTX, or PDF file is valid, editable, complete, and free of unwanted changes that require manual correction.<\/p>\n<div class=\"td-badge-row\"><span class=\"td-badge\">Deliverable<\/span><span class=\"td-badge\">Acceptance<\/span><\/div>\n<\/div>\n<div class=\"td-platform-card\">\n<h3>Usability Gate<\/h3>\n<p>It rejects the entire operation if a critical check fails, such as opening a file, the ability to edit it, or preserving its structure.<\/p>\n<div class=\"td-badge-row\"><span class=\"td-badge\">Gate<\/span><span class=\"td-badge\">Critical check<\/span><\/div>\n<\/div>\n<div class=\"td-platform-card td-platform-card--navy\">\n<h3>Task-completion score<\/h3>\n<p>It balances completed requirements against unwanted changes that increase the time or cost of manual repairs.<\/p>\n<div class=\"td-badge-row\"><span class=\"td-badge\">Rubric<\/span><span class=\"td-badge\">Repair burden<\/span><\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<p>On average, tasks require 2.32 hours of human labor, with a median of 2.03 hours and a maximum of 8.35 hours. So we\u2019re not talking about quick tasks like \u00abchange the color of a cell,\u00bb but rather workflows where planning, consistency, and error checking must be maintained over a long period of time.<\/p>\n<h2 id=\"apo-1715-protaseis-se-100-ergasies\">From 1,715 actual sentences in 100 graded assignments<\/h2>\n<p>The team started with 1,715 tasks proposed by professionals. After an initial review to ensure they met the criteria of a genuine operational need, a clear description, and a specific deliverable, 595 remained. Three senior experts independently assessed whether each task was sufficiently complex and long-term, yet still feasible for a human to complete. When at least two agreed, the task moved on to the next stage. This resulted in a pool of 282 candidates, from which 100 were ultimately selected.<\/p>\n<p>The instructions were depersonalized without losing their original intent, while subjective requirements that could not be reliably verified were removed. The input files were reconstructed using LLM and manually corrected for realism, privacy, copyright, and consistency. Final acceptance required unanimous approval from three senior experts.<\/p>\n<p>This process is valuable for any team building its own AI evaluation set: good tests are not just a random collection of prompts. The same principle explains why the <a href=\"https:\/\/twodots.gr\/every-eval-ever-ai-benchmarks-diafaneia\/\">AI benchmarks need transparency<\/a> about the dataset, the scoring, and the failures hidden behind an average. They require clear results, controlled inputs, the removal of sensitive data, and criteria that can be applied repeatedly.<\/p>\n<h2 id=\"office-stin-praxi-polytropiko\">In practice, Office is multimodal<\/h2>\n<p>The 100 tasks are accompanied by 220 input files and require 115 output files. The inputs include 77 images, 63 DOCX files, 31 PPTX files, 25 XLSX files, 14 PDF files, and 10 video or audio files. The outputs consist of 48 DOCX, 40 PPTX, 24 XLSX, and 3 PDF files. The distribution was not artificially balanced; it was maintained to reflect the needs identified in the dataset.<\/p>\n<p>This explains why a demonstration using plain text alone is not enough to demonstrate office automation. A real-world request may be brief, while the critical details are contained in images, videos, or embedded elements within a document. The agent must understand the content while simultaneously preserving the structure, formatting, layout, and editability.<\/p>\n<p>The taxonomy covers intents such as reformat, restructure, annotate, extract, compute, and beautify, across domains such as academic documents, education, financial data, technology, administrative tasks, and business operations. For an e-commerce or marketing department, these correspond to familiar needs: reorganizing reports, performing calculations in spreadsheets, creating presentation decks, and cleaning up material for delivery.<\/p>\n<h2 id=\"dyo-oikonomika-simata\">Two occupational codes instead of a general occupational title<\/h2>\n<p>Each task is associated with two metrics: human work time and a task price proxy. The time was measured by 20 annotators with basic office suite skills. At least two people completed each task without LLM assistance. When the times differed significantly, a third attempt was assigned. The final time was the average of the two fastest valid attempts.<\/p>\n<p>The price proxy was derived directly from practitioners for approximately 20% of the projects, when there was reliable prior outsourcing experience. For the remaining projects, three experts provided independent estimates based on project category, complexity, and typical compensation. The aggregation method limited the impact of any outlier estimate.<\/p>\n<p>The mean proxy value was $6.86 and the median was $5.11. The authors present the values in USD, converting CNY at an exchange rate of 1 CNY to 0.14609 USD, calculated as a 180-day average from January 19 to July 16, 2026. These amounts do not constitute a global labor price list; they are annotations for this specific benchmark and should be interpreted within that context.<\/p>\n<h2 id=\"christikotita-os-pyli\">Why Usability Acts as a Gateway<\/h2>\n<p>OmegaUse\u2011OfficeVal avoids both exclusively human grading and the LLM\u2011as\u2011judge approach. It uses code-based verifiers derived from detailed rubrics, reviewed by experts, and calibrated by comparing human judgments with the code\u2019s results.<\/p>\n<p>There are two dimensions. The first checks whether the artifact is valid, opens, remains editable, and can be used in practice. The second evaluates the requirements of the specific task, rewarding what has been completed and penalizing unwanted changes that increase the need for repairs.<\/p>\n<p>In total, there are 219 usability checks\u2014an average of 2.19 per assignment\u2014and 2,009 task-completion checks\u2014an average of 20.09. If even one usability check fails, the task receives a score of zero. Only a usable deliverable passes the weighted completion score. For businesses, this makes perfect sense: an \u00abalmost correct\u00bb spreadsheet that won\u2019t open is virtually worthless.<\/p>\n<h2 id=\"apotelesmata-oikonomia-ochi-poiotita\">The results: the economy has not yet become a source of quality<\/h2>\n<p>The GLM\u20115.2, Qwen3.7\u2011Plus, Kimi K2.6, DeepSeek\u2011V4\u2011Pro, and MiniMax M3 models were evaluated, using the best submission by a junior worker for each task as the human baseline. The human achieved a total score of 27.79. The best model, GLM-5.2, scored 17.91, while Qwen3.7-Plus, Kimi K2.6, DeepSeek-V4-Pro, and MiniMax M3 scored 17.51, 17.00, 14.48, and 13.82, respectively.<\/p>\n<p>On average, a human took 2.324 hours per task. The models ranged from 0.184 hours for DeepSeek-V4-Pro to 2.275 hours for MiniMax M3. The average cost per task was $6.8560 for the human proxy baseline, compared to $1.4823 for GLM\u20115.2, $0.2152 for Qwen3.7-Plus, $0.7719 for Kimi K2.6, $0.6111 for DeepSeek-V4-Pro, and $1.7572 for MiniMax M3.<\/p>\n<p>The conclusion drawn from the data is not that agents \u00abaren\u2019t worth it.\u00bb It is that their effectiveness has two distinct dimensions. They are already cheaper and generally faster, but this advantage has not translated into the quality of human-level output. Choosing based solely on token cost or latency can therefore lead to the wrong business decision.<\/p>\n<div class=\"td-chart td-chart--metrics\">\n<p class=\"td-chart-title\">Quality, time, and cost do not result in the same ranking<\/p>\n<p class=\"td-chart-intro\">The four values are taken from the main results table of the study and pertain exclusively to this specific benchmark and scaffold.<\/p>\n<div class=\"td-metric-grid\">\n<div class=\"td-metric-card\"><span class=\"td-metric-value\">27,79<\/span><strong>human score<\/strong><\/p>\n<p>The best submission by a junior worker for each task served as the human quality baseline.<\/p>\n<\/div>\n<div class=\"td-metric-card\"><span class=\"td-metric-value\">17,91<\/span><strong>best LLM score<\/strong><\/p>\n<p>GLM-5.2 had the highest average quality among the five agents evaluated.<\/p>\n<\/div>\n<div class=\"td-metric-card\"><span class=\"td-metric-value\">$0,2152<\/span><strong>lower cost<\/strong><\/p>\n<p>Qwen3.7-Plus had the lowest average inference cost per task.<\/p>\n<\/div>\n<div class=\"td-metric-card\"><span class=\"td-metric-value\">0.184 hours<\/span><strong>faster execution<\/strong><\/p>\n<p>DeepSeek-V4-Pro had the lowest average time per task.<\/p>\n<\/div>\n<\/div>\n<\/div>\n<h2 id=\"meso-score-oikonomiki-axia\">An average score and a financially significant project are not the same thing<\/h2>\n<p>The researchers report a score, a time-weighted score, and a price-weighted score. This shows not only how well a model performed on average, but also how much human time or task value its successes account for. The model with the best simple average is not necessarily the one that performs best on the most economically significant tasks.<\/p>\n<p>The GLM-5.2 had the highest simple score among the models, while the Qwen3.7-Plus had the highest time-weighted score, 34.73 versus 30.13, and a price-weighted score of 106.50 versus 93.51. These numbers do not directly predict a company\u2019s ROI, but they demonstrate that the ranking changes when the weighting of the evaluation changes.<\/p>\n<p>For a marketing team, this means that tests must reflect the actual portfolio of work. A model that makes many small corrections may have a good average, but could fail on a critical quarterly report. Conversely, a more expensive run may be worth it when it protects a high-stakes deliverable.<\/p>\n<h2 id=\"pou-emfanizontai-plireis-apotychies\">Where Do Total Failures Occur?<\/h2>\n<p>The distribution of scores shows the difference in reliability. Humans had the highest percentage of tasks scoring above 50, at 21%, and the lowest percentage of zero-score results, at 29%. GLM\u20115.2 had 14% tasks scoring above 50. Qwen3.7-Plus had the lowest zero-score rate among the LLM agents, 38%, which the authors interpret as greater resilience in producing partial progress.<\/p>\n<p>For DeepSeek-V4-Pro and MiniMax M3, the zero-score rates were 50% and 51%. Because the usability gateway filters out non-usable artifacts, these rates highlight a problem lurking behind impressive demos: the file may exist but may not constitute a reliable delivery.<\/p>\n<p>At the same time, scores declined as human work time increased, for both humans and LLMs, with a more pronounced drop observed in the models. The authors attribute this pattern to prolonged planning, detailed file manipulation, and more opportunities for structural errors to accumulate.<\/p>\n<h2 id=\"ti-allazei-etairiko-ai-workflow\">What's Changing in the Design of a Corporate AI Workflow<\/h2>\n<p>The first change is to define the deliverable before the prompt. The team must describe what \u00abopens,\u00bb \u00abremains editable,\u00bb \u00abdoes not break formulas or layout,\u00bb and \u00abmeets all requirements.\u00bb These can be converted into acceptance checks, even if there is no fully automated verifier.<\/p>\n<p>The second is the distinction between usability and completeness. First, we check whether the artifact can be used without requiring a risky repair. Then, we rate its completeness. The third is the recording of repair time. If an employee needs forty minutes to correct a low-quality AI output, that time is part of the actual cost. At this point, a clear process for <a href=\"https:\/\/twodots.gr\/ai-agents-poios-ftaiei-otan-aftomatopoiisi-apotygchanei\/\">Who takes over when an automation fails?<\/a>.<\/p>\n<p>The fourth is weighting by business value. Critical tasks should not get lost in the average. An evaluation set for e-commerce can assign greater weight to financial reports, catalog operations, or presentations sent to customers, provided that the weights are based on actual priorities rather than arbitrary assumptions.<\/p>\n<h2 id=\"exi-elegchoi-office-agent-paragogi\">Six checks before an Office agent goes live<\/h2>\n<p>The benchmark is not a ready-made procurement policy. However, it provides a practical framework for acceptance testing for teams that automate documents, spreadsheets, presentations, or PDFs.<\/p>\n<div class=\"td-step-list\">\n<p class=\"td-step-list-title\">Acceptance Checklist for Office AI Workflows<\/p>\n<ol>\n<li><span class=\"td-step-kicker\">Test 1<\/span><strong>Select actual deliverables<\/strong>\n<p>Build the evaluation set using representative DOCX, XLSX, PPTX, and PDF files from the group, after removing any personal and confidential data.<\/p>\n<\/li>\n<li><span class=\"td-step-kicker\">Check 2<\/span><strong>Here's the usability gate<\/strong>\n<p>Automatically reject files that won't open, lose their formatting or structure, can no longer be edited, or corrupt critical layouts.<\/p>\n<\/li>\n<li><span class=\"td-step-kicker\">Check 3<\/span><strong>Calculate claims and losses<\/strong>\n<p>Rate separately the items that have been completed and the unwanted changes that increase repair time or operational risk.<\/p>\n<\/li>\n<li><span class=\"td-step-kicker\">Check 4<\/span><strong>Record the repair time<\/strong>\n<p>Measure the time spent on human review, correction, and re-execution along with the inference cost, so that the cost is not artificially inflated.<\/p>\n<\/li>\n<li><span class=\"td-step-kicker\">Test 5<\/span><strong>Weigh by business value<\/strong>\n<p>Give greater weight to artifacts that affect customers, financial decisions, campaigns, or regulatory obligations.<\/p>\n<\/li>\n<li><span class=\"td-step-kicker\">Test 6<\/span><strong>Keep manual approval and regression tests<\/strong>\n<p>Require final approval by the owner for critical deliveries, and repeat the same tests whenever the model, prompt, tool, or file template changes.<\/p>\n<\/li>\n<\/ol>\n<\/div>\n<h2 id=\"sosto-symperasma-epicheiriseis-marketers\">The Right Conclusion for Businesses and Marketers<\/h2>\n<p>OmegaUse\u2011OfficeVal does not prove that a specific model will perform the same way in your environment. It tests specific models, with a specific scaffold and 100 specific tasks. However, it provides a robust measurement framework: real-world requests, final artifacts, usability gates, task-level rubrics, and economic costs.<\/p>\n<p>Today, safe adoption resembles a controlled collaboration between humans and agents more than it does a blind replacement. This complements the broader picture of <a href=\"https:\/\/twodots.gr\/ai-stin-ergasia-stin-eyropi-openai-ellinikes-epicheiriseis\/\">AI in the Workplace and Greek Businesses<\/a>, where true value depends on processes, skills, and responsible oversight. Agents can reduce time and inference cost, but the benchmark results show that final quality control remains essential, especially in time-consuming tasks where errors accumulate.<\/p>\n<p>For a business owner, the question isn\u2019t simply \u00abHow quickly did he complete the file?\u00bb It\u2019s \u00abhow much useful work did it deliver, how much rework was required, and what was the value of the work it completed?\u00bb That\u2019s where the conversation shifts from the AI demo to measurable productivity. For small teams, the same logic applies when the <a href=\"https:\/\/twodots.gr\/ai-proti-proslipsi-mikres-epicheiriseis\/\">AI is treated as a first-time hire<\/a>: Its role requires goals, boundaries, and quality control.<\/p>\n<div class=\"td-decision-band\">\n<p class=\"td-decision-label\">The decision to produce<\/p>\n<p><strong>Don't approve an Office agent just because they're fast or inexpensive.<\/strong><\/p>\n<p>Require a usable artifact, task-specific verification, recording of repair time, and a clearly designated human owner. If a file is not ready for delivery or requires extensive reconstruction, a low inference cost does not equate to productivity.<\/p>\n<\/div>\n<section class=\"td-service-cta\">\n<div class=\"td-service-cta-content\">\n<p class=\"td-service-cta-eyebrow\">Business automation and AI from TWO DOTS<\/p>\n<p class=\"td-service-cta-title\">Turn Office AI demos into controlled workflows.<\/p>\n<p>TWO DOTS designs automation solutions with acceptance criteria, human approval points, deliverable checks, and measurable repair costs for documents, reports, and recurring operations.<\/p>\n<div class=\"td-service-cta-actions\"><a class=\"td-service-cta-button\" href=\"https:\/\/twodots.gr\/aftomatismoi-epicheiriseon-ai\/\">See business automation with AI<\/a><\/div>\n<\/div>\n<\/section>\n<section id=\"sychnes-erotiseis\" class=\"td-faq-section\">\n<div class=\"td-faq\">\n<p class=\"td-faq-heading\">Frequently Asked Questions (FAQs)<\/p>\n<details class=\"td-faq-item\">\n<summary class=\"td-faq-title\">What is OmegaUse\u2011OfficeVal?;<\/summary>\n<div class=\"td-faq-content\">\n<p>It is an open benchmark featuring 100 time-consuming Office tasks, final deliverables, human time, a task price proxy, detailed rubrics, and code-based verifiers.<\/p>\n<\/div>\n<\/details>\n<details class=\"td-faq-item\">\n<summary class=\"td-faq-title\">What exactly does it evaluate?;<\/summary>\n<div class=\"td-faq-content\">\n<p>Evaluates the final DOCX, XLSX, PPTX, or PDF file as an artifact: whether it opens, remains editable, meets the requirements, and does not contain errors that increase the time required for manual correction.<\/p>\n<\/div>\n<\/details>\n<details class=\"td-faq-item\">\n<summary class=\"td-faq-title\">How large are the benchmark jobs?;<\/summary>\n<div class=\"td-faq-content\">\n<p>The mean is 2.32 hours, the median is 2.03 hours, and the maximum is 8.35 hours.<\/p>\n<\/div>\n<\/details>\n<details class=\"td-faq-item\">\n<summary class=\"td-faq-title\">Why does an entire project fail a usability check?;<\/summary>\n<div class=\"td-faq-content\">\n<p>Because the benchmark treats usability as a gateway. If the artifact does not open, cannot be edited, or fails another critical check, it is not considered a usable deliverable.<\/p>\n<\/div>\n<\/details>\n<details class=\"td-faq-item\">\n<summary class=\"td-faq-title\">Were the AI agents cheaper than the human baseline?;<\/summary>\n<div class=\"td-faq-content\">\n<p>Yes, in this particular setup, all the evaluated agents were cheaper and generally faster, but they fell short in terms of the quality of the output.<\/p>\n<\/div>\n<\/details>\n<details class=\"td-faq-item\">\n<summary class=\"td-faq-title\">Which model had the best overall quality?;<\/summary>\n<div class=\"td-faq-content\">\n<p>The GLM\u20115.2 had the highest simple score among the models, 17.91, compared to 27.79 for the human baseline.<\/p>\n<\/div>\n<\/details>\n<details class=\"td-faq-item\">\n<summary class=\"td-faq-title\">Can a company use these scores as its own ROI?;<\/summary>\n<div class=\"td-faq-content\">\n<p>Not directly. The scores apply to specific tasks, models, and scaffolds. A company must adapt the methodology to its own data, costs, risks, and repair times.<\/p>\n<\/div>\n<\/details>\n<details class=\"td-faq-item\">\n<summary class=\"td-faq-title\">What is the basic acceptance test for an Office agent?;<\/summary>\n<div class=\"td-faq-content\">\n<p>First, the file is checked to ensure it is usable and safe to process. Then, its completeness, unwanted changes, repair time, and the business value of the work are assessed.<\/p>\n<\/div>\n<\/details>\n<\/div>\n<\/section>\n<div class=\"td-source-list\">\n<p id=\"piges\" class=\"td-source-list-title\">Sources<\/p>\n<ul>\n<li><a href=\"https:\/\/arxiv.org\/abs\/2607.27155\" target=\"_blank\" rel=\"noopener\">Zhou et al.: OmegaUse\u2011OfficeVal \u2014 Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding<\/a><\/li>\n<li><a href=\"https:\/\/omegause-officeval.github.io\/\" target=\"_blank\" rel=\"noopener\">Official OmegaUse\u2011OfficeVal page with a leaderboard, analysis, and resources<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/baidu-frontier-research\/OmegaUse-OfficeVal\" target=\"_blank\" rel=\"noopener\">Official source code and documentation for OmegaUse\u2011OfficeVal<\/a><\/li>\n<\/ul>\n<\/div>","protected":false},"excerpt":{"rendered":"<p>OmegaUse\u2011OfficeVal measures quality, time, and cost across 100 Office tasks, demonstrating why AI agents require usability checks and human oversight.<\/p>","protected":false},"author":1,"featured_media":87782,"comment_status":"open","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"_gspb_post_css":"","content-type":"","footnotes":""},"categories":[199],"tags":[6506,19002,19001,19000,18999],"class_list":["post-87696","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-techniti-noimosyni","tag-ai-agents","tag-ai-axiologisi","tag-business-productivity","tag-llm-benchmarks","tag-office-automation"],"blocksy_meta":{"styles_descriptor":{"styles":{"desktop":"","tablet":"","mobile":""},"google_fonts":[],"version":7}},"_links":{"self":[{"href":"https:\/\/twodots.gr\/en\/wp-json\/wp\/v2\/posts\/87696","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/twodots.gr\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/twodots.gr\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/twodots.gr\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/twodots.gr\/en\/wp-json\/wp\/v2\/comments?post=87696"}],"version-history":[{"count":0,"href":"https:\/\/twodots.gr\/en\/wp-json\/wp\/v2\/posts\/87696\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/twodots.gr\/en\/wp-json\/wp\/v2\/media\/87782"}],"wp:attachment":[{"href":"https:\/\/twodots.gr\/en\/wp-json\/wp\/v2\/media?parent=87696"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/twodots.gr\/en\/wp-json\/wp\/v2\/categories?post=87696"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/twodots.gr\/en\/wp-json\/wp\/v2\/tags?post=87696"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}