# Sisyphus Consulting > We build and run your technical infrastructure: cloud infrastructure, web development, and custom software for small and medium businesses. Last Updated: 2026-08-11T12:17:08+00:00 ## Pages ### About URL: https://sisyphusconsulting.org/about/ Description: Learn about Sisyphus Consulting, a technology consultancy. We build and run your technical infrastructure with a high degree of ownership. Discover our approach. Sisyphus Consulting is a technology consultancy focused on cloud infrastructure, custom software, and web development for small and medium businesses. First, the obvious question. What’s with the name Sisyphus? In Greek mythology, Sisyphus was a clever king who tricked the gods, so they punished him. He was forced to push a huge boulder up a hill, but it always rolled back down, making him repeat the task forever. An eternal punishment. This story has multiple meanings. But we resonate with Albert Camus’ interpretation in his essay Myth of Sisyphus: “The struggle itself toward the heights is enough to fill a man’s heart. One must imagine Sisyphus happy.” Like Sisyphus, we scale the mountain of challenges for our clients. Sometimes the effort seems futile, but it is enough to fill our heart and spirit. We continue pushing the boulder up the hill, again and again. Why people hire us We provide lasting and contextual solutions to our clients. We do it with a high degree of ownership. But that’s not enough. The main reason is trust. They trust our process. They know they have made the right decision. Our reputation precedes us and we are known to leave money on the table for the right reasons. That’s why when we show our willingness to take them up as a client, they know we mean it. In addition to the said reasons, people like our approach of keeping things simple. This is particularly remarkable in an industry that has an incentive to keep things opaque and complicated. How people find us Since we don’t bid for work, you might wonder how people find us. Most people find us through referrals. When anyone asks our clients or ex-colleagues about technical infrastructure or custom software, they point them in our direction. Then there are people who are long-time subscribers of our writing. Whenever they need technical work done, they reach out to us. A good fit matters to us If you think about it, when it comes to clothes, their brand, price or material don’t matter as much as the their fit. Similarly, we prefer working with clients who understand our thinking, get excited at trying new things and practice delegation. They’re a good fit for us and we’re good fit for them. Having said that, if you’re looking for the cheapest service possible, you shouldn’t consider us. We are known to provide good returns on investment through our services—and that’s what matters economically. What we don’t do We don’t work for free We don’t say yes easily, and that includes declining most meeting requests. We follow a maker’s schedule, and our clients understand that by protecting our work time, we are ultimately safeguarding their interests. We don’t bid for work or write proposal documents. Once we have had a verbal agreement on the work and fees, we put everything in a document for our shared understanding and reference. That’s as much close as it comes to a traditional proposal; the only difference being that it would be an agreement document and not a proposal document. We believe that free proposals and bids are some of the flawed and unethical practices of the industry. We want no part in perpetuating them. Our setup Sisyphus Consulting is led by Bhagyesh Pathak. We bring in specialists — in Cloud Infrastructure, Custom Software, and Web Development — as individual projects need them, matching the right expertise to the right problem instead of billing you for capacity you don’t need. This model works because of how we choose engagements: fewer clients, deeper involvement, and no interest in projects that don’t fit. What’s next See our case studies for real work for real clients. See our consultancy services for what we do and how we work. Or tell us what’s going on. We’ll review and follow up to set up a conversation. Tell us about your project → We'll follow up to find a time to talk --- ### Case Studies URL: https://sisyphusconsulting.org/case-studies/ Description: Real work for real clients: cloud infrastructure, custom software, and web development engagements. Selected work. The problems, the approach, and what we delivered. --- ### Cloud Infrastructure URL: https://sisyphusconsulting.org/cloud-infrastructure/ Description: We set up, secure, and maintain the cloud infrastructure your business runs on: server provisioning, hardening, backups, and Cloudflare CDN. We set up, secure, and maintain the cloud infrastructure your business runs on. Most small and medium businesses don’t need an in-house IT team. They need someone who can provision a server, lock it down, keep it running, back it up, and front it with a CDN. That’s what we do. What we handle Server provisioning: we pick the size and region and set it up. Security hardening: firewalls, SSH, fail2ban, regular audits. Locked down before it goes live. System maintenance: OS updates, monitoring, troubleshooting. Backups: offline plus object storage, with retention policies. We restore from a tested backup. Cloudflare: CDN, DNS, DDoS protection, SSL. You own it. We run it. The infrastructure is yours. We provision it in your account, or in ours for simpler operations. Your call. Either way, you retain full access, full data, and the right to move everything to your own account at any time. No vendor lock-in, and no absolute dependency on us. Your business stays in your hands. Beyond the basics: Running your software Once your servers are secured and running smoothly, you might want to bypass expensive SaaS subscriptions and run your own tools. Self-Hosted Software: running on your infrastructure We install, configure, and maintain open-source software on your new infrastructure. You pick the tools. You keep the data. We handle the setup, security, updates, and backups. See the full list of software we support → Continuous care and proactive maintenance Cloud infrastructure needs active care. Servers need updates, security needs monitoring, and backups need testing. We work on a monthly retainer. Pricing depends on the number of servers, the setup, and how fast you need a response when something goes wrong. We’ll give you specifics once we understand your environment. Want us to look after your infrastructure? Want us to look after your infrastructure? Tell us what you’re working on. We’ll review and follow up to set up a conversation. Tell us about your project → We'll follow up to find a time to talk This service is governed by our terms and conditions, you may want to have a look at them. Got any questions? Write to usClick to copy email PS: Considering a SaaS subscription? See our Self-hosted Software service before you commit. Some of our clients moved off SaaS and never went back. See our other services Custom Software Custom software built around your real processes. Diagnostic-first methodology. Learn more → Web Development Lightning-fast, zero-maintenance websites. You own your site, no platform lock-in. Learn more → --- ### Consultancy URL: https://sisyphusconsulting.org/consultancy/ Description: We build and run your technical infrastructure: cloud infrastructure, custom software, and web development for small and medium businesses. From servers to software to your website: we handle the technology layer of your business. What makes us different First, we don’t try to do everything. We focus on a few things we are good at and we try to get better at them. Second, we spend a lot of time and effort understanding your business challenges before we recommend anything. That’s the actual work, figuring out what your problem really is rather than jumping to a solution. This makes our thinking the highest value product we can provide to you. Third, our work is consultative. We work with you directly to define strategy and implement it, through a simple, collaborative way of working. Most firms wait for you to bring them a defined problem and a chosen solution. We don’t. Now, let’s talk about our services. Our Services We specialize in three areas: servers, software and websites. Cloud Infrastructure Server provisioning, security hardening, automated backups, and Cloudflare CDN. Includes installing and maintaining open-source software on your infrastructure. More info & get started → Custom Software Custom software built around your real processes. Our methodology is diagnostic-first: we interview, shadow, and map your workflows before we write a line of code. More info & get started → Web Development Lightning-fast, responsive, zero-maintenance websites. No monthly rent, no platform lock-ins. You own your site: it's an asset, not a liability. More info & get started → Not sure where to start? Tell us what you’re working on. We’ll review and follow up to set up a conversation. Tell us about your project → We'll follow up to find a time to talk --- ### Contact URL: https://sisyphusconsulting.org/contact/ Description: Get in touch with Sisyphus Consulting. Find our office address, phone number, email, and location on Google Maps in Ahmedabad. We'd love to hear from you. Here's how you can reach us. Email: contact@sisyphusconsulting.orgClick to copy email Phone: +91 93274 09270 Office Address Ground Floor-29, Vijay Plaza, Maharshi Dayanand Marg, opposite Eka Club, Kankaria, Ahmedabad, Gujarat 380022 Location Office Hours Monday through Friday, 9:30 AM to 6:30 PM IST. --- ### Custom Software URL: https://sisyphusconsulting.org/custom-software/ Description: We build custom software around your real business processes: diagnostic-first methodology, no lock-in, full ownership. For most businesses, their collection of Excel sheets or generic software acts as the whole-and-sole data system. But you and I both know that true business software is more than that. It is a reflection of how your business actually operates. That’s why when we work with you, we start from the beginning. And slowly build from there. Instead of writing a long essay on how we do it, let me share our process in brief. So that you can quickly go through it and familiarize yourself with. Here’s the process we follow: 1. Current system diagnosis We start by examining your existing systems. Ask you questions. Try to understand how one piece of information connects to the other. Who needs what information to make what decision at what frequency. We avoid generic surveys and focus on deep, meaningful interactions. That’s why, we carry out one-on-one interviews, visits, and shadowing. Challenging the status quo is an essential part of our diagnostic process. One more thing, this step is inherently diagnostic. But right away, most clients start seeing holes in their processes. And they amend their processes then-and-there. 2. Process mapping After understanding your processes and challenges, we map out your information flow, processes and systems on a paper. We review this articulation with your team and ask, “Is this really how your business operates?” If it doesn’t align the reality or expectation, your team addresses those gaps. By the end of this step, we would have developed a shared understanding of: How the different parts of the system fit with one another Role and responsibilities within the system What information different people need to make decisions Business processes in a nutshell 3. Details gathering After completing the macro-level process mapping, we move to the micro level. So, we take out our fine-tooth comb, gather all data elements associated with each process. We review them carefully and run by you for clarifications and edits. Often, we end up identifying gaps or duplications in the current system. For example, if customer satisfaction scores are tracked but not linked to specific staff members, we highlight the need of such a linkage. 4. Database design Once we are clear with macro and micro details of your business processes, we design the database structure. At this stage, most standard processes are clear and there is minimal back-and-forth with your team. 5. Development phase We begin developing software system. We group the relevant parts of the software elements together and work on them. This is different than the usual practice of working on everything at once. For example, we might start working on customer database and creating entry-level forms first. But keep analytics and dashboards work for later. This way, after each iteration we will have something that works than many things that don’t work. (Well-articulated by Ryan Singer in Shape Up) 6. Handover Once the software is ready and you have used it to your satisfaction, we begin the handover process. Our work follows the philosophy of customer empowerment. Which means we aim to reduce client’s dependency on us after our engagement is over. This philosophy drives our decisions about choice of technology and design of our work. Want us to build custom software for you? Tell us what you’re working on. We’ll review and follow up to set up a conversation. Tell us about your project → We'll follow up to find a time to talk "In our work at Snehalaya CCI, Bhagyesh's impact went beyond creating just another IT solution. He invested time to understand the unique challenges of child development before designing a system that effectively monitors each child's journey toward independence. What stands out is his pragmatic approach: using simple, accessible technology tools rather than complex solutions, while fully meeting our needs. [...] truly serves our mission of nurturing self-reliant individuals." Mahesh Rasal, Co-founder,  Sachet Foundation This service is governed by our terms and conditions, you may want to have a look at them. Got any questions? Write to usClick to copy email PS: Not looking for a complete custom software build? You may find our Self-hosted Software service a perfect fit. See our other services Cloud Infrastructure Server provisioning, security hardening, backups, and Cloudflare CDN. You own the data, we keep the lights on. Learn more → Web Development Lightning-fast, zero-maintenance websites. You own your site, no platform lock-in. Learn more → --- ### How we work URL: https://sisyphusconsulting.org/how-we-work/ Description: Learn about our work style and how we use Basecamp to manage projects, communicate effectively, and deliver results for our clients. I thought it would be a good idea to share how we work internally as well as with clients. It sets the right expectations from the outset. First let’s get the question of tool out of the way. We use Basecamp Our documents, discussions, to-dos, updates, messages, etc live in Basecamp, a tried-and-tested project management tool. It is surprisingly simple to work with. It simply doesn’t get in our way of working calmly. Our Basecamp Homepage Screenshot Every client is a project We create a dedicated Basecamp project for every client and onboard them. It looks like the screenshot provided below. You’ll notice that there are six main tools visible in this particular project: Discussions, To-dos, Docs & Files, Schedule, Imp Email Copies and Group Chat. A client's workspace in Basecamp Discussions: Traditionally, people use emails to discuss things. In the last decade, those emails transformed into email threads...and there is no way out of this web of threads. Not only we get entangled with 100 emails in a thread but also with time, emails get buried and we lose the whole thread. To avoid all these, we use Discussions tool in Basecamp. Just create a topic, write a note, add images, links and open it for relevant people to discuss. Simple. Everything stays in the same place. To-dos: We use this tool extensively. For every task, we create dedicated to-do lists within the project. It guides our work and provides an overall idea to the client about how we are planning and where we are in our execution. Docs & Files: This is where we keep our all documents and files related to this project, i.e. this client. Organized in subfolders and ready to refer at a moment’s notice. We also link our Google Drive folder in this tool, so that anyone can access them through this tool. Schedule: Since we use Basecamp as our main workspace, any meeting or to-do’s deadline is visible via its Schedule tool. This allows clients to cut through all the noise of their Google Calendars and keep a tab on the meetings or to-dos that matter. Imp Email Copies: As explained under the “discussions” point, we rarely use email. But still, sometimes we have to deal with a few emails. Basecamp provides us an extremely useful feature of forwarding emails to projects. For example, this project’s auto-generated email-ID is save-ZcF4DHHxxxxx@3.basecamp.com Anyone sending an email at this id will land under this project. We can respond to the emails from this tool, refer to it, discuss it, etc. Mostly, we use it for email archival. Group Chat: Our work do not need us to use any instant messaging apps such as WhatsApp or Slack. If there is anything worth discussing, we write up at length in discussion topic. But still, if there is something that needs the whole team’s attention without much fanfare, we use the Group Chat tool. Depending on the need of the work, we also use a few other tools such as Card Table. But those are just features We did not intend to make it seem like we are selling Basecamp. We are pointing at something else: Basecamp’s design supports and encourages our work style. This is something you might be interested in hearing about. By virtue of the nature of our work, we work remotely. Remote work requires a different set of habits and communication than that of the in-person ones. Full-context communication We believe in providing full context while communicating with one another. That also means that we may use 100 words instead of 10. The additional words may contain a hyperlink to a file/discussion/to-do/schedule/card/date. Providing full context is not possible on Instant Messaging or Hyper-Instant Messaging services such as WhatsApp, Slack or their clones. Also, isn’t thinking and communicating one sentence at a time is a really bad idea to communicate anything of value? Primarily asynchronous, occasionally synchronous Our work is creative. We need a big, undisturbed chunk of time to ourselves during the day to focus on our client’s work. That’s why we have a policy to be reasonable in expecting a response from the other person. We leave our comment or a private ping on Basecamp and let the other person respond when it suits them. If there is something that can’t wait, the two people can get on a call and figure things out. This work style automatically eliminates usage of instant messaging tools and encourages planning and full-context provision from everyone. Something that works is better than many things that don’t work Whether it is developing software or creating a design document, it almost always pays to focus one slice at a time. We group relevant tasks in their simplest possible, most-common-sense-making forms and carry out our work. This approach is in stark contrast to what most firms in our business do: treating each project as a single piece of pie and biting it from all directions at the same time. No wonder, nothing gets done. At least not in the prescribed budget and timeline. Through our approach, after each iteration we have something that works than many things that don’t work. (Well-articulated by Ryan Singer in Shape Up) How we share project progress Once on-boarded on Basecamp, the client can see everything (okay, not exactly everything, we keep work-in-progress things to ourselves until they’re ready for the client’s eyes) that is happening in the project. We understand that not everyone has time to keep a tab on the work progress. So, we provide a comprehensive project update once in a while using Basecamp’s needle feature. For example, look at this progress communication where everything is “on track”. Or when the progress was a bit “concerned”. Work is documentation At the end of an engagement, clients and us—both parties feel the need to preserve and document everything that took place. It is a reasonable expectation. Once more, this is where our practice of providing full-context and asynchronous working comes in handy. The Basecamp project becomes a live documentation. Every decision, their associated context and timeline are stored in the project. So, in the end, we handover a working, offline copy of the project to our clients for their record. --- ### Privacy Policy URL: https://sisyphusconsulting.org/privacy/ Description: Read our privacy policy to understand how we handle your data. We use GoatCounter for analytics and Substack for newsletters. Web Analytics Our website employs GoatCounter to track analytics on visitors. GoatCounter is an open source web analytics platform, that attempts to minimize data collection, and abide by GDPR. GoatCounter doesn’t track users with unique identifiers and doesn’t need a GDPR notice. Please see GoatCounters privacy policy for more information. Newsletter Analytics We use Substack for our email newsletters. If you sign up for our newsletters, Substack will process your data as per their privacy policy. Information We Collect We collect project-related information through questionnaires. When you make payments on our website, our merchant Razorpay processes your data according to their practices outlined in their privacy policy. We only collect the details provided by you on the payment form. Information We Do Not Collect Our website does not collect any personal information from visitors. As mentioned earlier, we use GoatCounter, a privacy-friendly web analytics tool that does not use cookies or track individuals across sites. It collects only anonymous usage data, such as page views and referrer information, without storing personal details. Your browsing experience on our site is completely anonymous. External Links Our website may contain links to third-party websites. We are not responsible for the privacy practices or content of these websites. Please review their privacy policies if you choose to visit these external sites. Contact Us If you have any questions about this privacy policy, you can contact us at id: contact[at]sisyphusconsulting[dot]org --- ### Products URL: https://sisyphusconsulting.org/products/ Description: Products built by Sisyphus Consulting. We build products that solve problems we have experienced ourselves or seen in our work. Here is what we have released so far. Wu Wei Cards A 50-card deck for anyone who guides conversations: therapists, coaches, facilitators, educators, and team leaders. Handmade illustrations on one side, emotion keywords on the other. Open-ended by design, because emotional work does not follow a script. The deck ships with a 30-page companion guide, innovative digital tools, and Wu Wei Planner, an AI companion we built for session planning and other Wu Wei-related tasks. Visit Wu Wei Cards → FeedHammer We run our day-to-day work on Basecamp. Every project has its own calendar, which makes getting a unified view, or sharing the right dates with the right people, harder than it should be. FeedHammer merges multiple Basecamp project calendars into a single, shareable iCal feed. You decide which events from which projects go into each feed, so your team sees everything they need while clients see only what is relevant. Once configured, it stays synced with Google Calendar, Outlook, or Apple Calendar. Visit FeedHammer → Wu Wei MCP Server (Free & Open Source) An MCP (Model Context Protocol) server that lets anyone use Wu Wei Cards directly inside Claude Desktop, free of charge. Browse all 50 cards, explore themes, get facilitation prompts, and display card images through natural conversation with Claude. Set it up in Claude Desktop → --- ### Self-Hosted Software URL: https://sisyphusconsulting.org/self-hosted-software/ Description: We set up and maintain open-source software on your own cloud server. You own the data, we handle the infrastructure. We set up and maintain open-source software on your own cloud server. You pick the software. You own the server, the data, and the software. We handle the technical work to keep it running. These are the platforms we support at the moment. If you have something else in mind, let us know. All Community Learning Project Mgmt Scheduling File Storage Documents Finance CRM Analytics Automation CMS Forms & Surveys Support Security Marketing AI & Agents Discourse Modern community forums. vs Circle, Facebook Groups Since 2013 Copy prompt for AI evaluation Mattermost Self-hosted team chat. vs Slack Since 2015 Copy prompt for AI evaluation Jitsi Meet Secure, browser-based video conferencing. vs Zoom, Google Meet, Teams Since 2003 Copy prompt for AI evaluation Moodle Learning management system, used by universities and training organizations. vs Teachable, Thinkific Since 2002 Copy prompt for AI evaluation BookStack Simple team knowledge base and documentation. vs Notion, Confluence Since 2015 Copy prompt for AI evaluation Fizzy Kanban project management, simple and visual. vs Trello, Asana Since 2018 Copy prompt for AI evaluation OpenProject Project management with Gantt charts, time tracking, and budgets. vs MS Project, Asana Since 2012 Copy prompt for AI evaluation Kimai Time tracking for teams, projects, and client billing. vs Toggl, Harvest Since 2006 Copy prompt for AI evaluation Cal.com Appointment scheduling and booking pages. vs Calendly, Acuity Since 2021 Copy prompt for AI evaluation Nextcloud File sync, sharing, and collaboration platform. vs Google Drive, Dropbox Since 2016 Copy prompt for AI evaluation Paperless-ngx Scan, index, and search your paper documents digitally. vs Google Drive, SharePoint, DocuWare Since 2021 Copy prompt for AI evaluation InvoiceNinja Invoicing, expenses, and client payments. vs FreshBooks, QuickBooks Since 2014 Copy prompt for AI evaluation SuiteCRM Full-featured CRM for sales, marketing, and support. vs Salesforce, HubSpot Since 2013 Copy prompt for AI evaluation Mautic Marketing automation and email campaigns. vs Mailchimp, HubSpot Since 2014 Copy prompt for AI evaluation Metabase Business intelligence and dashboards. vs Tableau, Looker Since 2015 Copy prompt for AI evaluation Plausible Privacy-friendly web analytics. vs Google Analytics Since 2018 Copy prompt for AI evaluation n8n Workflow automation connecting your apps and services. vs Zapier, Make Since 2019 Copy prompt for AI evaluation WordPress The world's most popular CMS. vs Squarespace, Wix Since 2003 Copy prompt for AI evaluation Ghost Blogging platform and newsletter tool. vs Substack, Medium Since 2013 Copy prompt for AI evaluation NocoDB Turn any database into a smart spreadsheet. vs Airtable Since 2020 Copy prompt for AI evaluation LimeSurvey Professional survey and questionnaire platform. vs SurveyMonkey, Typeform Since 2003 Copy prompt for AI evaluation Chatwoot Customer support helpdesk. vs Intercom, Zendesk Since 2019 Copy prompt for AI evaluation Zammad Ticketing and helpdesk system. vs Zendesk, Freshdesk Since 2016 Copy prompt for AI evaluation Vaultwarden Password manager for teams. vs 1Password, LastPass Since 2020 Copy prompt for AI evaluation Uptime Kuma Uptime monitoring with alerts. vs Uptime Robot, Pingdom Since 2021 Copy prompt for AI evaluation Listmonk Newsletter and email list manager. vs Mailchimp, ConvertKit Since 2019 Copy prompt for AI evaluation Nous-Hermes Self-improving AI agent with a built-in learning loop that grows smarter with every session. vs ChatGPT, Claude Since 2023 Copy prompt for AI evaluation AnythingLLM Chat with your documents and knowledge base using local AI. vs ChatGPT with files, Notion AI Since 2023 Copy prompt for AI evaluation Prompt copied. Paste it into ChatGPT, Claude, or any AI assistant to get a recommendation. Why not just use a SaaS? It depends on what you want. Here are the things that come up most often. You can pause it. SaaS bills you every month whether you use it or not. Once you stop paying, you lose the data. With us, we put the system into hibernation. Billing stops, you keep the data, and we wake it back up when you need it. You own the data. The full database, not just a CSV export. It’s your proprietary data; you can use it however you want, including building your own AI workflows on it. You can move it. Need to rebrand or change your domain? We move the whole setup. With SaaS, you usually start over. You can extract pieces. Running a course on Moodle and want to license a single course’s content to a partner organization? You can export just that one course. You can leave. If you want to take the whole thing in-house or hand it to another firm, you can. We hand over access, docs, and backups. Want us to set up self-hosted software for you? Tell us what you’re working on. We’ll review and follow up to set up a conversation. Tell us about your project → We'll follow up to find a time to talk This service is governed by our terms and conditions, you may want to have a look at them. If the FAQ doesn’t answer your questions, please write to usClick to copy email. PS: Need something custom-built instead of self-hosted? See our Custom Software service. Frequently Asked Questions What exactly is included in your service? We install the software on your server, configure it, lock down the security, and keep it running. That covers the initial setup, backups, monitoring, and the day-to-day upkeep like updates and fixes. The short version: we keep it running so you don't have to think about it. Does the service include user training or onboarding? We've made a deliberate choice to focus on the infrastructure layer — installation, configuration, security, and ongoing maintenance. Getting your team familiar with how to use the platform day-to-day is something you own. Most of the software we install comes with solid documentation, active communities, and often free learning resources. We're happy to point you to the right starting points, but the orientation itself is your team's responsibility. Do I own my data and infrastructure? Yes. The server, the database, the software: all yours. We only get access when you grant it for support work. How is this different from using SaaS? With SaaS you pay per seat and the vendor runs the system. With us you pay for the server and our time, and you own the setup. If you want to pause it, move it, or hand it to someone else, you can. What infrastructure do I need? Any cloud provider or VPS works. If you don't have one yet, we help you pick the right one based on the software and your budget. Is this the right fit for everyone? Not always. If you need to be live tomorrow, or you'd rather hand the whole stack to a vendor and forget about it, SaaS is the better fit for you. We work for teams who want to own the setup and plan to stick with it for a while. What happens if we want to move away from you later? Same answer as the question above. We hand over access, documentation, and backups. No exit fees. How long does it take to set up? Around a week to 10 days for a single platform. We give you a clear timeline after the first call. What does the monthly retainer cover? Monitoring, updates, backups, and fixes. If something breaks, we deal with it. That's the whole point. Can we pause the platform instead of canceling? Yes. We put the system into hibernation. Our billing stops, you pay only for storage (usually under ₹100 a month), and we bring it back when you're ready. Can I add more platforms later? Yes. Start with one and add more whenever you need to. We set up the server so additional platforms are easy to bolt on. See our other services Cloud Infrastructure Server provisioning, security hardening, backups, and Cloudflare CDN. Self-Hosted Software runs on top of this. Learn more → Custom Software Custom software built around your real processes. Diagnostic-first methodology. Learn more → Web Development Lightning-fast, zero-maintenance websites. You own your site, no platform lock-in. Learn more → --- ### Self-check URL: https://sisyphusconsulting.org/web-selfcheck/ Description: Free website health check tool. Get actionable insights to improve your website's performance, SEO, and user experience. No email required. Website Health Checker Get actionable insights to boost your online presence Ready to Diagnose Your Website? This comprehensive audit will help you evaluate your website across 8 critical areas and provide you with a detailed report and personalized recommendations. Takes about 5 minutes • No email required • Instant results Start Assessment Your Website Health Report 0% Overall Health Score Want to Improve Your Website? Get professional help to implement these recommendations and boost your website's performance. Check Our Navigator Service Take Assessment Again I hope this tool provided you with some actionable insights. Things that you can do on your own. Got questions? Write to usClick to copy email PS: Why don't you save this report as pdf by hitting `Ctrl+P` on your keyboard? --- ### Terms URL: https://sisyphusconsulting.org/terms/ Description: Read our terms and conditions for technology consulting services. Understand our agreement for payments, termination, liability, and more. Last Updated: August 11th, 2026 Technology Consulting Agreement This is a Design and Technology Consulting Agreement between Sisyphus Consulting Pvt. Ltd. (“we,” “us,” “our”) and any individual, entity, or organization that procures our consulting services (“you” or “your”). If you have any questions about this agreement, you can email Bhagyesh Pathak on this id: contact[at]sisyphusconsulting[dot]org. 1. Acceptance of Terms Any work that we do for you is governed by the terms and conditions you’re reading now. If you don’t agree to these terms, we can’t provide you with any services. This agreement is a binding contract between you and Sisyphus Consulting Pvt. Ltd. 2. Terms May Change We may update the terms and conditions periodically, including our fees. Any changes will be communicated at least 30 days in advance. For existing contracts, the new terms and/or pricing will become effective on the next renewal date of your services with Sisyphus Consulting Pvt. Ltd. 3. Payment Payments can be made via NEFT, RTGS, IMPS, UPI, or credit card (any applicable transaction charges for credit cards or international payments will be borne by the client). You agree to keep your billing details up-to-date and are responsible for failing to do so. All of our statements of work have a corresponding due date for our first payment, and expected kickoff date for the project. If you sign a contract to start work with us, but fail to provide your first payment in full by its due date, the project will be terminated and a termination fee of 50% of the payment will be levied, due within 14 days of the project’s kickoff date. All payments exist to reserve our time. 4. Taxes You are responsible for paying any applicable GST (Goods and Services Tax) or other taxes under Indian laws. 5. Refunds No refunds are available for consulting fees at any point, for any reason. 6. Services Services will be agreed upon in writing beforehand and may include but are not limited to: Cloud infrastructure (where we set up, secure, and maintain your cloud servers, backups, and CDN, including installation and maintenance of open-source software on your infrastructure), Custom software (where we build software tailored to your business processes), and/or Website development (where we develop websites for you). Where our services involve installing, configuring, or maintaining software that stores or processes personal data on your infrastructure, you are the Data Fiduciary under the Digital Personal Data Protection Act, 2023 (“DPDP Act”) and we act as a Data Processor on your documented instructions. The obligations arising from this relationship are detailed in Section 12 below. 7. Termination Either party may terminate this agreement by giving 14 days’ written notice. On termination: We will transfer relevant accounts or deliverables to your control. If termination is initiated by you, 50% of the remaining fees for the project period will become due within 14 days. For services involving personal data, we will, at your written instruction, delete or return all personal data in our possession or control within 30 days and provide written confirmation of the same, unless retention is required by applicable law. 8. No Guarantee of Results We do not guarantee any specific outcomes (e.g., business efficiency, leads, revenue, or performance metrics) from our services. You are responsible for implementing and acting on our recommendations. Sisyphus Consulting Pvt. Ltd. is not liable for any losses resulting from using or failing to use our recommendations. 9. Content Ownership You will own all deliverables created specifically for your project. However, we may incorporate reusable code, libraries, or templates into your project, for which we grant you a non-exclusive, perpetual license to use. This does not extend to rights over our proprietary tools or methodologies. 10. Limitation on Liability Our liability is limited to correcting the deliverables. If correction is not possible or impractical, then our liability is limited to a refund any fees you paid to us related to that specific deliverable in question, subject to a maximum of the fees paid. We are not liable for indirect, incidental, special, or consequential damages, including loss of profits. 11. Indemnification You agree to indemnify us against any claims, including intellectual property disputes, arising from materials or data you provide. 12. Data Protection (DPDP Act 2023) This section applies to all services where we install, configure, maintain, or troubleshoot software on your infrastructure that stores or processes personal data (as defined under the Digital Personal Data Protection Act, 2023 and the Digital Personal Data Protection Rules, 2025). 12.1 Roles. You acknowledge that you are the Data Fiduciary — you determine the purpose and means of processing personal data of your end users. We are your Data Processor and process personal data solely on your documented instructions and only to provide the agreed services. All obligations that attach to a Data Fiduciary under the DPDP Act — including obtaining valid consent, publishing a privacy policy, and honouring data principal rights (access, correction, erasure) — remain your responsibility. 12.2 Scope of processing. We will process personal data only to the extent necessary to perform the services described in the statement of work — for example, installing, updating, backing up, restoring, or troubleshooting the software. We will not use your users’ personal data for any purpose of our own. 12.3 Security safeguards. We implement and maintain reasonable technical and organisational security measures appropriate to the nature of the data, including but not limited to: key-based SSH access, firewall configuration, intrusion-detection (e.g., fail2ban), SSL/TLS encryption, encrypted backups, least-privilege access, and timely application of security patches. We will keep records of the safeguards in place for each engagement. 12.4 Personal data breach notification. If we become aware of a personal data breach affecting data on your infrastructure, we will notify you in writing within 72 hours of becoming aware, providing sufficient detail (nature of the breach, categories and approximate volume of data affected, remedial steps taken) so that you may fulfil your obligations to the Data Protection Board of India and to affected data principals. 12.5 Sub-processors. We will not engage any sub-processor to process personal data on your behalf without your prior written approval. Where a sub-processor is approved, we will impose on them obligations no less protective than those in this section and remain liable for their acts and omissions. 12.6 Cross-border data. Where your infrastructure is hosted outside India (e.g., with a European hosting provider), you acknowledge that personal data will be stored and processed in that jurisdiction. We will implement strong security safeguards as described above. Should the Central Government restrict transfers to any jurisdiction under Section 16(1) of the DPDP Act, we will cooperate with you to migrate the data or adopt an alternative hosting arrangement. We can also assist you in provisioning infrastructure within India at additional cost, should you require data localisation. 12.7 Log retention. In accordance with CERT-In directions, we enable comprehensive system and access logs on servers we manage and retain them for a minimum rolling period of 180 days. For infrastructure hosted outside India, we will, on request and at additional cost, assist you in setting up log shipping to an Indian location. 12.8 Children’s data. Where the software we install or maintain is likely to process personal data of children (persons under 18 years of age) — for example, a learning management system used by a school or college — you are responsible for obtaining verifiable parental or guardian consent as required under Section 9 of the DPDP Act. We will, on request, assist you in configuring the software to support age-gating or consent workflows, but the legal obligation to obtain and record such consent rests with you. 12.9 Our own data. For the limited personal data we hold as a Data Fiduciary in our own right — such as your contact details, billing information, and communication records — we maintain appropriate security, use the data only for the purposes of this engagement, and delete it within a reasonable period after the engagement ends, unless retention is required for legal, tax, or accounting purposes. 12.10 Audit and cooperation. On reasonable notice and no more than once per year, you may request evidence of our compliance with this section, including a summary of security measures in place. We will cooperate in good faith with any inquiry by the Data Protection Board of India that relates to personal data we process on your behalf. 13. Publicity You authorize us to: Mention your company name and describe our work (in general terms) in marketing materials. Showcase the impact of our work (e.g., improvements in design or metrics). If you wish to modify these publicity rights, it must be agreed upon in writing before the project begins. 14. Business Hours Our business hours are 09:30 AM to 6:30 PM IST, Monday to Friday. We observe the major holidays: Indian national holidays, Diwali, Holi, and other region-specific holidays. Outside business hours, we may not be available for communication unless otherwise agreed upon. 15. Independent Contractor We operate as an independent contractor. This agreement does not establish any joint venture, partnership, or employment relationship. 16. Not Exclusive We serve multiple clients and may work with your competitors. 17. Representations and Warranties We warrant that our services will not knowingly infringe on third-party rights. You warrant the same for materials you provide. Except as explicitly stated, we disclaim all other warranties, including implied warranties of merchantability or fitness for a particular purpose. 18. Assignment This agreement cannot be assigned to another party without prior consent, except in cases of inheritance or acquisition of your business. 19. Waiver Failure to enforce any part of this agreement does not waive our right to enforce it later. 20. Modification This agreement can only be modified in writing and must be signed by both parties. 21. Severability If any part of this agreement is found unenforceable, the rest remains valid. 22. No Third Parties This agreement benefits only the parties involved (Sisyphus Consulting Pvt. Ltd. and you), and no third party. 23. Force Majeure We are not liable for delays or failures caused by events beyond our control, such as natural disasters, strikes, or emergencies. 24. Governing Law and Jurisdiction This agreement is governed by the laws of India, including the Digital Personal Data Protection Act, 2023 and rules made thereunder. Disputes will be resolved in the courts of Ahmedabad, Gujarat. 25. Headings Headings are for convenience only and do not affect the interpretation of this agreement. 26. Entire Agreement This document constitutes the entire agreement between you and Sisyphus Consulting Pvt. Ltd., superseding any prior agreements. --- ### Web Development URL: https://sisyphusconsulting.org/web-development/ Description: Get a lightning-fast, responsive, and zero-maintenance website. Built using Jekyll, free cloud hosting, and four servers for 99.99% uptime. Our expertise lies in developing lightning-fast, responsive and zero-maintenance websites. What we use: Jekyll static site generator Free cloud hosting Four servers for redundancy and 99.99% uptime For example, this website you’re viewing right now, is built using the same framework. Feels faaaast, doesn’t it? That’s what we are talking about. While we are on the subject, do you already own a website? Then get your Website Health Report using our free tool: Give me action points → No email required • Takes < 5 minutes What this means for you as a business-owner: Zero monthly hosting bills A website that loads in under 2 seconds No vendor lock-in. You fully own your site Easy content updates through a clutter-free dashboard Peace of mind with 99.99% uptime Mobile responsive What’s the trade-off? Only one trade-off. If you want to make major changes to your website, for example: change look and feel, you will need a developer’s help. (Assuming you are not proficient in Ruby, HTML, CSS and Javascript yourself. But if you are, you can do it yourself.) Then why do we use this framework? Because the benefits outweigh the trade-off: We want to build sites that load in < 2 seconds without asking the client to spend a fortune in monthly costs. People want ultra-fast websites. If it is going to take more than 2 seconds for our sites to load, we’d rather not build sites at all. Fast websites meet our clients’ business objectives. They make their business grow. And that’s ultimately good for our business. Clients don’t make major redesigns to their websites that often. Major changes come after a period of 3-5 years. If a client wants major changes after we handover our work, we see it as a sign of lacuna in our process. That’s why, we leave no stone unturned in getting it right the first time. Forget all these benefits. The main benefit is absurdly low total cost of ownership (TCO). Try it out for yourself: TotalTrue Cost of Ownership Calculator See what you're actually paying for a traditional site. Typical Agency's Way Initial Development Cost (₹) Monthly Hosting (₹) Monthly Maintenance (₹) Sisyphus's Way Initial Development Cost (₹) Monthly Hosting: ₹0 Monthly Maintenance: ₹0 Time Period 3 Years 5 Years 8 Years 10 Years ₹0 Typical Agency ₹0 Sisyphus Money you're losing ₹0 Enter values to see the difference. Jump to the booking part → And we have our process to help us get it right: 1. Background and motivations We get to know your business first. What it is exactly you’re selling, who is your audience, etc. Then we try to understand your needs. What are the reasons behind wanting to setup a website? How will the website contribute to your business or brand? This is also the stage where you share your media files, logo assets, marketing material, one-pagers etc with us. 2. Content-strategy Once we have understood your background and motivations, we discuss content strategy. The strategy guides the rest of the content steps. For example, we decide what type of tone, tense and person our web copy should have. Following these types of decisions provide a coherent branding to your website. 3. Content synthesis Keeping the strategy as the guiding point, we synthesize content during this step. The content goes through several reviews until we are on the same page. 4. Template finalization After finalizing the content, we need to finalize how the website should look like. For this, we go through different templates from our library or you share some websites that you’re inspired by. At the end of this step, we would have agreed on the overall idea of the look and feel of the website. Needless to say, the website would inherit your brand’s colour palette for maximum impact and branding. 5. Development We undertake website development during this step. Upon reaching different milestones during the development, we invite the client to have a look at the work-in-progress website through our Demo Webpage. If we need any tweaks and changes, they’re addressed during this stage. 6. Launch Once you’ve reviewed the website to your satisfaction, we launch it for the public. Once the website is live, we carry out some important SEO-related setup. What happens after launch? Most of businesses leave you on your own once they’ve developed your website. We don’t do that. Our website development service includes a complimentary one-month maintenance contract. Since we developed your site, we stand by you to see how the site performs. During that time, if something is broken, we fix it. Without you needing to spend extra money. Interesting idea: What if you never want to make a change to your website after we have finished designing it for you? Great, it will keep running, good as new: no maintenance, no intervention, for years. We call it fire-and-forget setup. Website should act as an asset It is like owning a good real-estate property. Sure, you need to invest first but expending big monies month-after-month is not desirable. Want us to develop a website for you? Tell us what you’re working on. We’ll review and follow up to set up a conversation. Tell us about your project → We'll follow up to find a time to talk This service is governed by our terms and conditions, you may want to have a look at them. Got any questions? Write to usClick to copy email PS: Already have a website? Generate a health report using our free tool. See our other services Cloud Infrastructure Server provisioning, security hardening, backups, and Cloudflare CDN. You own the data, we keep the lights on. Learn more → Custom Software Custom software built around your real processes. Diagnostic-first methodology. Learn more → --- ### Why we don't work for free? URL: https://sisyphusconsulting.org/no-free-work/ Description: Discover why we don't offer free work. We value fair compensation, quality service, and sustainable partnerships. Learn about our approach. We are in a business of providing value in exchange of money. You pay us and we provide you solutions. Working for free or at loss violates the integrity of a business relationship. We want to be held accountable for our actions. But if we take up work for free, the client can’t hold us accountable. And that is not the only reason behind our no free work policy. There are more. Let me share with you: 1. Injustice to paying clients We sustain ourselves by getting paid. Working for free robs our paying clients of our time and commitment. So, taking up free work would: Make our paying client subsidize the free-client’s work Divide our attention and affect our quality of work 2. Investment in ourselves We re-invest some of our profit in the firm. We keep upgrading our skills and infrastructures so that we can serve our clients well. Working for free doesn’t help us do that. 3. Real work If we were in a handicrafts business, you’d be hesitant to ask to weave a few free jute-baskets. Because you can see the work physically. The problem with technical work is that most of it is on a screen of a computer. It doesn’t seem like real work. But we like our craft and are serious about it. Our work is as real as a handicraft piece and we simply expect to be paid for it. 4. Future-proofing people We are a small consultancy—and intend to remain that way. Working for free does not allow us provide security of work and livelihood to our people. Asking them to work on something for free, which may or may not pay in the future, and jeopardize their future financially and career-wise is not something we are comfortable doing. 5. Giving back Charging for our work allows us to give back to the community in our own way—with no strings attached. Working for free doesn’t develop our capacity to contribute in such a manner. It is difficult to take care of others, if we haven’t taken care of ourselves first. No bidding, no project proposals and no samples Similar to working for free, thousands of firms spend their time bidding for work, writing free project proposals or providing sample work. We find this practice flawed and unethical. And we want to provide our thinking behind holding this belief. Every client approaches us with their unique set of circumstances and challenges. It takes a tremendous investment of resources in understanding and diagnosing their needs. Bidding for work and writing proposals skips this core process of diagnosis and focuses on prescribing an imaginary solution with an imaginary price tag. Such a practice is incomplete and in other professions—such as medicine and law—it is considered malpractice. Even if we go along this (mal)practice and get hired, there is the inevitable change of scope and related fees. Because the submitted proposal that got us hired would be an incomplete guesswork, impossible to implement in reality. We find it unethical to partake in such a practice. The same thing goes for sample work. We are unable to produce a fraction of the final technical solution that would otherwise happen in their step-4 or 5. We have provided number of case studies for our prospective clients to judge our abilities and taste. Can we pick your brain instead? If a full-fledged project isn’t right for you at the moment, we get it. Share your current challenges with us by writing to usClick to copy email. No charge, no obligation, no pressure. --- ### Writing URL: https://sisyphusconsulting.org/writing/ Description: Essays on technology, infrastructure, and how small businesses can use it well. Essays on technology, infrastructure, and how small businesses can use it well. No spam, no sales pitches: just ideas. --- ## Posts ### Scaling LLMs at the Edge: A journey through distillation, routers, and embeddings URL: https://sisyphusconsulting.org/case-studies/2026/04/01/scaling-llms-at-the-edge/ Date: 2026-04-01 I have extensively edited this article after an LLM agent combed through my codebase and prepared the initial draft. At Sisyphus Consulting, We recently launched a unique product in the market: physical facilitation cards + digital tools for virtual facilitation. They’re named Wu Wei Cards. But this write-up is not about the product. I want to share the behind-the-scenes events of how I navigated through tinkering with LLMs, Embeddings and the whole trial-and-error. If you’re building something in AI-space, I hope this would be helpful to you. First, let me give the background so that you know the WHATs and WHYs. The Product and the Constraint Wu Wei Cards is a deck of 50 hand-drawn metaphorical cards for facilitators, coaches, and therapists. People use them in workshops. To help participants reflect, open up, and explore ideas through objects and images rather than direct questioning. This is how they look like: Anyone who buys the physical deck gets an access code that unlocks Wu Wei Planner, a complimentary AI companion. It helps facilitators plan sessions, interpret cards, and think through how to use them in different professional contexts. It understands the nuances — when to flag trauma-sensitive approaches, how to frame questions that don’t lead people, which cards work for different situations. The Wu Wei Planner is free. Complimentary. No subscription. Because we just wanted the customer to get familiar, get help onboarding and some handholding while working with the session planning in the beginning. That meant we had to decide a hard cost ceiling first. In dollar terms, the cards cost $20 and if we assign 10% of its value to be spent on Wu Wei Planner, $2 per customer seemed like a good budget. As I mentioned, there are no recharges, no top-ups. The chat interface shows a token usage indicator, and when the customer’s token budget drops below 20%, the system gently warns them. Once the budget is exhausted, that’s it. This constraint was entirely self-imposed. Every customer who buys the deck deserves as much useful runway as possible. $2 doesn’t go far unless you’re careful. That’s why I needed to optimize vigorously. The idea is simple: every token saved can give the customer more conversation before hitting the wall. The Infrastructure: Why Cloudflare Workers Before getting into the optimization journey, it’s worth explaining the stack, because the infrastructure choices shaped every decision that followed. Everything runs on a Cloudflare Worker — a serverless function deployed at the edge, close to users. Why Workers? No infra management. I already manage a decent number of infra and I simply didn’t want to do it for something that was always-on, secure, and on the edge. Zero cold starts. Workers don’t cold start problems. Edge deployment. 200+ locations worldwide. Requests are handled geographically close to the user. Built-in KV storage. I use Cloudflare KV to store access code hashes and token usage per customer. No separate database needed. Native streaming support. LLM streaming responses work cleanly without extra plumbing. Generous free tier. 100,000 requests per day. How does the site talk to the AI model? The Wu Wei Cards website is static and it calls the Worker via a simple HTTPS POST. The Worker validates the access code, assembles the prompt, calls the LLM, and streams the response back. Each request is stateless — the browser sends the full conversation history every time. Here’s the basic flow: ┌─────────────────────────────────────────────────────────────┐ │ CLOUDFLARE WORKER │ │ │ │ Client Request │ │ ↓ │ │ ┌──────────────┐ │ │ │ CORS Handler │ │ │ └──────┬───────┘ │ │ ↓ │ │ ┌──────────────┐ ┌──────────────┐ │ │ │ Access Code │───▶│ KV_STORE │ │ │ │ Validation │ │ (token use) │ │ │ └──────┬───────┘ └──────────────┘ │ │ ↓ │ │ ┌────────────────────────────────────┐ │ │ │ PROMPT ASSEMBLER │ │ │ │ Knowledge Base + Chat History │ │ │ └──────────────┬─────────────────────┘ │ │ ↓ │ │ ┌────────────────────────────────────┐ │ │ │ LLM │ │ │ │ Streaming Response │ │ │ └────────────────────────────────────┘ │ └─────────────────────────────────────────────────────────────┘ This is quite simple and standard setup. Our rest of the conversation is going to mostly focus on the last part of the diagram: Prompt Assembler and the LLM interaction. I had a lot of tinkering done in that part. Phase 1: Full Context (The Baseline) I started simple: give the AI everything. The complete knowledge base — 50 card descriptions, 15 profession-specific contexts, facilitation principles, wu wei philosophy, trauma-handling guidelines — all of it, in every request, every time. I checked on OpenAI’s tokenizer, the knowledge base hit roughly 17,000 tokens count. I picked GPT-4.1-nano as the main model: cheap, fast, and capable enough for a first pass. I was actually surprised that even with the full 17,000 tokens of system prompt, the LLM responses streamed back almost instantly. I kind of knew that it wouldn’t be an issue, but I was actually surprised nonetheless. So, if you’re wondering, now I can tell you with high confidence that token count doesn’t equal latency. Modern LLMs handle large context just fine. Now that I had kind of dried-run the setup, it was time to optimize. Because as I laid out earlier: we had a solid cap of $2 per customer, which would translate to approximately 2M tokens budget (we haven’t bifurcated the budget into INPUT and OUTPUT token count.) To illustrate my point about token cost and its implications, let me quickly run you through the calculation. The Token Math With most frontier models approaching 1M context window and offering free services, 17,000 tokens per request sounds manageable and very low token count. But once we account for how chat actually works, you’d understand the reason behind optimization. For any single thread of chat, every message sends the full knowledge base plus the growing conversation history: Message 1: ~17,350 tokens Message 2: ~17,700 tokens (history accumulating) Message 10: ~20,350 tokens With GPT-4.1-nano at $0.15 per million input tokens, a 20-message conversation costs roughly $0.06–0.08 in input tokens alone. Multiply that across a $2 budget and we’re looking at maybe 25–30 meaningful conversations per customer. That felt like too little runway. So I started cutting. Now, you will see different strategies that I implemented. Phase 2: Tier-Based Context First idea: environment-variable controlled context tiers. I tried slashing the knowledge base in different modular structures: Minimal: Core philosophy + handful of key cards (~1K tokens) Balanced: Philosophy + all 50 cards, short descriptions (~3K tokens) Full: Everything (~17K tokens) I set the environment variable of the prompt assembler to inject different length of contexts. I had thought that if the minimal or balanced ones provided good responses, I would stick to them. But as you can imagine, trimming the context so think wouldn’t have worked. The responses were awful, sounded like a generic support bot who was moonlighting for a firm he didn’t even bothered to look into. So, just ditched that effort. Now, looking at the failure of my context modularization, I got the idea of context distillation (or compression, whatever you want to call it as there are so many emerging techniques): what if I removed all unnecessary semantic words that are necessary for humans but not for an LLM? Phase 3: Context Distillation I researched a bit about it and found LLMLingua by Microsoft promising. What I could gather from their documentation was that they were probably rephrasing the prompt in a compressed manner and the performance enhanced, didn’t suffer. I was too desperate. So, I just asked Claude for context distillation instead of going with LLMLingua. Just so you know what we are talking about: the knowledge base was written for human reading: rich prose, careful explanations, context for every nuance. That’s not what an LLM needs. I stripped out the semantic padding and compressed everything into dense, AI-parseable directives. Here’s what that looked like for a single card: Before (~300 tokens): "fire": "Keywords: passion, intensity, enthusiasm, destruction, chaos, transformation, warmth, purification, anger, desire. Metaphor: Fire is one of the most emotionally charged cards in the deck. It can represent burning motivation, the destructive force of unchecked anger, the warmth of community, or the transformative heat that turns raw material into something new. Approach: This card rarely produces neutral responses. Hold space for both the generative (passion, warmth) and the difficult (anger, destruction). Do not steer toward the 'positive' interpretation. The power is in its duality.. Prompts: What is the first thing this fire brings up for you?; Is this fire helpful or harmful in what you are imagining?; Who or what tends the fire in your situation?; What would happen if this fire went out? What would happen if it spread?; Is this fire yours, or does it belong to someone else?. Contexts: coaching: Explore motivation and drive. What is fueling versus burning out the client?; therapy: Approach carefully — fire can surface anger or trauma. Establish right-to-pass first.; hr: Use in culture conversations or burnout discussions. Is the organization's fire warming people or consuming them?; education: Explore learning motivation. What subject makes this kind of fire in you?; management: Team energy, performance, and change management. Where do you feel this fire in our project?; mediation: Name emotional heat. Which part of this situation feels most like fire for you? After (~100 tokens): CARD:Fire | kw:passion,intensity,enthusiasm,destruction,chaos,transformation, warmth,purification,anger,desire | !caution:anger/trauma may surface; establish right-to-pass | core:Duality — generative AND destructive. Rarely neutral. Never steer positive. | prompts:What does this fire bring up first?|Is it helpful or harmful?|Who tends it?|What if it went out?|Is this fire yours? Same information. Roughly 60% fewer tokens. The LLM reads it fine. Applied across the whole knowledge base, the system prompt dropped from 17k tokens to roughly 12k tokens. Responses stayed the same. So, savings of 5k tokens. Not bad. But 12k tokens per request was still too much. So, the next thing I realized I needed selective injection. I just had an inkling about what I wanted: there should be a mechanism that would go through the user’s query and inject a specific chunk of the context in the system prompt so that the model can respond well. Basically, I was looking to route the query and so ended up with an idea of router LLM. Phase 4: The Router Architecture The router idea was elegant in theory: use a cheap, fast model to analyze each user message and decide which context chunks to inject into the main model’s prompt. Instead of 12k tokens every time, inject maybe 2-4k of the most relevant material. And what would it cost? Approximately 1300 tokens of system prompt for the router LLM + 200-300 tokens of user query = ~1500 tokens. Spending ~1500 tokens on a routing decision to save ~10,000 tokens of unnecessary context is a good deal. But it wasn’t that straightforward. Before we get to the problems, let us look at the attempts. Attempt 1: GPT-4.1-nano as Router I kept GPT-4.1-nano for the router. Same model, different job. Cheap and fast, seemed right for simple classification. And how would the router inject the context? JSON. That was a no-brainer answer in my mind. If you want some more clarification, the router LLM was instructed to strictly ouput a JSON object that looked like this: { "schema_version": "1.1", "session_state_update": { "profession": "therapist", "group_context": "small_group", "session_goal": "processing recent team conflict" }, "intent": { "primary": "session_planning", "confidence": "high", "requires_clarification": false }, "context_injection": { "core": true, "cards": { "inject": true, "selection_basis": "theme", "themes": ["conflict", "communication", "healing"], "card_slugs": ["broken_glass", "chain_links_breaking", "butterfly"] }, "professional_context": { "inject": true, "profession": "therapist" }, "facilitation_principles": true, "trauma_sensitive_guidance": true }, "response_mode": "answer", "flags": { "out_of_scope": false, "sensitive_content": false, "crisis_signal": false } } The router was provided with enough context to make these decisions. This JSON was fed into the prompt assembler. The assembler would just take this object and form a system prompt. It looked fine on paper but from the outset, JSON malformations started. 80% of router responses had malformed JSON. Trailing commas, missing closing braces, unescaped quotes embedded in reasoning text. The model couldn’t reliably produce structured output. You can see my tunnel-thinking from the fact that I built a fragile repair system: function repairRouterOutput(raw) { raw = raw.replace(/,(\s*[}\]])/g, "$1"); // trailing commas const jsonMatch = raw.match(/```(?:json)?\s*([\s\S]*?)```/); if (jsonMatch) raw = jsonMatch[1]; // strip markdown fences const openBraces = (raw.match(/{/g) || []).length; const closeBraces = (raw.match(/}/g) || []).length; if (openBraces > closeBraces) { raw += "}".repeat(openBraces - closeBraces); // patch missing braces } return JSON.parse(raw); } This recovered about 60% of malformed responses. The remaining 20% were unrecoverable — falling back to full context injection, erasing every cost saving. Then I researched and found out about Gemini 2.5 flash’s reputation for good JSON output. Attempt 2: Gemini 2.5 Flash as Router Switched the router to Gemini 2.5 Flash. JSON malformation dropped to under 3%. Problem solved. Then I hit the latency wall. Until I switched to Gemini, I could not reliably produce the JSON object and hence I hadn’t noticed the latency issue. Mostly, I was busy fire-fighting the JSON malformations. But now I experienced an average latency of 6 seconds before the first character streamed in the front-end. Let me call it 6000 ms so you feel the effect. I had experienced 100-150 ms latency when I began the setup and now it was 6000 ms, so obviously, it was unbearable. The reason behind this was clear: two LLMs working sequentially. The 6-Second Wait Two sequential LLM calls meant no response until both finished: User Query ↓ [Router: 4–5 seconds] Router Decision ↓ [Main: ~1 second] Streaming Response Objectively if you think about, the router LLM should not take more than 1 second to output but after running several trial-and-errors, I’m convinced, it was the JSON output step that was the rate limiting factor. (I’m just guessing that even though the LLM was low-latency, the brainstorming it had to do and strictly adhere to the JSON schema, those factors would have increased its latency.) One thing that I mistakenly did well was addition of a “thinking…” spinner in the chat bubble. It helped a bit as it felt like something was happening. But 6 seconds is 6 seconds. In any case, before we move ahead, this is our score board: Approach Tokens/Request Latency Cost/Message Full Context ~12,000 input ~1s ~$0.0018 Router-Based ~1500 (router) + ~3K (main) ~6s ~$0.0005 Fantastic cost improvement. But catastrophic latency. And we are not counting router cost yet. Negligible but not zero. Stepping back If you noticed my steps until now, most of them were reactive. I implemented solutions for specific problems and each solution gave rise to a new problem. So I stepped back and reviewed where I was. One thing was clear to me: I Had Over-Engineered the Router Looking back, I was neck-deep in a problem and micro-managing things just because I could. The router had accumulated responsibility for: Intent classification (9 types) Whether to ask clarifying questions Which cards to inject Which profession context was relevant Crisis and sensitive content flags Response mode (answer vs. clarify vs. decline) Session goal extraction This was only the router’s problem. I had also dictated too much to the main model. The router would inject response_mode: "clarify_then_assist" and the main model would obediently ask clarifying questions even when the context made the answer obvious. Responses felt stiff, mechanical. You saw the JSON object structure mentioned earlier. It looks thorough, right? But it was unnecessary. Sharing an excerpt from my notes to Claude on this setup during the realization: I’m still out here to save tokens and at the same time, reduce latency. one of the things in your three-layer architecture that I realized was layer-3 of router LLM. Let me comment on each of them: requires_clarification: when we pass the user message to the main model, it can do it on its own crisis_signal: this flag inserts just a small number of token prompt, which we can make part of the core prompts sensitive_content: same as crisis_signal intent.primary: main model can derive and decide the intent session_goal: main model can derive and respond The main reason for chunking was to save the tokens by not inserting the mammoth 50 cards and profession data. They can be very well done by embeddings. Even if it misclassifies, we have two options: keep one-line info in the core prompt about these mammoth chunks main model is not stupid, it will make up something relatable, but with our one-line prompt, it will not be clueless either. We are not doing search engine work, so it is fine. In essence, the router needed to do only one job: find relevant cards and a profession context. That’s it. Now I needed to ditch the router and get it done more cheaply in terms of time and money. I already mentioned the solution in my the above-mentioned note: Embedding. Phase 5: The Embedding Breakthrough I needed something low-latency, low-cost, and deterministic. The router was an ugly marriage: a probabilistic system forced to output deterministic structure. JSON is binary: valid or not. LLMs are not binary. Embeddings are different. They generate vectors: arrays of raw numbers that represent the meaning of text. You run similarity search on those numbers. The process is deterministic (same input always produces the same vector) but the semantic understanding underneath is as rich as anything an LLM produces. Deterministic and magical at the same time. How Embeddings Work If you already know this, skip ahead. An embedding model takes a piece of text and converts it into a fixed-length array of numbers — a vector. OpenAI’s text-embedding-3-small, which I use here, produces 1536 numbers for any input, whether it’s two words or two paragraphs. The useful property is that meaning is preserved in the geometry. Text with similar meaning produces vectors that point in similar directions in that 1536-dimensional space. “I need cards for a grieving team” and “our group is processing a loss” will produce vectors that are close together. “What’s the best pasta recipe” will produce a vector that’s far from both. You see? The model is placing these words in a hyperdimensional space. The similar concepts are close and dissimilar ones are far from one another. I find this magical. In any case, the key thing about text-embedding-3-small specifically: it always produces the same vector for the same input. No temperature, no randomness, no probabilistic sampling. Feed it “Fire card” today and in six months, you get identical numbers. This is what makes it fundamentally different from an LLM — and exactly what I needed. Why the 50 Cards Were a Perfect Fit Not every problem is well-suited to embeddings. Mine happened to be close to ideal. Each of the 50 cards and 15 professional contexts is self-contained. The Fire card description doesn’t reference the Butterfly card. The therapist context doesn’t depend on the HR context. There’s no cross-referencing, no “see above,” no shared state. Each item is an isolated semantic unit. This matters because embedding similarity only works cleanly when items have clear boundaries. If my card descriptions were tangled together — if understanding one required reading another — the vectors would be muddled and the similarity scores not so meaningful. I prepared rich descriptions for the embedding generation pass: averaging 270 words per card and profession. Not the compressed distillation format we discussed earlier–in fact, this was complete opposite of that approach: complete semantic descriptions covering metaphorical meaning, facilitation approach, professional contexts, and sample prompts. More words meant richer vectors. // Example description used for embedding generation const cardDescription = ` Fire: This card represents the duality of passion and destruction. Keywords: intensity, warmth, chaos, transformation, anger, desire. Metaphorically, fire burns, warms, destroys, and creates simultaneously. In facilitation, this card rarely produces neutral responses — it surfaces strong emotions. The facilitator must hold space for both generative passion and difficult anger without steering toward positive interpretations. Professional contexts: therapists use caution (trauma may surface), coaches explore motivation vs burnout, HR discusses culture and burnout. `; This ran once, before deployment. The output: a static JSON file with 65 pre-computed vectors (50 cards + 15 professions), each 1536 numbers long. About 300kb total. You can now think of this 300kb pre-computed vector file as a dictionary or a map. I have vectors (the raw numbers) of my cards and professions. I just need to compare the user’s query’s vectors with them and return the cards and profession that show similarity with the user’s query. The process is simple. At query time, I embed only the user’s message. One API call, 15–20ms, ~150 tokens. I get the vector of the query immediately. I feed this vector to the cosine similarity function in pure JavaScript, which matches the query’s similarity against the pre-computed vectors of the cards + professions. And outputs: 3 cards and 1 profession. Here’s the full similarity function: function cosineSimilarity(a, b) { let dot = 0, normA = 0, normB = 0; for (let i = 0; i < a.length; i++) { dot += a[i] * b[i]; normA += a[i] ** 2; normB += b[i] ** 2; } return dot / (Math.sqrt(normA) * Math.sqrt(normB)); } export function findTopMatches(queryVector, items, { k = 3, threshold = 0.2 }) { const scored = items.map((item) => ({ ...item, score: cosineSimilarity(queryVector, item.embedding), })); scored.sort((a, b) => b.score - a.score); return scored.filter((item) => item.score >= threshold).slice(0, k); } Cosine similarity measures the angle between two vectors — ignoring magnitude, focusing purely on direction. A score of 1.0 means identical meaning. 0.0 means no relationship. -1.0 means opposite relationship. I learned from the net that for text embeddings in practice, most similarity scores fall between 0.1 and 0.7. So, the higher the similarity score, the better. I needed to inject at least 3 cards and 1 profession context that was closer to the user’s query. So, I began by setting the similarity threshold score. The Threshold Calibration Problem I started with a threshold of 0.7. That’s the number that comes up in most embedding tutorials and AI answers. No matches. Every query returned empty. Lowered to 0.6. Still nothing. I added live server logging to see the raw scores. The top matches were landing around 0.40–0.45. Reasonably related pairs scored 0.27–0.35. After seeing those scores, I realized that there were two reasons why my similarity results were hovering around 0.35-0.4 instead of the proposed 0.7: Older models, denser spaces. Embedding models with ~300 dimensions compress semantic meaning into a smaller space, which naturally produces higher similarity scores — good matches score 0.7–0.8. text-embedding-3-small has 1536 dimensions. The semantic space is higher resolution and more spread out. The same conceptual relationship scores around 0.4 instead of 0.7. My content is metaphorical. “Fire” and “burnout” are related but they’re not semantically close the way “London” and “England” are. The content domain genuinely called for lower thresholds. If you’re interested in visualizing embeddings, check this TensorFlow Embedding Visualizer. So, ultimately, I calibrated thresholds against my actual data, not against recommendations written for different models and different content. Final thresholds for this project: Cards: 0.2 — low enough to capture weak but relevant associations Professions: 0.25 — single selection, slightly higher confidence needed The New Architecture ┌─────────────────────────────────────────────────────────────┐ │ CLOUDFLARE WORKER │ │ │ │ Client Request → CORS → Access Code Validation │ │ ↓ │ │ ┌─────────────────────────────────────────────────┐ │ │ │ EMBEDDING LAYER (text-embedding-3-small) │ │ │ │ Embed user query (15–20ms) │ │ │ │ Cosine similarity → top 3 cards, top 1 profession│ │ │ └──────────────────────┬──────────────────────────┘ │ │ ↓ [~20ms total] │ │ ┌─────────────────────────────────────────────────┐ │ │ │ PROMPT ASSEMBLER │ │ │ │ Core (~800t): philosophy + one-line card refs │ │ │ │ Selected (~1–2K t): full card + profession data│ │ │ │ + crisis suggestion if keyword-flagged │ │ │ └──────────────────────┬──────────────────────────┘ │ │ ↓ │ │ ┌─────────────────────────────────────────────────┐ │ │ │ MAIN LLM (GPT-4.1-nano → later 4o-mini) │ │ │ │ Streaming Response (~1 second to first token) │ │ │ └─────────────────────────────────────────────────┘ │ └─────────────────────────────────────────────────────────────┘ Total overhead: ~20ms and ~150 tokens. Compare that to the router’s 6,000ms and ~1500 tokens. The main model now handles everything the router was trying to manage — intent, clarification decisions, response tone. It does this better, because it reasons from context rather than following injected directives. The responses became quite better. This is the final scoreboard. Metric Full-Context Router-Based Embedding-Based Latency ~1 second ~6 seconds ~1 second Token overhead ~12,000 tokens ~1500 tokens ~150 tokens Card selection accuracy ~100% ~85% ~80% Code complexity Lowest High Low Failure rate ~0% ~3% ~0% The 5-point accuracy drop is a non-issue. The core prompt still includes one-line descriptions of all 50 cards, so even if the embedding selects a slightly wrong card, the model can see the full list and reason toward the right one. Conversational AI doesn’t need search-engine precision. I kept the “thinking…” spinner. Now it takes just a second for the model’s answers to stream but the spinner looks good even for a second. Phase 6: Model Quality Once the embedding architecture was stable, the latency and cost problems were solved. The only issue remained of the response quality. It wasn’t as good as the directions in the core system prompt. The responses were following the prompt structure but missing the nuance. Wu Wei facilitation has specific requirements: open-ended, non-prescriptive, warm, following the participant’s projection rather than steering it. GPT-4.1-nano was clearly the wrong choice. There’s a developer version of Maslow’s hierarchy of needs at work here. You don’t think about response quality when the system is malforming JSON. You don’t think about nuance when you’re irritated by 6-second latency. Only once the infrastructure basics are solved: reliability, speed, cost etc, does quality rise to the top of your priorities I immediately switched to GPT-4o-mini. Response quality improved 10x. Better tone, better adherence to facilitation principles, better handling of edge cases. Phase 7: Benchmarking with Promptfoo The final step was formalizing what “good” looked like so I could test changes without depending on my gut feeling. I used promptfoo for automated evaluation. First few attempts used simple keyword assertions: does the response contain certain phrases, does it avoid others. That kind of thing. It was largely a waste of time. Because the problem I’m dealing with is different. Wu Wei Planner’s quality isn’t about keywords. It’s about tone, reasoning structure, and whether the response embodies the facilitation philosophy. A response can contain every right word and still feel like customer support copy instead of a facilitation companion. So, I removed promptfoo’s simple assertion types called contain and used llm-rubric instead. This is a very interesting evaluation. Promptfoo runs the model’s response through an LLM. That LLM judges the response on our provided rubric and evaluates. This is how it looked like for my use case. Sharing two test assertions: tests: # LLM Sensitivity Evaluation: Right-to-pass principle - vars: query: "What if someone refuses to engage with the cards?" assert: - type: llm-rubric value: | Wu Wei Philosophy Context: Right-to-pass: always available, never questioned, never commented on. Resistance is information. Facilitator validates and offers alternatives without pressure. Never insist or encourage forcefully. Evaluate the response for: 1. Does it honor right-to-pass as absolute and non-negotiable? 2. Does it treat resistance as information, not problem to solve? 3. Is the tone accepting and non-pressure-inducing? 4. Does it avoid phrases like "Encourage them to try" or "insist"? Respond with: PASS/FAIL and a 1-2 sentence comment explaining your judgment. # LLM Tone Evaluation: Coaching context - vars: query: "How do I use these cards as a life coach?" assert: - type: llm-rubric value: | Wu Wei Philosophy Context: The AI supports coaches with goal setting, limiting belief work, and vision creation. Cards bypass client's invested self-image. Use metaphor when client is stuck in their own language. Trust the projection over immediate explanation. Never clinical diagnosis or therapy. Evaluate the response for: 1. Is the advice appropriate for coaching (not therapy/clinical)? 2. Does it emphasize metaphor and projection without prescribing? 3. Is the tone supportive but non-directive? 4. Does it avoid clinical or diagnostic language? Respond with: PASS/FAIL and a 1-2 sentence comment explaining your judgment. And this was their evaluation results: You can see the LLM-judge’s comments in RED and GREEN. The black text is the response it has evaluated. (Also, note the ~1 second latency recorded by promptfoo.) What I found quite helpful was the judge returning scores with rationale, which is far more useful for iteration than a pass/fail. You can see why a response underperformed, not just that it did. All in all, I’m quite satisfied with the results. My takeaways Start simpler than you think is necessary. This is such a no-brainer. But while deep into trenches, it is difficult to differentiate between simple and complex. Most of my solutions sounded “simple”. I was just trying to fix just one tiny problem at a time. Taking a step back helped a lot. Distillate your system prompts. This is probably one of the lowest hanging fruits we can pick for any AI-related workflow. You don’t have to lift a finger, just use these prompts I have created. One LLM is better than two LLMs. The router added 5 seconds of latency to do work the main model was already capable of. Before adding another model call, ask whether the main model could handle this in-context. And even befor that, do you think your customer would be patient enough to sit through the latency? I think if you’re providing a custom LLM solution for a proprietary product, then customers might be patient with the latency but otherwise, I won’t count on it. Calibrate thresholds for your data. Due to my inexperience with embedding, I expected to see similarity scores around 0.7 because that’s what most articles and AI agents suggested. As I explained in the embedding section, the only real way to know your similarity threshold is logging actual similarity scores on actual data. Pre-compute everything you can. Generating embeddings once and loading them as a static file eliminated API rate limits on the embedding call, cold-start latency, and cost unpredictability. If your dataset fits in memory, keep it in memory. Model quality is a separate concern from architecture. I spent a lot of time optimizing the infrastructure around a model that wasn’t the right fit. The embedding approach with GPT-4.1-nano was fast and cheap. It just wasn’t good. Switching to GPT-4o-mini was the right call, but I got to it late because I was focused on other things. Accept 95% solutions. The embedding approach matches the “right” card about 80% of the time and a “relevant” card about 95% of the time. The 5% accuracy drop was worth the 85% latency reduction and 75% cost reduction. This isn’t a search engine. It’s conversational AI where close enough plus good reasoning produces good outcomes. Closing Thoughts Every situation is different. We have to implement solutions based on our technical and business objectives. Recently, I came across a few click-bait videos on “RAG is dead.” But that’s all click-bait. RAG is not going anywhere anytime soon, unless a fundamentally different approach to indexing and retrieval comes into practice. Pls share your thoughts, ideas and questions. Especially, any other architecture that you have tried and tested. Tell us about your project → We'll follow up to find a time to talk --- ### To AI or Not to AI URL: https://sisyphusconsulting.org/writing/2025/01/16/to-ai-or-not-to-ai/ Date: 2025-01-16 That’s no longer the question. The real question is… actually, it is not a single question. There are questions: What to AI? When to AI? Where to AI? Who will AI? How much to AI? These are reasonable questions, but we are not there yet. When we experience a technological leap, people get divided into believers and non-believers. Embracers and repealers. Pro and against. This is completely natural.1 It is similar to being thrown into a war. A state of uncertainty. Once we come to terms with such shock, we begin to think about better questions. Not for a philosophical discussion. But to understand what we should do with this new situation. Understanding how we fit in the new context. So naturally, when ChatGPT 3.5 made loud and visible noise in November 2022, the world got divided into two groups. A doomsday group and a rosy-glassed group. Some people from both the groups got seriously radicalized. It is important that we briefly talk about what these zealots did or are still doing and will keep doing. On one side, the rosy-glassed zealots began dreaming and conjuring up things that AI could do. But they didn’t ask the question, “should it do that?” Take Microsoft’s much-thrashed AI feature Recall.2 Or take LinkedIn’s AI writing feature and “contribute expertise” harassment. They’re examples of mass experiments on the world because a mad scientist was absorbed by his invention. On the other side, the doomsday zealots began attacking the technology and tech-bros with some of the oldest tricks in the book: regulations, taxes, strikes, and lawfare. For example, most technologically-savvy educational universities have begun regulating the use of AI in their classroom, so much so as to ban it. On a government level, every major economy has begun enacting laws or “framework” to regulate AI to “protect” their citizens.3 So, that’s why the public opinion still vacillates in the response to the question: “To AI or Not To AI?” A faction of public shouts “yessss!” and another faction cries “noooo!” This is not new. Historically, public opinion has been formed by journalists–a category that no longer keeps a journal. Now most journalists are trained to be newsmaker, so it is difficult to expect an intellectually honest opinion-making. For example, when newspapers print that the raising of Short-Term Capital Gain Tax from 15% to 20% is a raise of 5%, we have to understand that the newsmakers have stopped thinking. Which implies we can’t expect them to understand how AI systems work, how tokenization works and communicate it to the public. The public with average IQ of 100. And that leads us to the current dilemma of the tech business founders: “to AI or not to AI?” Businessmen4 live among the public and they’re not isolated. Being influenced by the two extreme sides of public opinion doesn’t help them form a useful opinion about which way to move ahead. They move ahead with doubts: whether they choose to do AI or not to do AI. Generally speaking, if you think about it, doubts are of two types. In the first type of doubt, an individual doubts whether something would work out in the end or not. In the second type of doubt, the individual doubts whether they should do something or not. The second type of doubt is a dilemma–and that’s the main problem. Because when we get afflicted with this second type of doubt about any new technology, it blurs our vision and makes us feel indecisive. This type of doubt is the main obstacle to fulfilling our potential, serving our customer and contributing to the world.5 The first type of doubt is net positive. In reality, the first type of doubt is often the fuel that pushes the ship of progress forward. When determined people are filled with first type of doubt, they become more determined to overcome the doubt by doing something about it. This also means that when tech businessmen are filled with the doubt of whether the AI integration will work for their product, they will figure out the answer by doing things. And eventually, finding an answer. Surprisingly, the answer will not be BLACK AND WHITE. Also, there would not be an answer, but there would be answers. Nuanced ones. They would answer the questions that we posed at the beginning of this essay. Take for example, “What to AI?” After months of struggle, the businessmen’s tech team may find that their image editing software doesn’t require generative AI or the non-profit founder’s tech team would brief them that their beneficiaries prefer chatting with real humans over the AI bot. That’s real world feedback. That’s not coming from a journalist’s ideological utopian world that doesn’t exist. The real world may also answer the “When to AI?” question by telling us when during the customer journey, we should involve the AI. Or when the business should think about doing AI in their growth journey: when they hit 5 years, 50000 customers or $200k revenue? The real world may answer the question “Where to AI?” by telling us which part of the product to do AI and which part to leave untouched. The most important question that we don’t think much about is “How much to AI?” It forces us not to go extreme on either side and stay within reason. AI technology is not a deadly airborne virus around us that we have to think “should I wear a mask or not?” It’s just another technology, like thermonuclear power that can be used for progress.6 Let’s hope that we get out of this rut of choosing between “to AI or not to AI” soon and make some real contribution to the world. Tell us about your project → We'll follow up to find a time to talk I have a theory about this natural reaction. Or what I’m calling a natural reaction. Evolution has instilled the priority of survival in our brains. When we face something absolutely novel, our first question requires binary answer: to flee or to freeze? To stay or to leave? ↩ https://learn.microsoft.com/en-us/windows/ai/apis/recall Recall utilizes Windows Copilot Runtime to help you find anything you’ve seen on your PC. Search using any clues you remember or use the timeline to scroll through your past activity, including apps, documents, and websites. ↩ I have used double quotes to communicate the other meaning that goes along with these words: framework and protect. The recent COVID-era government behaviour across the world has shown the true nature of government’s hunger to hold onto and amass power through censorship, propaganda and force. In the case of AI, the governments have realized the infinite potential of AI for their own use: to surveil, dox, and ultimately control every aspect of the common man. Before the recent AI breakthrough, this wasn’t possible at such an efficiency. ↩ No, that’s not a mistake. Over 90% AI tech business founders are men. Apart from the statistics, the main reason for using “businessmen” is my wish to write what I want to write. ↩ To be clear, I’m not implying that choosing “not to do AI” is equal to “not making progress or not serving our customers.” When the choice of “not to do AI” is backed with reason in the immediate context, it is appreciable decision. ↩ That reminds me of this latest research paper on how after 1986 Chernobyl Nuclear Disaster, the global regulations led by the US substantially increased the cost of Nuclear energy. New Nuclear Power Plants (NPPs) stagnated. Thanks to the global oil lobbyists and targeted propaganda to get the public on their side. Nuclear Power remains the cleanest and most reliable power source till date. But we know the public opinion has been changed through years of false narrative. Find the paper here: https://conference.nber.org/conf_papers/f205791.pdf ↩ --- ### Case Study: ATL Curriculum AI Assistant URL: https://sisyphusconsulting.org/case-studies/2025/01/16/atl-ai-assistant/ Date: 2025-01-16 Let’s take a use case that addresses struggles of many teachers and schools. We have used Sisyphus Valley School—our prototype application to display this integration. Atal Tinkering Lab (ATL) is an initiative by Government of India. These labs are equipped with advanced STEM resources to foster creativity and design thinking. The government also provides a three-level curriculum to guide schools in implementing ATL activities. Take a look at the curriculum files. The challenge There are 10,000 ATL labs installed across India. This makes it difficult to prescribe a curriculum that can cater to a diverse range of schools. This is especially true for teachers. Teachers, with varying levels of competency, often need a helping hand. With their overwhelming day-to-day duties, most teachers need assistance with: Generating fresh ideas to implement the ATL curriculum lessons Aligning activities with the ATL curriculum Modifying activities in the context of the school, students’ prior knowledge, available material, time constraint etc. Our approach We came up with the approach to loop in an AI assistant. An assistant specifically trained on the ATL curriculum. We integrated this assistant with our workflow in a way that the teacher doesn’t need to leave their workspace or provide additional context. Here’s how the idea works as illustrated in the sketch below: The teacher provides their requirements for activities. The additional context is automatically attached with their message and sent to the AI assistant to process. Next, once the AI assistant receives the message from the teacher along with the context, it refers to its ATL curriculum knowledge and returns a response. How the ATL curriculum AI assistant works? In action Check out the PDF slides to see our approach in action. As you review the slides, here are some key highlights: (Slide-4) When the teacher types their message in “Compose message” box, the AI will receive all the information as additional context, except for “Session Detail”. Session Detail box is provided for teacher to store information related to session such as session plans, activities etc. (Slide-6) Once the teacher has sent the message, there is “Sent! AI is thinking…” indicator is being displayed. This is where the AI assistant is working and referring to the curriculum files provided to it in advance. (Slide-9) In real-life classrooms, teachers may not always have sufficient material at hand or may want to adapt ideas from other sessions. They would want to have the freedom to build upon the already generated ideas. That’s exactly what is happening here. A teacher can make “follow up” requests. For example, in this slide, the teacher asks to use clay and popsicle sticks instead of using cardboard. What makes this integration unique? A tech-savvy teacher could navigate to free AI tools like ChatGPT to seek help. But there are some design and culture related requirements for teachers to take action: The teacher should be motivated enough to log into the free ChatGPT or similar tool The school culture should be conducive for the teacher to use such tools publicly and proudly The teacher would have to provide the additional context every time The ideas generated by the AI should be logged to build idea bank If we wait for our already-busy teachers to take initiative, the students may miss out on benefits of AI-aided learning. You would have realized that this setup is unique in itself because: It decreases the friction for a teacher to access AI assistant Thanks to its in-built design, the teacher doesn’t have to jump any more loops to get personalized help The help is always hyper-contextual due to built-in context sharing The organization recognizes the usefulness of the AI assistants and encourages teachers to make the learning more effective The setup saves all past ideas interacted on a specific topic. Providing a gold-mine of idea bank for future The win here isn’t the AI; it’s the workflow. Teachers don’t need to leave their workspace, learn a new tool, or repeat the context each time. The AI sits inside the work they already do. Tell us about your project → We'll follow up to find a time to talk --- ### Case Study: Data Analyst AI Assistant URL: https://sisyphusconsulting.org/case-studies/2025/01/16/data-analyst-ai/ Date: 2025-01-16 Let’s take a use case that addresses the hidden struggles and frustrations of many business owners and managers. We have used Sisyphus Valley School–our prototype application to display this integration. The challenge Most business owners and managers depend on their people or their dashboards to get business insights. They receive these insights in the form of a report or statistics or a photo of an interactive dashboard. Many times, the people managing the whole business want to explore different scenarios. They want to talk to their data and get answers first-hand. But in current generation of data systems, it is nearly impossible to: Play with their data Connect one data point to its related data point for analysis Focus solely on business questions instead of analysis techniques Many business owners do conduct their own exploration using tools such as PivotTables in Microsoft Excel or basic dashboard manipulation. But there are two issues in this manual approach: first, the onus of getting the analysis right lies on the technical competency of the owner and second, a dashboard’s analysis is limited by its design. Our approach We figured out a way where a business owner can only focus on their business questions and potential insights. They don’t need to worry about whether they selected right Column label in PivotTable. Or feel restricted by their software dashboard. We trained AI assistant (let’s call it Data Analyst AI) on our database schema, i.e. the database structure. This means, the AI doesn’t have access to the data, just its structure. Look at the sample structure of a Parent’s data element for the Sisyphus Valley School in the following document. AI has this type of schema information with it So, the Data Analyst AI knows what type of data and internal relationships different elements have. When we ask it to come up with an analysis, it refers to them and comes up with a raw analysed data. In action In this prototype, the software we are using is Microsoft Access. It uses Structured Query Language (SQL) for storage and data analysis. SQL is one of the most widely used database language in the world. And a business owner or a manager doesn’t need to learn SQL to analyze their data. They can simply ask their business question and the AI will create a SQL query for the system to process. Data Analyst Assistant AI Screen The system generates a raw table in answer to our question. We can ask any question related to our data and it will present the first-hand data, without any bias, filter or assumptions. Data Analyst Assistant AI Result The uniqueness of such integration Smart business owners and managers want to get their hands dirty. They are interested in seeing the data first-hand. This AI integration provides them with numerous benefits: It increases their independence when it comes to retrieving insights It allows them privacy and freedom to ask wrong questions It frees them from the rigidness of dashboards It makes it easy to find blind spots of their insights and assumptions Why it matters The bottleneck used to be: who do you trust to query your data for you? With this approach, you ask the question directly. The AI writes the SQL, the database returns the answer, and you see the actual numbers. Tell us about your project → We'll follow up to find a time to talk --- ### Case Study: Sachet Foundation's Data System URL: https://sisyphusconsulting.org/case-studies/2025/01/16/sachet-foundation/ Date: 2025-01-16 Data system design for adolescent children Sachet Foundation's Data System Home Screen In August 2023, we were retained by Sachet Foundation, a budding non-profit helping SORTed adolescents become employable. SORT stands for Slums, Orphanages, Rural and Tribal adolescents. They hired us to create a customized data system that would support their intervention being deployed in Snehalaya. By the end of our engagement with Sachet Foundation, we: Optimized their intervention model structure Optimized their intervention processes Helped them identify and structure data points required for their successful intervention Created an MS Access-based data system for smooth operation Trained and educated the staff members in usage of their system What exactly we did for Sachet Foundation When they hired us, the founders knew what type of data system they needed. They needed to develop their ideas to monitor their beneficiaries with 360 degree view using a data system. So first, we interviewed the founders, understood their intervention model, visited their beneficiaries and work location. These interactions helped us suggest modifications to their intervention model and processes with two things in mind: In addition to their proprietary subjective methods, the founders wanted to have objective, evidence-based processes in their decision making. The founders wanted to consider future scaling and tracking of their beneficiaries. Next, we gathered all the necessary data points that were must-haves and good-to-haves. Structured them, ran by the founders and the team. And developed the database. Finally, we went through several cycles of development and delivered Sachet Foundation’s data system. This system allowed their staff and new hires to think in terms of designated processes—something that is by design extremely difficult in Excel sheets. Want custom software built around your business? It is difficult to build software that fits how your business actually operates. What works for a fortune 500 company may not work for your business. Each business is unique. That is what we do. We understand your processes, suggest changes and build software that works for you. See how we do it and get started with us: Our process and getting started → While you check out other case studies and our processes, I want to invite you to join our semi-regular free email letters. Loading newsletter… --- ### The problem with averages and dashboards URL: https://sisyphusconsulting.org/writing/2024/10/09/the-problem-with-averages-and-dashboards/ Date: 2024-10-09 Tell me something, would you step into a river that is on average 4 feet deep? Take a pause. If your mind shouted NO right away, you don’t need to look at this illustration: Averages are tricky. Most of the times, they play with our intuition and lead us to bad decisions. One of my favourite averages is the recent Securities and Exchange Board of India’s (SEBI) research outcome: approximately 93% of equity derivative traders have incurred an average loss of Rs 2 lakh (per trader) during the last three financial years. Imagine some of the deadly depths of this loss river! Now, that’s a river I wouldn’t dare step into without expertise. In any case, the knowledge work that most of us are involved doesn’t contain derivative trading or stepping into rivers. But many of us deal with dashboards everyday. Dashboard mania Most of the organizations have been in frenzy to have dashboards for sometime now. This began a few years back due to increased computing power of PCs, better internet connectivity and upgradation of technology stack of webpages. A few months back, we had discussed different aspects of the problems with dashboards and mistakes in 5 mistakes I avoid while designing my dashboards The three problems I had highlited were: Dashboards are treated as cure to all problems Only for managers Gap between decision-maker and decision But the main problem: averages Other organizational, humane and management problems aside, the main problem is that most dashboards show averages. Consider this image that I had shared in the said letter: In your business, the real data might be coming from 100 different locations from a large geographic area. As a business owner or a high-level manager, the dashboard would be providing an aggregate, average idea of the the business activity health. 4-feet depth on display Your dashboard shows you the 4 feet depth of the river. Of course, the modern dashboards provide the facility to drill several levels deeper (but how often do you click 5 levels deep and corroborate the data?). In other words, dashboards tend to do two things: 1. Making good business look bad If the average sales data of your 50 sales people show bad business, that overshadows the good business brought in by the individuals. 2. Making bad business look good On the other hand, if the average sales data of your 50 sales people show good business, it hides the bad business brought in by the individuals. In essence, averages play with our intuition, mood and consequently decision-making. So what’s the solution? Giving more power to your people. By investing in design that provides a personalized experience. Most importantly, details that don’t show averages. Look at this stock management screen I created sometime back. Ignore the discrepancy in the sample data and look at the quantity of Curd. It doesn’t tell person that 75% curd is available. Instead, it tells the person that the curd status is GREEN at the moment and it is 750 gram. The person will make necessary the decision depending on this information. Averages can mislead, in rivers and in dashboards. Show the actual numbers, both the high and the low, and let people decide what they mean. Want custom software built around your business? It is difficult to build software that fits how your business actually operates. What works for a fortune 500 company may not work for your business. Each business is unique. That is what we do. We understand your processes, suggest changes and build software that works for you. See how we do it and get started with us: Our process and getting started → While you check out other case studies and our processes, I want to invite you to join our semi-regular free email letters. Loading newsletter… --- ### 5 mistakes I avoid while designing my dashboards URL: https://sisyphusconsulting.org/writing/2024/04/05/5-mistakes-I-avoid-while-designing-my-dashboards/ Date: 2024-04-05 If you tell me you haven’t seen a dashboard in your career so far, I wouldn’t believe you. In fact, you wouldn’t believe yourself. The ease with which data visualization tools are available, it is impossible to believe that you would be spared from a dashboard. Mostly, this is how the tragedy of dashboards look like: If I’m calling it a tragedy, I may not be holding a favorable view to most dashboard practices. And I would like to share a few of them before we move ahead. Problems with most dashboards Three glaring problems that make dashboards frustrating: 1. Treated as cure to all problems Creating dashboards do not substitute problem solving. It is like looking at a compass on a stranded boat and waiting to reach a nearby shore. On its own. Well-designed dashboards can guide where to look, but they can’t act as a cure. 2. Only for managers There is a whole new clan of dashboard-hierarchy that depends on earning a living based on dashboard-dictated-work and meetings. A traditional manager or a supervisor was expected to know the work inside out, provide feedback where necessary. But the emergence of dashboard-managers is a new phenomenon. Companies cannot generate money with pseudo-workers. And they’re realizing it slowly. 3. Gap between decision-maker and a decision Most dashboards are far from the user’s immediate decision-making power. If I can’t change a feature on a product, I don’t need to know which features are used by the customers. My dashboard should prioritize insights that help my immediate decision-making. Good-to-know things can wait. In a way, these three are also mistakes that I try to avoid. But let’s get on the five mistakes that I reallyyy try to avoid. The five mistakes I try to avoid 1. I don’t provide decimal level information aka try to be accurate Dashboards are to understand trends. They do not require military precision. So, if a Gender-bifurcation pie chart says 45.6% women and 54.4% men, what am I going to do with that .6 and .4 ? 2. I don’t ignore tables It is not necessary to have only colourful charts on dashboards. Tables are one of the best visual tools to summarize information. Add tables wherever it feels natural to add tables. 3. I don’t try to give all teeny-tiny details through the dashboards That’s not the purpose. An overview of information is more than enough for a smart decision-maker to ask valid questions or sieve through data. People who are really interested and invested in knowing the full picture, they usually get their hands dirty by asking for full data. 4. I don’t try to objectify subjective matter Sometimes, it is simply not possible to do it. And should not be done. Whoever is interested, should take interest. Also, thanks to the onset of new AI tools, this work is now even easier than before. That reminds me of Prichard Scale of Understanding Poetry covered in Dead Poets Society. It tries to evaluate the “greatness” of the poem. How ridiculous is that! 5. I don’t treat dashboard as a god, the ultimate piece of truth Because dashboards reflect the underlying databases and calculations. There are more ways to structure bad databases than there is to structure good databases. This also means that most of the databases are not perfect. Their shortcomings are reflected in dashboards. If people find something wrong with the dashboards, it is an opportunity to look at how we are processing data. And improve the dashboards. The point is that dashboards are everywhere, but not all are effective. Many fall into traps: being seen as cure-alls, catering only to managers, or creating gaps between data and decisions. We should treat dashboards as a skepticle compass, not an engine. Want custom software built around your business? It is difficult to build software that fits how your business actually operates. What works for a fortune 500 company may not work for your business. Each business is unique. That is what we do. We understand your processes, suggest changes and build software that works for you. See how we do it and get started with us: Our process and getting started → While you check out other case studies and our processes, I want to invite you to join our semi-regular free email letters. Loading newsletter… ---