面试准备一些资料
This commit is contained in:
168
Cloud DevOps/ITOM ESM Cloud Service Catalog.md
Normal file
168
Cloud DevOps/ITOM ESM Cloud Service Catalog.md
Normal file
@@ -0,0 +1,168 @@
|
||||
# ITOM Cloud Service Delivery : ITOM ESM Cloud Service Catalog
|
||||
|
||||
Created by Wei Shen, last modified by Adina Lehene on Sep 05, 2025 EDT
|
||||
|
||||
## Introduction
|
||||
|
||||
This document lists all the cloud services that are currently supported by ESM Cloud Service Team.
|
||||
|
||||
If the product team needs to add new Cloud services to customers, please refer to: [ITOM Cloud Service Delivery Approval Process for New Services](ITOM-Cloud-Service-Delivery-Approval-Process-for-New-Services_688996646.html)
|
||||
|
||||
## Service Catalog
|
||||
|
||||
### Product Cloud Services
|
||||
|
||||
| | | | | | | | | |
|
||||
|---|---|---|---|---|---|---|---|---|
|
||||
|Category|Serivce Name|Description|Audience|Service Request method|Related Documents|Service Level Target|Service Owner|Status|
|
||||
|Trial Service|**Request a SMAX Premium Trial**|This Offering is used for requesting a SMAX Premium trial tenant|Opentext Internal|Submit Service Request via X4X|No ops eng intervention|3 business days|CSD|IN USE|
|
||||
|**Request an ESM SaaS Trial/PoC Tenant (SMAX, AMX, CMS, OO)**|If you require a Tenant that includes more than just SMAX, you can use this offering to request. ( for HCMX please use the other offering in our Service Catalog)|Opentext Internal|Submit Service Request via X4X|[FAQ for SaaS Product Trials](https://us2-smax.saas.microfocus.com/saw/ess/viewResult/1369241?TENANTID=202385354)<br><br>[How to provision](https://confluence.opentext.com/pages/viewpage.action?pageId=686088202) <br>[X4X offering](https://us2-smax.saas.microfocus.com/saw/ess/offeringPage/2307287?TENANTID=202385354)|3 business days|CSD|IN USE|
|
||||
|**Request an ESM SaaS Trial/PoC Tenant (UD Standalone)**|If you require a Tenant that includes only Universal Discovery and CMDB (no native SACM, SAM, no more post configuration), you can use this offering to request.|Opentext Internal|Submit Service Request via X4X|[FAQ for SaaS Product Trials](https://us2-smax.saas.microfocus.com/saw/ess/viewResult/1369241?TENANTID=202385354)<br><br>[X4X offering](https://us2-smax.saas.microfocus.com/saw/ess/offeringPage/2827026?TENANTID=202385354)|3 business days|CSD|IN USE|
|
||||
|**Request a HCMX SaaS Trial/PoC Tenant**|Please use this offering to request a HCMX Trial/PoC tenant|Opentext Internal|Submit Service Request via X4X|[FAQ for SaaS Product Trials](https://us2-smax.saas.microfocus.com/saw/ess/viewResult/1369241?TENANTID=202385354) <br>[How to provision](https://confluence.opentext.com/pages/viewpage.action?pageId=686088202) <br>[X4X offering](https://us2-smax.saas.microfocus.com/saw/ess/offeringPage/2307287?TENANTID=202385354)|3 business days|CSD|IN USE|
|
||||
|**Request Customer Web Trial**|This Offering is used for requesting a SMAX only web trial for external customer|External Trial Customer|[Opentext web portal](https://www.microfocus.com/en-us/products/service-management-automation-suite/free-trial)|[FAQ for Customer Web Trial Process](https://us2-smax.saas.microfocus.com/saw/ess/viewResult/1363194?TENANTID=202385354)|3 business days|CSD|IN USE|
|
||||
|**Request IT Operations Aviator Trial (SMAX/OPSBRIDGE)**|Please use this offering to request Aviator either to be added to an existing SMAX Trial / PoC Tenant or to request Aviator for OpsBridge.|Opentext Internal|Submit Service Request via X4X|[Add the Aviator capability](https://us2-smax.saas.microfocus.com/saw/ess/offeringPage/2337847?TENANTID=202385354) <br>[Aviator-OpsB Confl. art.](https://confluence.opentext.com/pages/viewpage.action?spaceKey=ICSD&title=Aviator+widget+on-boarding+tasks+for+OpsB)|3 business days|CSD|IN USE|
|
||||
|**Extend duration for a specific Trial (SMAX, AMX, HCMX)**|Request to extend the duration for specific Trial tenant.|Opentext Internal|Submit Service Request via X4X|[FAQ for SaaS Product Trials](https://us2-smax.saas.microfocus.com/saw/ess/viewResult/1369241?TENANTID=202385354)|2 business days|CSD|IN USE|
|
||||
|**Request Tenant Decomission**|Collect the tenant that is not used any more. Submit request to decommission the Trial/Poc tenant|Opentext Internal|Submit Service Request via X4X|[FAQ for SaaS Product Trials](https://us2-smax.saas.microfocus.com/saw/ess/viewResult/1369241?TENANTID=202385354) <br>[X4X offering](https://us2-smax.saas.microfocus.com/saw/ess/offeringPage/1628360?TENANTID=202385354)|5 business days|CSD|IN USE|
|
||||
|Version Upgrade Service|**Product Major Version Upgrade on ESM Cloud Farms**|Planned Standard Change|External Paid Customers|Submit Change Request in OT SM9|[Product Version Upgrade](https://confluence.opentext.com/display/ICSD/Product+Version+Upgrade) <br>[OT Doc SMAX 25.3 sample](https://docs.microfocus.com/doc/423/25.3/upgrade)|According to the change plan|CSD|IN USE|
|
||||
|**Product Patch Upgrade on ESM Cloud Farms**|Planned Standard Change|External Paid Customers|Submit Change Request in OT SM9|- [Patch Cloud Deployment Process](Patch-Cloud-Deployment-Process_686087749.html)|According to the change plan|CSD|IN USE|
|
||||
|**Apply Urgent Hotfix on ESM Cloud Farms**|Unplanned Production Change|External Paid Customers|Submit unplanned production change request via X4X|- [Request Unplanned Change in Cloud Production Environment Process](Request-Unplanned-Change-in-Cloud-Production-Environment-Process_686070239.html)|According to the change plan|CSD|IN USE|
|
||||
|**EKS AMI Rotation on ESM Cloud Farms**|This service is used to periodically update the EKS worker node server AMI to meet security requirements|External Paid Customers|Submit Change Request in OT SM9||According to the change plan|CSD|IN USE|
|
||||
|Cloud Deployment Service|**Deploy a new ESM (SMAX+CMS+OO+HCMX/FinOps) Cloud Farm**|This service is based on the business need to deploy a new ESM Farm and complete all the productionized tasks which include:<br><br>- Cloud Applicaiton monitoring<br>- Service Availbility monitoring<br>- Configure AWS backup plan<br>- Operation Automation to support this new farm<br>- Tenant Provision Automation to support this new farm<br>- etc.|External Paid Customers|Business Driven|- [ESM Cloud Farm Construction](ESM-Cloud-Farm-Construction_688988187.html)|According to the business plan<br><br>It takes about 1 month from project establishment to budget approval to final farm delivery|CSD|IN USE|
|
||||
|Monitoring Service|**Product Major Functionalities Service Avalibility Check**|This service means that when there is a new farm, or a new capbility in the product that requires a customized Service Availability check, the Cloud Servcie team needs to assist the product team and work with the APM Scripting team, Service Center team to configure and implement the APM monitoring.|External Paid Customers|For new capability or product that need to have service availbility check please contact Cloud Service team|- [APM Monitoring Business Flow](APM-Monitoring_686073667.html)|24x7|CSD|IN USE|
|
||||
|**ESM Service Health Web App**|- Real-time major functionalities service healthy status<br>- Publish monthly SLA result<br>- Publish major incident report<br>- Publish planned maintanance window schedule|External Paid Customers||- [Service Health Page](Service-Health-Page_686084001.html)<br>- [ESM Monthly SLA Result](Monthly-SLA_686070031.html)|24x7|CSD|IN USE|
|
||||
|Disaster Recovery Service|**Disaster Recovery Service**|According to the SaaS service description, when a disaster occurs in the region where the cloud application is located, cloud application recovery and data recovery can be performed in other available areas based on the remote-region backup data. Recover the customer's business within the committed RPO/RTO|External Paid Customers||(Document need to be updated)|24x7|CSD|IN USE|
|
||||
|**Disaster Recovery Integrity Testing and Report**|According to the SaaS service description, conduct DR validation testing regularly and provide relevant testing report|External Paid Customers||(Document need to be updated)|N/A|CSD|IN USE|
|
||||
|**Cross Region & Cross AWS Account data backup service**|Cross-account, cross-region data (AWS RDS, AWS EFS, K8S velero, AWS S3) backup service for all deployed customer production farms|External Paid Customers||(Document need to be updated)|N/A|CSD|IN USE|
|
||||
|Security Service|**Provide service to apply OS security patch or apply remediation to all Cloud farms or AWS account according to security scan result**|Implement Qualys, Prisma scans on each production AWS account and based on the results of the scans perform the necessary remediation work according to priority|External Paid Customers||- [Process on how to handle Security Issues found by Qualys Scan](Process-on-how-to-handle-Security-Issues-found-by-Qualys-Scan_688996390.html)|N/A|CSD|IN USE|
|
||||
|**Update Cloud Application WAF rules, and provide WAF logs**|Regularly update Cloud Application WAF rules, and provide WAF logs to facilitate product engineering team observation and improve WAF protection mechanisms|External Paid Customers||(Document need to be updated)|N/A|CSD|IN USE|
|
||||
|Cost Optimization Service|**Conduct regular multi-dimensional audits of account spending and identify where savings can be made**|From time to time, based on the results of daily audits and review to find opptunity to save AWS cost and covert to tasks to be executed by Cloud Service team|External Paid Customers||- [Cloud Cost Optimization/FinOps](686065517.html)|N/A|CSD|IN USE|
|
||||
|Cloud Readiness Service|**New product/Product new capability Cloud Readiness Review Service**|This service is used to discuss and review the cloud-readiness of new product or new capabilities derived from products.|Opentext Internal|Business Driven|- [Cloud Application Cloud Readiness Check List](Cloud-Application-Cloud-Readiness-Check-List_682933055.html)|According to the business plan|CSD|IN USE|
|
||||
|
||||
### Customer Cloud Services
|
||||
|
||||
| | | | | | | | | |
|
||||
|---|---|---|---|---|---|---|---|---|
|
||||
|Category|Serivce Name|Description|Audience|Service Request method|Related Documents|Service Level Target|Service Owner|Status|
|
||||
|Customer Cloud Service Offering|**SMAX: Add an integration user**|This service is to assist customer to create integration user for various external intergrations with SMAX tenant|External Paid Customers|Submit Service Request via PCS|- [Create Integration Users](Create-Integration-Users_686065319.html)|2 business days|CSD|IN USE|
|
||||
|**ESM: License update**|This service is to assist customer to renew license or update license allocation|External Paid Customers|Submit Service Request via PCS|- [Apply license to ESM customer tenant](Apply-license-to-ESM-customer-tenant_688996779.html)|5 business days|CSD|IN USE|
|
||||
|**SMAX: Configure SAML authentication**|This service is to assist customer to configre SAML SSO authentication for SMAX tenant.|External Paid Customers|Submit Service Request via PCS|- [Configure SAML authentication for SaaS Customer](Configure-SAML-authentication-for-SaaS-Customer_686065288.html)|5 business days|CSD|IN USE|
|
||||
|**SMAX: Configure OAuth authentication**|Configure the OAuth authentication for the OpenText SaaS customer account|External Paid Customers|Submit Service Request via PCS|- [Add OAuth authentication - Ops+Customer tasks](684947018.html)|5 business days|||
|
||||
|**SMAX: Configure custom domain (NLZ)**|This service is to assist customer to configure custom domain for SMAX tenant|External Paid Customers|Submit Service Request via PCS|- [Configure SMAX custom domain (New Landing Zone)](686065305.html)|5 business days|CSD|IN USE|
|
||||
|**SMAX: Configure custom mail sender**|This service is to support customer configure their own custom email address as sender in SMAX tenant but also for customers who want to change the default "Email From Name"|External Paid Customers|Submit Service Request via PCS|[Configure custom mail sender, dedicated AWS SES users](686065263.html) <br>[OT Doc](https://docs.microfocus.com/doc/423/25.2/configcustomemailaddress)|5 business days|CSD|IN USE|
|
||||
|**SMAX: Customize login/logout screen**|This service is to configure the theme settings of the login and logout pages for customer tenant to suit customer company's look and feel|External Paid Customers|Submit Service Request via PCS|- [Customize the login and logout pages](Customize-the-login-and-logout-pages_686065324.html)|5 business days|CSD|IN USE|
|
||||
|**SMAX: Decommission customer tenant**|This service occurs when a customer exits ESM Cloud and needs to comply with the relevant customer exit process to perform relevant tasks.|External Paid Customers|Submit Service Request via PCS|- [ESM Cloud Customer Exit Process](ESM-Cloud-Customer-Exit-Process_686070016.html)<br>- [ESM Customer Tenant Decommission](ESM-Customer-Tenant-Decommission_688996785.html) <br> (Document need to be updated)|5 business days|CSD|IN USE|
|
||||
|**ESM: Enable ESM capabilities (CMS/OO/FinOps/AC)**||External Paid Customers|Submit Service Request via PCS|- [Enable ESM capabilities (UCMDB/OO/FinOps/AC/OP/ODL)](688996783.html)|5 business days|CSD|IN USE|
|
||||
|**SMAX: Maintain customized language package**||External Paid Customers|Submit Service Request via PCS|- [SMAX maintain custom language packs](SMAX-maintain-custom-language-packs_688996787.html) <br> (Document need to be updated)|5 business days|CSD|IN USE|
|
||||
|**SMAX: Modify allowed attachement file types**|This service is to support customer to add allow attachment types in their tenant|External Paid Customers|Submit Service Request via PCS|- [Allowable SMAX Attachment Extensions](Allowable-SMAX-Attachment-Extensions_686065217.html)|5 business days|CSD|IN USE|
|
||||
|**SMAX: Modify maximum attachement size**||External Paid Customers|Submit Service Request via PCS|- [SMAX modify maximum attachement size](SMAX-modify-maximum-attachement-size_688996790.html) <br> (Document need to be updated)|5 business days|CSD|IN USE|
|
||||
|**SMAX: Customer domain certificate renewal**|Based on the customer SMAX tenant that has configured the custom domain, if the custom domain certificate expires, can use this service to renew domain certificates|External Paid Customers|Submit Service Request via PCS|- [Configure SMAX custom domain (New Landing Zone)](686065305.html)|5 business days|CSD|IN USE|
|
||||
|**SMAX: Configure IP allowing list service**|For customers who have purchased the IP allowlisting service, please use this offering to provide the list of IP addresses or IP ranges to be configured by OpenText to control the access to your tenant.|External Paid Customers|Submit Service Request via PCS|(Document need to be updated)|5 business days|CSD|IN USE|
|
||||
|**ESM: Request a new ESM Dev/QA tenant**||External Paid Customers|Submit Support/Service Request via PCS|(Document need to be updated)|5 business days|CSD|IN USE|
|
||||
|**ESM: Request Power BI gateway**|This service is to assist customers in generating FinOps reports using Power BI|External Paid Customers|Submit Support/Service Request via PCS|- [Integrate with Power BI to create FinOps reports](Integrate-with-Power-BI-to-create-FinOps-reports_686065345.html)|5 business days|CSD|IN USE|
|
||||
|**UCMDB: Disable Native SACM and enhanced CI in SaaS**|This service is to disable Native SACM feature between SMAX and UCMDB and also disable enhanced CI in UCMDB|External Paid Customers|Submit Support/Service Request via PCS|- [Disable NSACM and enhance CI lifecycle in SaaS](Disable-NSACM-and-enhance-CI-lifecycle-in-SaaS_688987700.html)<br>- [Disable Native SACM manually](Disable-Native-SACM-manually_686073918.html)|5 business days|CSD|IN USE|
|
||||
|**Aviator: Enable ITOM Aviator for ESM cloud tenant**|Enable ITOM Aviator AI Service on existing customer's SMAX/UCMDB tenant|External Paid Customers|Submit Support/Service Request via PCS|- [Enable ITOM Aviator for ESM tenant](Enable-ITOM-Aviator-for-ESM-tenant_688996800.html) <br> (Document need to be updated)|5 business days|CSD|IN USE|
|
||||
|**Aviator: Setup Configuration for hybrid mode**|For customers that have purchased Hybrid Aviator (Connecting our OpenText Public Cloud Aviator Service to your on premise or private cloud SMAX tenant), please submit your details below so we can start the provisioning process.|External Paid Customers|Submit Support/Service Request via PCS|- [Enable ITOM Aviator for SMAX on-premise customer](Enable-ITOM-Aviator-for-SMAX-on-premise-customer_688996802.html) <br> (Document need to be updated)|5 business days|CSD|IN USE|
|
||||
|**Operation Platform: Enable Operation Platform capability (OP/UIS/ODL)**|Enable Operation Platform on top of existing ESM customer tenant|External Paid Customers|Submit Support/Service Request via PCS|- [Operations Platform tenant enablement](Operations-Platform-tenant-enablement_688996278.html)<br>- [Enable Optic Data Lake](Enable-Optic-Data-Lake_688996343.html)|5 business days|CSD|IN USE|
|
||||
|Customer Communication|**Customer Communication**|Customer communication via emails, service health page, PCS news to communicate planned standard changes|External Paid Customers|Send email notification via PCS<br><br>Publish news in Service Health Page|- [Send email notification to SaaS customers via PCS](Send-email-notification-to-SaaS-customers-via-PCS_686069617.html)|N/A|CSD|IN USE|
|
||||
|Customer Cloud Onboarding|**New SaaS Customer Onboarding Service**|- SaaS order Fulfillment<br>- Product License Generation<br>- Provision Customer Tenants (Prod/Dev)<br>- Allocate license to customer tenants<br>- Tenant initial configuraiton<br>- Send customer notification with tenant detail<br>- Welcome call (Handled by CSM)|External Paid Customers||- [ESM Tenant Provisioning Automation](ESM-Tenant-Provisioning-Automation_686079418.html)<br>- [ESM license generation detail](ESM-license-generation-detail_686070325.html)<br>- [ESM products licensing provisioning (SMAX/HCMX, UCMDB/CMS/UD, OO)](686070266.html)|2 business days|CSD|IN USE|
|
||||
|**Off-Cloud customer migrate to ESM Cloud Farms**|Operation tasks to support off-cloud customer migrate to ESM cloud farm. Including tenant data import, EFS data import etc.|External Paid Customers||(Document need to be updated)|According to the business plan|CSD<br><br>CLOUD RND|IN USE|
|
||||
|Incident Management|**Respond to Major Incident and follow established runbooks to quickly restore SaaS services**|When a service outage occurs in the cloud environment, intervene as soon as possible and restore the service as quickly as possible according to the existing ops runbook.|External Paid Customers||[Major Incident Management Process](Major-Incident-Management-Process_686083938.html) <br>[RnD WIki SaaS Coverage](https://rndwiki.houston.softwaregrp.net/confluence/pages/viewpage.action?spaceKey=SMAXaaS&title=ESM%20SaaS%20RnD%20Coverage)|24x7|CSD|IN USE|
|
||||
|**Major Incident RCA Tracking & Customer Notification**|Root cause analysis for major incidents occurring in the Cloud environment. Work with the engineering team to develop corrective action & preventive action|External Paid Customers||- [ESM Cloud Incident Tracking List](ESM-Cloud-Incident-Tracking-List_686083932.html)|24x7|CSD<br><br>CLOUD RND|IN USE|
|
||||
|
||||
### Internal Cloud Services
|
||||
|
||||
| | | | | | | | | |
|
||||
|---|---|---|---|---|---|---|---|---|
|
||||
|Category|Serivce Name|Description|Audience|Service Request method|Related Documents|Service Level Target|Service Owner|Status|
|
||||
|Internal Cloud Service for product engineering|**Uplanned Production Change Service**|This service is used to handle the current unplanned changes on Cloud production environments other than planned regular major version pgrade & patch deployment . According to the current practice, these actions include but are not limited to:<br><br>- Emergency unplanned hotfix<br>- Data change in production database (not included in patch/upgrade)<br>- Unplanned configuration change in production application<br>- Unplanned application K8S configuration change (Adjust pod number, pod size, yaml configuration etc.)<br>- Unplanned WAF Change<br>- etc.|Product Engineering|Submit unplanned production change request via X4X|- [Request Unplanned Change in Cloud Production Environment Process](Request-Unplanned-Change-in-Cloud-Production-Environment-Process_686070239.html)||CSD|IN USE|
|
||||
|**Operational Document Review & Approval Service**|This service is used to review and approve a variety of Ops documents submitted from the RnD Team|Product Engineering|Submit Ops doc|- [ITOM Cloud Service Ops Doc Management Process](ITOM-Cloud-Service-Ops-Doc-Management-Process_686069689.html)||CSD|IN USE|
|
||||
|**Review & Approval for New Service**|This process applies to all new cloud services introduced due to product upgrades, feature additions, or customer requirements. No service shall be made available to customers without prior approval from the Cloud Service Delivery team.|Product Engineering|Submit new service request form|- [ITOM Cloud Service Delivery Approval Process for New Services](ITOM-Cloud-Service-Delivery-Approval-Process-for-New-Services_688996646.html)||CSD|IN USE|
|
||||
|
||||
### EU-Managed Cloud Services
|
||||
|
||||
| Category | Serivce Name | Description | Audience | Service Request method | Related Documents | Service Level Target | Service Owner | Status |
|
||||
| ------------------------ | ---------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -------- | ---------------------- | ----------------- | -------------------- | ------------- | ------ |
|
||||
| EU-Managed Cloud Service | **EU-Managed Cloud Service** | EU-Managed cloud service specifically refers to the special cloud service provided for specific EMEA customers, with specialized EMEA Cloud Ops engineer (EU residents) to perform various operations. | | | | | CSD | IN USE |
|
||||
|
||||
### FedRAMP Cloud Services
|
||||
|
||||
|Category|Serivce Name|Description|Audience|Service Request method|Related Documents|Service Level Target|Service Owner|
|
||||
|---|---|---|---|---|---|---|---|
|
||||
|||||||||
|
||||
|||||||||
|
||||
|||||||||
|
||||
|||||||||
|
||||
|||||||||
|
||||
|
||||
Contents
|
||||
|
||||
- [Introduction](#ITOMESMCloudServiceCatalog-Introduction)
|
||||
- [Service Catalog](#ITOMESMCloudServiceCatalog-ServiceCatalog)
|
||||
- [Product Cloud Services](#ITOMESMCloudServiceCatalog-ProductCloudServices)
|
||||
- [Customer Cloud Services](#ITOMESMCloudServiceCatalog-CustomerCloudServices)
|
||||
- [Internal Cloud Services](#ITOMESMCloudServiceCatalog-InternalCloudServices)
|
||||
- [EU-Managed Cloud Services](#ITOMESMCloudServiceCatalog-EU-ManagedCloudServices)
|
||||
- [FedRAMP Cloud Services](#ITOMESMCloudServiceCatalog-FedRAMPCloudServices)
|
||||
|
||||
Document Information
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
|**Product:**|ESM|
|
||||
|**Applicable Product Version:**|All Versions|
|
||||
|**Document Author:**|[Wei Shen](https://confluence.opentext.com/display/~wshen)|
|
||||
|**Created Date:**|01 May 2024|
|
||||
|**Reviewed By:**|[Adina Lehene](https://confluence.opentext.com/display/~alehene)|
|
||||
|**Reviewed Date**|05 Sep 2025|
|
||||
|**Approved By:**||
|
||||
|**Approved Date:**||
|
||||
|**Document Authority:**|CLOUD OPS RND|
|
||||
|
||||
Change History
|
||||
|
||||
|Version|Published|Changed By|Comment|
|
||||
|---|---|---|---|
|
||||
|**[CURRENT](viewpage.action?pageId=688996649) (v. 18)**|Sep 05, 2025 06:23 EDT|**[Adina Lehene](/display/~alehene)**||
|
||||
|[v. 17](viewpage.action?pageId=718117022)|Sep 05, 2025 04:59 EDT|**[Adina Lehene](/display/~alehene)**||
|
||||
|[v. 16](viewpage.action?pageId=718116762)|Sep 05, 2025 03:50 EDT|**[Adina Lehene](/display/~alehene)**||
|
||||
|[v. 15](viewpage.action?pageId=718116580)|Sep 03, 2025 04:38 EDT|**[Paul Badiu](/display/~pbadiu)**||
|
||||
|[v. 14](viewpage.action?pageId=716273937)|Mar 17, 2025 20:32 EDT|**[Wei Shen](/display/~wshen)**||
|
||||
|
||||
[Go to Page History](/pages/viewpreviousversions.action?pageId=688996649)
|
||||
|
||||
**Related pages**
|
||||
|
||||
- Page:
|
||||
|
||||
[ESM Cloud Farm Version Tracking](/display/ICSD/ESM+Cloud+Farm+Version+Tracking)
|
||||
|
||||
- Page:
|
||||
|
||||
[How to get an Opentext Confluence account](/display/ICSD/How+to+get+an+Opentext+Confluence+account)
|
||||
|
||||
- Page:
|
||||
|
||||
[ITOM APM AppPluse Cloud Farm Information](/display/ICSD/ITOM+APM+AppPluse+Cloud+Farm+Information)
|
||||
|
||||
- Page:
|
||||
|
||||
[ITOM Cloud Service Ops Doc Management Process](/display/ICSD/ITOM+Cloud+Service+Ops+Doc+Management+Process)
|
||||
|
||||
- Page:
|
||||
|
||||
[ITOM ESM Cloud Service Catalog](/display/ICSD/ITOM+ESM+Cloud+Service+Catalog)
|
||||
|
||||
- Page:
|
||||
|
||||
[ITOM OpsB NOM Cloud Service Catalog](/display/ICSD/ITOM+OpsB+NOM+Cloud+Service+Catalog)
|
||||
|
||||
- Page:
|
||||
|
||||
[OpsB and NOM Cloud Deployments Version Tracking](/display/ICSD/OpsB+and+NOM+Cloud+Deployments+Version+Tracking)
|
||||
|
||||
|
||||
|
||||
|
||||
Document generated by Confluence on Sep 15, 2025 22:24 EDT
|
||||
|
||||
[Atlassian](https://www.atlassian.com/)
|
||||
40
Cloud DevOps/Production Outages.md
Normal file
40
Cloud DevOps/Production Outages.md
Normal file
@@ -0,0 +1,40 @@
|
||||
#production #outage #incident #saas #p0
|
||||
1. App issues
|
||||
|
||||
1. Customer API integration taking too much workload (Happens without notification, when 2-3 or even more customers enable the API integration, it directly impacts the performance of the farm)
|
||||
|
||||
2. RDS out of credits (queue depth 2→20, no issue, 20→80, dramatic slow down)
|
||||
|
||||
3. EC2 out of network bandwidth (logging causing session hanging)
|
||||
|
||||
4. Server hanging - Load is not well distributed
|
||||
|
||||
5. Search engine ingrest hangs the normal app operation (IO intensive operation block others)
|
||||
|
||||
6. SAML login failure for 50% of users (If it’s OK, it’s always OK, IDM server ok, time sync issue of IDP server)
|
||||
|
||||
7. Attack - IDM - outage - (Find the source IP via agent, then apply WAF country block)
|
||||
|
||||
8. SMAX-Octane integration Timeout (NAT gateway: 350s)
|
||||
|
||||
9. Out of memory ( Some server pods hanging there, not responding, )
|
||||
|
||||
10. Redis high scan issue
|
||||
|
||||
11. XIE(Integration Engine) Infinite loop (High soft interrupt, short term: rolling restart; long term, fix)
|
||||
|
||||
12. Xmpp (chat) need to be restarted in a specific schedule otherwise it will hanging there
|
||||
|
||||
13. Helm regression ( losing dependency during upgrade)
|
||||
|
||||
14. ALB ingress controller outage (forgot to update add-on)
|
||||
|
||||
15. EFS running out of credits( the entire system gradually slowing down )
|
||||
|
||||
16. Instantly overload ( auto healing )
|
||||
|
||||
2. Infra & cloud issue
|
||||
|
||||
1. Cross zone HA (Frankfurt one zone failure)
|
||||
|
||||
2. DNS issue (Internal DNS resolving issue → CP FW → Network Account → South America DNS)
|
||||
@@ -1,327 +0,0 @@
|
||||
---
|
||||
title: SOC 2 in Practice — Interview-Ready Experience Notes
|
||||
source: "[[SOC2 Audit in Practice]]"
|
||||
tags: [soc2, compliance, audit, interview, cloud-devops]
|
||||
---
|
||||
|
||||
# SOC 2 in Practice — Interview-Ready Experience Notes
|
||||
|
||||
Interview-oriented rewrite of my hands-on SOC 2 notes: what I owned, how each control actually ran, what evidence it produced, and how to tell the story in English in 30 seconds, 2 minutes, or 20 minutes depending on the room.
|
||||
|
||||
## 0. Contents
|
||||
|
||||
1. Fast facts (the fact sheet I keep in my head)
|
||||
2. 30-second pitch (say it verbatim)
|
||||
3. 2-minute STAR narrative
|
||||
4. Control deep dives by Trust Services Category
|
||||
5. The evidence model (why the controls survive an audit)
|
||||
6. Numbers to quote
|
||||
7. Likely interview questions and short answers
|
||||
8. Gaps and what I would improve (say this — it signals seniority)
|
||||
9. Vocabulary bank (English phrasing for this domain)
|
||||
10. Transcription corrections vs. the raw notes
|
||||
|
||||
---
|
||||
|
||||
## 1. Fast facts
|
||||
|
||||
- **Role:** Cloud Service Delivery Manager, and designated **account owner** for the AWS account estate.
|
||||
- **Scope:** approximately **23 AWS accounts**, shared across Operations, DevOps domain teams (Network, CCOE) and Architecture.
|
||||
- **Access model:** AWS access is federated with the corporate **single sign-on** (email-based authentication), so the authoritative list of who can reach a given account lives with the IT organisation at a higher level — I request it, I do not own it.
|
||||
- **Platform:** multi-tenant B2B SaaS application running on **AWS EKS** (Kubernetes), **RDS** and **EFS**.
|
||||
- **Regulatory landscape:** SOC 2 across all five Trust Services Categories, plus **FedRAMP** for the US instance and **GDPR** for the EU environment.
|
||||
- **Recurring cadence:** quarterly access reviews · monthly hardened-AMI/patch cycle · backups every 6 hours · 7-day retention · two DR tests per year.
|
||||
- **Committed service levels:** **RPO 6 hours / RTO 24 hours**, published in the SaaS Service Description.
|
||||
|
||||
---
|
||||
|
||||
## 2. 30-second pitch (say it verbatim)
|
||||
|
||||
> "I'm the Cloud Service Delivery Manager for a multi-tenant SaaS platform on AWS, and I'm the account owner for about 23 AWS accounts. A large part of my job is being the **control owner** for our SOC 2 programme rather than just a participant in it: I run the quarterly access recertification across every account, I own the segregation rule that keeps Operations out of the product source code, I run vulnerability management at the OS layer through Qualys and our hardened AMI baseline, and I own cross-region backup and twice-yearly disaster recovery testing against a published 6-hour RPO and 24-hour RTO. What I really built is the evidence trail — so that when the auditor samples any quarter, the source list, the decision, the action and the sign-off are already there."
|
||||
|
||||
Why this works: it names the domain, the scale, the controls owned, the numbers, and the differentiator (evidence-by-default, not audit-by-fire-drill).
|
||||
|
||||
---
|
||||
|
||||
## 3. 2-minute STAR narrative
|
||||
|
||||
**Situation.** A B2B SaaS business running on AWS with customers who require a SOC 2 Type II report before they will sign. Five Trust Services Categories in scope. The cloud estate spans ~23 AWS accounts with several engineering groups — Operations, DevOps (Network team, Cloud Center of Excellence), Architecture — plus product engineering teams that own the application. Compliance had to be produced by people who were already running the platform, across regions with different regulatory regimes (FedRAMP in the US, GDPR in the EU).
|
||||
|
||||
**Task.** I owned the recurring operational controls and the evidence they generate: access control and recertification, source-code segregation, vulnerability management, backup and disaster recovery, the non-production data rule, and the customer-exit data-deletion process — each with a named approver and a retained record.
|
||||
|
||||
**Action.** I turned each one into a documented, repeatable loop rather than an annual scramble: role-based IAM personas with least privilege (Operations broadest, DevOps domain teams scoped to their domain, read-only for Architecture); a quarterly six-step access review driven off the SSO access list with escalation to peer managers and immediate revocation for leavers; a permissions check against GitLab proving Operations cannot reach product source code; a severity-filtered vulnerability pipeline that splits OS-level findings (fixed via the monthly CCOE hardened AMI) from library-level findings (routed to product teams, re-scanned and diffed against the previous review); AWS-native backup every 6 hours with cross-region replication and two tiers of restore strategy; two DR tests a year — one integrity-only, one full production failover; and a staging environment that is provably free of customer data and PII.
|
||||
|
||||
**Result.** The controls run on cadence, each producing a raw list → decision record → action record → management sign-off chain that the auditor can sample. The audit becomes a reporting exercise over evidence that already exists, instead of a project that consumes a quarter of engineering time. *(Fill in your own closing metrics here — e.g. number of findings closed per cycle, audit outcome, deals unblocked. The raw notes do not record them; do not invent them in an interview.)*
|
||||
|
||||
---
|
||||
|
||||
## 4. Control deep dives by Trust Services Category
|
||||
|
||||
### 4.1 Security
|
||||
|
||||
#### A. AWS account access review — quarterly access recertification
|
||||
|
||||
**Control objective.** Only authorised, currently-employed, appropriately-scoped people can reach production AWS accounts, and that is revalidated on a fixed cadence.
|
||||
|
||||
**Design — three IAM persona tiers, least privilege by default:**
|
||||
|
||||
- **Operations** — broadest privilege. They operate and upgrade resources on AWS and make network changes.
|
||||
- **DevOps domain teams** (Network team, CCOE) — narrower, scoped to their area. CCOE owns the OS-level AMI update pipeline.
|
||||
- **Architecture** — read-only. They need to pull information out of the environment but must never modify anything in the accounts.
|
||||
|
||||
**Execution — the repeatable six-step loop:**
|
||||
|
||||
1. Every three months, obtain the **account access list for the trailing three months** from the primary account owner / IT organisation (access is SSO-backed, so the authoritative user list sits at that higher IT level).
|
||||
2. **Classify** every identity in the list by group: Operations, DevOps, Architecture.
|
||||
3. **Re-validate the permission set** each identity actually holds against the IAM role baseline for that group.
|
||||
4. **Verify employment status** for my own Operations team; for identities I neither recognise nor own, escalate to the relevant DevOps or Architecture group manager for confirmation.
|
||||
5. **Revoke immediately** — anyone who has left the company, or who no longer needs the access, is cut off at once. This is the remediation step: access revocation.
|
||||
6. **Produce a summary report** and obtain sign-off from my Director/Manager and senior management, confirming the review was performed.
|
||||
|
||||
**Evidence retained:** the raw access lists, the revocation record, and the confirmation emails exchanged with peer managers.
|
||||
|
||||
**Interview line:** *"The control isn't the review — the control is the paper trail. Pulling the list and reclassifying it takes a day; what takes discipline is capturing the source list, the exception decision and the sign-off so that a sample from any quarter tells the same story to the auditor."*
|
||||
|
||||
#### B. Segregation — Operations must not touch product source code
|
||||
|
||||
**Control objective.** Operations admins must not be able to read or commit to the product source code repositories, so they cannot alter product logic or security-relevant behaviour.
|
||||
|
||||
**Check:**
|
||||
- The Product Team owner/manager is the **repository owner** and holds the authoritative list of identities permitted on those repos. I request that list.
|
||||
- I then verify our **Operations identities against the GitLab permission set** — the expected result is that no Operations identity appears in the authorised list.
|
||||
|
||||
**Exception handling.** If an Operations engineer is found with access, the permission is revoked. The raw list, the revocation record and the summary report are retained and signed off by senior management.
|
||||
|
||||
#### C. Vulnerability management (Qualys + Prisma/Defender)
|
||||
|
||||
**Sources.** Qualys scans the **cloud application runtime OS** (Linux) and surfaces risks and vulnerabilities at the OS layer; Prisma/Defender provides additional scanning. As account owner I receive the periodic policy reports — they are very large and cover many facets of the OS.
|
||||
|
||||
**The remediation pipeline:**
|
||||
|
||||
1. **Filter by severity**, then separate **OS-level findings** from **application/library-level findings**.
|
||||
2. **OS-level** — remediate through the **CCOE hardened standard Linux AMI**. We started on AWS-native Linux hardening and later adopted the CCOE standard AMI to meet specific customer and security requirements. CCOE publishes roughly **monthly**, with the latest patches, and tests before release; we consume the tested image and schedule the upgrade.
|
||||
3. **Library-level** — findings that will not be fixed by an OS upgrade route to the **product/development team** as a request to upgrade the affected library version.
|
||||
4. **Define the plan** — which findings are fixed in the next cycle, prioritised by severity.
|
||||
5. **Re-scan and diff** — after the AMI upgrade and the product patches land, compare the new policy report against the previous review: mark what is fixed, highlight what is not, and re-plan the remainder by severity.
|
||||
6. **Document the whole chain** — findings, filtering, review, fix, testing, sign-off — for the audit.
|
||||
|
||||
**Worked example to tell in an interview:** Qualys surfaced findings caused by an **out-of-date Kubernetes version**; the remediation was planned as an **EKS version upgrade** in the next release cycle.
|
||||
|
||||
#### D. Risk assessment — *not yet written up in the source notes*
|
||||
|
||||
The raw notes leave this empty. Prepare to speak to whatever you actually have: a dated annual risk assessment, a risk register with owner and treatment decision, and evidence that the register fed the control set. If that does not exist, treat it as a gap (§8) rather than claiming it.
|
||||
|
||||
---
|
||||
|
||||
### 4.2 Availability
|
||||
|
||||
#### A. Backup strategy — roughly 90% AWS cloud-native
|
||||
|
||||
- **What is not backed up:** container images. They are build artifacts published by R&D to GitHub on every release, so backing them up adds no recovery value.
|
||||
- **What is backed up:** the **Kubernetes configuration** — the YAML manifests that describe the web application, its pod layout and worker-node distribution. Restoring those lets us rebuild the containerised application quickly and consistently.
|
||||
- **Data tier:** **RDS and EFS** are backed up through **AWS Backup** with defined backup plans; the cadence is **every 6 hours**.
|
||||
- **Cross-region:** additional scripts replicate RDS backups into a second remote region — **Oregon → North Virginia**, and **Frankfurt → Ireland**.
|
||||
- **Retention:** **7 days**.
|
||||
- **Driver:** the disaster-recovery commitments published in the **SaaS Service Description**; the objective is to keep data exposure within roughly **6 hours**.
|
||||
|
||||
#### B. Replication and restore strategy — two tiers, cost-driven
|
||||
|
||||
- **Tier 1 — cold / remote backup (default).** Cross-region snapshots. Cheap, slower to activate. Chosen for cost reasons.
|
||||
- **Tier 2 — warm standby (faster RTO).** Periodically **restore** the RDS and EFS snapshots directly into a live database in the secondary (North Virginia) environment, rather than leaving them as snapshots. This does **not** run a full runtime instance, but it materially shortens the time needed to bring the cloud service back in the second region. It costs more.
|
||||
- **Selection:** the tier is chosen per customer according to their requirements — a deliberate cost-versus-RTO trade-off.
|
||||
|
||||
#### C. Disaster recovery testing — twice a year (DR integrity testing)
|
||||
|
||||
**Test 1 — lightweight, no impact on production.** A data-integrity validation. We take the backup data — device configuration files plus the RDS data backup — and restore it into a **backup environment in the same region as production but isolated from it**. We measure how long the restore takes. There is no cutover and no customer impact.
|
||||
|
||||
**Test 2 — full DR / production failover.** This one exercises the **remote-region backups**. It typically runs as a weekend cutover:
|
||||
|
||||
1. Stop the production environment.
|
||||
2. Restore the remote-region instance using the latest backup delta.
|
||||
3. Cut production traffic over to the remote-region environment and resume data flow so customers are served.
|
||||
4. Roughly **one week later**, replicate the data back to the original production environment.
|
||||
5. Run a **service load/pressure test**, then fail back to the original environment.
|
||||
|
||||
This test is materially more demanding — it is a real failover and failback, run against a real customer commitment.
|
||||
|
||||
**Acceptance criteria.** The RPO and RTO committed to customers in the SaaS Service Description: **RPO 6 hours, RTO 24 hours**. The strategy exists to satisfy that commitment, not an internal preference.
|
||||
|
||||
#### D. Processing capacity and multi-location strategy — *not yet written up in the source notes*
|
||||
|
||||
---
|
||||
|
||||
### 4.3 Confidentiality
|
||||
|
||||
#### A. Confidential information in non-production environments
|
||||
|
||||
- We maintain a **dedicated staging environment** used for application upgrades, patch and hotfix deployment, cloud-infrastructure changes and automated testing.
|
||||
- **The assertion we have to prove:** staging contains no production or customer confidential data, and no **PII**.
|
||||
- **How it is evidenced:** provide the staging environment's **tenant names**, plus a live walkthrough/demonstration for the auditor, showing that only synthetic test data is in use and that real customer data never reaches staging.
|
||||
|
||||
#### B. Data deletion and removal practices — the customer exit process
|
||||
|
||||
**The contractual promise (Service Decommissioning, per the SaaS Service Description):** on expiry or termination of the order term, the provider may disable all customer access to the SaaS, and the customer shall promptly return or destroy any provider materials. The provider makes available any SaaS data in its possession in the format generally provided, within the **Termination Data Retrieval Period SLO**. After that period the provider has no obligation to maintain or provide the data, which is deleted in the ordinary course.
|
||||
|
||||
**Communication and coordination**
|
||||
- Notify all relevant internal teams — support, billing, account management, cloud service.
|
||||
- Designate a **single point of contact** to manage the transition and answer queries promptly; in practice this is usually the CSM.
|
||||
|
||||
**Data management**
|
||||
- Ensure the customer can export their data easily, and assist where needed.
|
||||
- Plan **secure deletion** of customer data after the agreed period, in line with data-protection regulation and the data-retention policy.
|
||||
- Give a clear timeline for how long data remains accessible after service termination.
|
||||
|
||||
**Security and compliance**
|
||||
- **Revoke access** — disable all user accounts associated with the customer.
|
||||
- **Compliance check** — confirm the termination process satisfies relevant legal and regulatory requirements such as **GDPR** or **CCPA**.
|
||||
|
||||
**Detailed steps**
|
||||
|
||||
1. **Customer raises a service request** — the exit project is triggered by a request in **PCS**, and all related communication stays in PCS until every task is complete and the user account is closed there. The request must state: whether the customer wants existing **ESM/SMAX transaction data exported**; the date they want all tenant data completely emptied; the exact date by which the provider commits all relevant data — **including backups** — is cleaned out; and the date the PCS support channel is to be closed.
|
||||
2. **Assist with data export.**
|
||||
- **SMAX** — customers can use **OData export**; Cloud Ops helps by running the existing out-of-the-box OData export script to export SMAX transaction data per tenant.
|
||||
- **CMS / HCMX / OO** — not supported at the time.
|
||||
- **PCS data** — not supported at the time.
|
||||
3. **Plan data deletion.**
|
||||
- Notify the customer when the tenant will be terminated and all data deleted; Cloud Ops drives this notification from PCS.
|
||||
- **Scope of deletion:** tenant data, user data, account data (in BO) and inactive PCS entitlements.
|
||||
- **Retention:** farm-level data retention is **7 days** — after 7 days, customer data is permanently removed from the cloud environment.
|
||||
|
||||
---
|
||||
|
||||
### 4.4 Processing Integrity
|
||||
|
||||
*(from the linked process notes: `[[Major Incident Management Process]]`, `[[Cloud Change Management Process]]`)*
|
||||
|
||||
#### A. Cloud change management
|
||||
|
||||
- **Definition of a change:** anything — hardware, software, system components, services or processes — deliberately introduced into production that may affect an SLA or the functioning of the environment or one of its components. Drivers include user requests, vendor-recommended changes, regulatory changes, upgrades, failures, infrastructure modifications, unforeseen events and periodic changes.
|
||||
- **Planned change:** scheduled **at least 2 weeks in advance** when customer action is required, or **at least 4 days in advance** otherwise.
|
||||
- **Emergency change:** a critical change to prevent loss of service functionality or availability. Requires approval from the **Cloud Delivery Manager, TO Manager or CS Manager**, and is scheduled at least 1 day ahead unless it is critical to resolve a major incident immediately.
|
||||
- **Change record:** every change is recorded in the **Essentials** system.
|
||||
- **CAB review:**
|
||||
- *No CAB required* — routine, frequently executed, pre-approved by an executive, low likelihood of disruption (e.g. monthly patch upgrade, routine EKS upgrade).
|
||||
- *CAB required* — non-exempt changes, typically maintenance-window changes involving more than one executor (e.g. major product version upgrade, AWS infrastructure change, landing-zone migration).
|
||||
- **Customer notification:** a centralised notification system and Service Health portal publish current availability, upcoming planned maintenance, outage reports and historical SLO data.
|
||||
|
||||
#### B. Major incident management
|
||||
|
||||
- **Detection:** automated monitoring for anomalies and performance issues, plus user reports through designated channels.
|
||||
- **Definition of a major incident:** service outage (users cannot access the application at all), performance degradation (evident in monitoring or user feedback), or major functionality impact (tracked through APM).
|
||||
- **Response:** rapid triage by a cross-functional team spanning development, operations and support; impact analysis across users, systems and business operations; structured communication and tracking until closure.
|
||||
|
||||
---
|
||||
|
||||
### 4.5 Privacy
|
||||
|
||||
#### A. Data controller and region-restricted personnel access
|
||||
|
||||
Two environments in the estate carry hard residency and personnel constraints, and the control is enforced on **who may touch the data**, not just on network boundaries:
|
||||
|
||||
- **United States instance (FedRAMP).** Only **US citizens physically located in the US** may touch any data in that environment. Engineers in China, India and Europe are deliberately excluded. Only US-based Operations engineers are authorised to operate, access and maintain that environment.
|
||||
- **European environment (GDPR).** Only **European engineers** may access those specific environments; engineers from other regions have no access.
|
||||
|
||||
**How this shows up in the control set:** the regional restriction is designed into the access model, and the quarterly access review is where it is verified — the "who" has to keep matching the "where".
|
||||
|
||||
**Reference notes:** `[[GDPR]]`, `[[FedRAMP Basics Understanding Federal Cloud Security Standards]]`.
|
||||
|
||||
---
|
||||
|
||||
## 5. The evidence model — why these controls survive an audit
|
||||
|
||||
Every recurring control I ran produces the same **four-artifact chain**. This is the part worth explaining in an interview, because it is what distinguishes an operator from a control owner:
|
||||
|
||||
1. **Source of truth (raw input)** — the three-month account access list; the GitLab authorised-user list; the Qualys/Defender policy report.
|
||||
2. **Decision record** — my classification and triage with the rationale: who stays, who is revoked, which findings are deferred and why.
|
||||
3. **Action record** — the revocation; the permission removal; the AMI upgrade; the remediation plan and the re-scan diff.
|
||||
4. **Sign-off** — the summary report signed by Director/Manager/senior management, plus the confirmation emails from peer managers for identities I do not own.
|
||||
|
||||
The auditor does not need to reconstruct the quarter; the quarter is already on file.
|
||||
|
||||
---
|
||||
|
||||
## 6. Numbers to quote
|
||||
|
||||
- **~23 AWS accounts** owned and reviewed as account owner.
|
||||
- **Quarterly** access recertification; each review covers a **3-month** access window.
|
||||
- **3 IAM persona tiers** — Operations, DevOps domain teams, read-only Architect.
|
||||
- **Monthly** hardened-AMI release cadence from CCOE; **2 scanning platforms** (Qualys, Prisma/Defender).
|
||||
- Backups **every 6 hours** for RDS/EFS; **7-day** retention; **2 cross-region pairs** (Oregon→N. Virginia, Frankfurt→Ireland).
|
||||
- **2 DR tests per year** — one integrity-only, one full failover with a weekend cutover and a ~1-week failback.
|
||||
- Committed **RPO 6 hours / RTO 24 hours**.
|
||||
- Change management: planned changes **≥2 weeks** (when customer action is needed) or **≥4 days**; emergency changes **≥1 day** with named approvers.
|
||||
- Customer exit: farm-level data retention **7 days**, then permanent removal.
|
||||
|
||||
---
|
||||
|
||||
## 7. Likely interview questions and short answers
|
||||
|
||||
**"How did you keep quarterly access reviews across 23 accounts from becoming a rubber stamp?"**
|
||||
Two things. First, I never reviewed only my own team — unknown identities were escalated to the owning manager for confirmation, which forced a real answer rather than a default approval. Second, every review had to end with a revocation decision or an explicit "no change", signed off by senior management. The output is a report, not a checkbox.
|
||||
|
||||
**"What is the most common finding in an access review?"**
|
||||
Stale access — people who have left, or moved roles, but whose identity still appears on the list. Because AWS access is federated to corporate SSO, the identity can still exist in the list even when the person is gone. That is exactly why the review cross-checks the list against current employment rather than trusting the list.
|
||||
|
||||
**"How did you implement least privilege?"**
|
||||
Role-based IAM personas rather than per-person grants: Operations broadest because they operate and upgrade the platform and make network changes; DevOps domain teams scoped to their area, like CCOE with the OS/AMI pipeline; read-only for architects who need visibility but must not modify anything.
|
||||
|
||||
**"How did you prove your staging environment held no production data?"**
|
||||
We gave the auditor the staging tenant names and walked them through the environment live, showing that only test data was in use — no customer data and no PII.
|
||||
|
||||
**"What happens when a vulnerability can't be fixed in the current cycle?"**
|
||||
It doesn't disappear; it gets re-planned by severity and stays visible. After remediation we re-run the scan and diff it against the previous review, so anything still open is highlighted and carries into the next plan. A finding caused by an out-of-date Kubernetes version, for example, became a planned EKS upgrade in the following cycle.
|
||||
|
||||
**"How did you trade off cost against recovery time?"**
|
||||
Two restore tiers: a cold cross-region backup as the default because it is cheap, and a warm option where snapshots are periodically restored into a live database in the secondary region — no full runtime, but a much shorter restore. Which one a customer gets is decided by their requirements and their RTO commitment, not by a blanket policy.
|
||||
|
||||
**"How do you test a DR plan without hurting customers?"**
|
||||
Two different tests. A lightweight integrity test restores into an isolated environment in the same region as production and only measures the restore — zero customer impact. The full test is a real weekend cutover to the remote region, running for about a week before a load-tested failback. Both are measured against the published 6-hour RPO and 24-hour RTO.
|
||||
|
||||
**"What is a Type II report and what does it mean for your day job?"**
|
||||
Type II is an opinion over an observation period, not a point-in-time snapshot — the auditor samples evidence from the whole window. That changes behaviour: evidence has to be produced contemporaneously at the moment the control runs, because it cannot be reconstructed later.
|
||||
|
||||
**"How did you handle FedRAMP and GDPR residency requirements?"**
|
||||
As a personnel control, not just a technical boundary. For the US instance, only US citizens physically in the US could touch the data — engineers in China, India and Europe were excluded. For the EU environments, only European engineers had access. The quarterly review is where the "who" gets checked against the "where".
|
||||
|
||||
**"What would you do differently?"** — go straight to §8. Answering this with specifics is usually the strongest part of the interview.
|
||||
|
||||
---
|
||||
|
||||
## 8. Gaps and what I would improve
|
||||
|
||||
Saying these out loud is a credibility play — it shows you know where the control set is thin:
|
||||
|
||||
- **Risk assessment is the weakest link.** A formal, dated annual risk assessment with a risk register — owner, treatment decision, and a traceable link from each risk to a control — would tighten the CC3 area far more than another access review.
|
||||
- **Access recertification is still manual and email-driven.** Tooling the campaign (automated attestation workflow with the owning managers) would remove the peer-manager round trips and shrink the review cycle.
|
||||
- **Processing capacity and multi-location strategy are documented far less rigorously than backup and DR**, even though availability depends on all of them.
|
||||
- **Confidential information classification** is implied by the staging-data rule but never written as a standalone classification policy — an easy win.
|
||||
- **Customer-exit data export was unsupported for CMS/HCMX/OO and PCS at the time.** That is a real gap in the exit promise and deserves a documented compensating process rather than an implicit "we'll handle it".
|
||||
|
||||
---
|
||||
|
||||
## 9. Vocabulary bank — English phrasing for this domain
|
||||
|
||||
**Access control:** access recertification · quarterly access review · access revocation · least privilege · role-based IAM personas · read-only access · privileged access · federated identity / single sign-on · segregation of duties · leaver access · stale entitlement · escalation to the owning manager.
|
||||
|
||||
**Source code control:** segregation from the source code repositories · no commit rights · authorisation list · repository owner · permission removal.
|
||||
|
||||
**Vulnerability management:** vulnerability scan · severity-based triage · hardened AMI baseline · patch cadence · remediation plan · re-scan and diff against the previous review · risk acceptance · outstanding finding · compensating control.
|
||||
|
||||
**Availability:** RPO (recovery point objective) · RTO (recovery time objective) · cold backup · warm standby · cross-region replication · failover and failback · DR integrity test · production cutover · load/pressure test · retention window.
|
||||
|
||||
**Confidentiality and privacy:** data classification · non-production environment · synthetic test data · PII · data retention and disposal · decommissioning · customer exit · data export (OData) · entitlement · data controller · data residency.
|
||||
|
||||
**Process integrity:** change record · CAB (change advisory board) · emergency change · planned maintenance window · service health portal · incident triage · impact analysis · cross-functional response.
|
||||
|
||||
**Audit language:** Trust Services Criteria (TSC) · Type I vs Type II · observation period · control owner · evidence retention · management sign-off · sampling · subservice organisation (AWS) · user entity · service auditor.
|
||||
|
||||
---
|
||||
|
||||
## 10. Transcription corrections relative to the raw notes
|
||||
|
||||
- "calibrate access write" / "cataloger" → **revoke access rights** (the remediation action in the access review and the GitLab check).
|
||||
- "Processor Integrity" → **Processing Integrity** (the fifth Trust Services Category).
|
||||
- "Recovery product objective" → **Recovery Point Objective (RPO)**; RPO **6 hours**, RTO **24 hours**.
|
||||
- "抵押的标准" → **the standard we committed to** in the SaaS Service Description, i.e. the published RPO/RTO.
|
||||
- "Prisma Defender" → written as **Prisma / Defender** here; confirm the exact product name you cite in an interview.
|
||||
- Indicative mapping to SOC 2 criteria (CC6.x access, CC7.x operations/vulnerability, CC8.1 change management, A1.2/A1.3 backup and recovery testing, C1.1/C1.2 confidentiality, P4.x privacy) is my own reading — verify against the actual audit report before quoting criteria numbers to an interviewer.
|
||||
@@ -1,181 +0,0 @@
|
||||
## Security
|
||||
### Access Control
|
||||
|
||||
**AWS Account Access Review**
|
||||
现在来说一下 SOC2 audit 在我们实际的工作当中是怎么样来执行的。先开始第一个:security 里
|
||||
面的 access control。讲的例子主要是我们要怎么样来管理 AWS account 的 access control。因为我作为整个 cloud service delivery 的 manager,所以我管理了差不多有 23 个 AWS account。我是这些 AWS account 的 account owner。所以我会每个季度进行一次 quarterly review 来检查这些 account 里面访问的用户的一些权限,包括检查一些已经离开公司的人员是否还有访问权限。先我们在 AWS account 里面事先会设计不同的 IAM 角色,然后给这些角色赋予不同程度的权限。比如 operation team 的 member 的权限比较大,因为他要在 AWS 上面去操作很多资源来进行一些升级,包括一些网络的修改等等。
|
||||
|
||||
其次我们还有一些 DevOps 团队。这些团队其实有些是 Network Team,有一些是 CCOE Team,他负责的是一些操作系统的 AMI的更新。所以他们的权限相对来说比 operations 要稍微小一点,主要集中在他们所对应的相关领域里面。我们还设计了一些 read-only 的权限,这个主要是给到一些 architect。他们可能要在我们的环境里面去抓取一些相关信息,但是他们不允许在我们的 AWS account 里面做一些修改,所以我们会设定一些 read-only 的政策。当然,这个里面权限设定会比较复杂,还有其他各种各样的权限都是根据不同的角色来设定的。你的操作方法呢是:我们每三个月会从整个 AWS account 的 primary owner 那边去拿到一个 account 的 access list。因为我们的 AWS account 访问是和 company 的 single sign-on 绑定的,所以我们是通过 email authentication 来登录 AWS account。所以这些信息会在更高级别的这个 IT 团队里面有,访问具体某个 AWS account 的 user 的一个 list。我拿到了过去3个月的访问记录以后,我会逐个地对这些人员进行分类:
|
||||
- operation 团队的人
|
||||
- devops 的
|
||||
- 其他的 architect group 的
|
||||
分好类以后我会检查他们所对应的权限。如果出现一些人员我并不认识或者是我并不组织的,我会发给相应的 devops 的 manager 以及 architect 的 group 的 manager,请他们帮忙来进行确认。 包括我自己的 operation 团队,我也会检查所有人员是否是当前在职人员。如果有一些人员已经离职了,但是他的访问权限还出现在这个 access list 里面,那我们就必须采取行动立即停止这些访问权限。 这个动作称之为“calibrate access write”。
|
||||
当我完成了每三个月一次的 access review 以后我会做一个 summary report,然后把这个 report 发给我的 director、manager,甚至高级别的 high-level manager,来进行 sign-off,来确认我们完成了这样的一个动作。
|
||||
|
||||
在这个过程当中所有最初的原始人员的名单、我进行 calibrate 的 report,以及我去跟其他的 manager 进行确认的沟通邮件,都会被记录下来作为 evidence,来应对后续的 SOC2 audit。
|
||||
|
||||
**Restrict to access product source code repository**
|
||||
还有一个项目也是会定期来做检查的。这个的目的也是从安全的角度考虑。在我们现在的这个组织里面我们是不允许 operation 的管理员有权限去访问整个产品的 source code repository 的。
|
||||
|
||||
operation 它应该不应该去放到 product 的 source code 不能做任何的 code commit 来修改产品里面的任何一些逻辑,包括一些安全方面的东西。它是不能涉及到这个 source code 的。所以在这个地方我们也有额外的 check:我们会检查我们的 GitLab 的权限,确保我们的 operation engineer 没有任何权限去访问我们的 source code 的这些 repository。
|
||||
地方呢,我们会找到相应的 product team 的 owner、manager,作为这些 code repository 的 owner。我们会要求他能够提供这么一个名单,然后我们会来检查我们的 operating ID 是否在这个名单里面。正常的应该不在那个名单里面。
|
||||
|
||||
如果我们发现有些工程师是有权限可以访问的,我们会做一些 cataloger 去把这些权限可能拿掉。同样的整个过程当中所有的记录,原始的记录包括 cataloger 的记录,以及后续的一些 summary 的 report,都会保存下来,也会给high level manager sign off。
|
||||
|
||||
**Risk Assessment**
|
||||
|
||||
|
||||
**Vulnerlubility Management**
|
||||
这里介绍的是一个关于 vulnerability 的管理。我的介绍的 case 是我们的商业应用在云上的环境下面,我们定期会有 security 的 scan。其中就有 Qualys 的 scan 和 Prisma Defender 的一些 scan。
|
||||
|
||||
Qualys 的 scan 主要是针对我们 Cloud Application Runtime 环境上面的操作系统,比如 Linux,包括会有哪些 risk 和一些漏洞,这些都会被 Qualys 进行扫描出来。我们是这样来进行管理的。
|
||||
|
||||
作为整个 account 的 owner,我会定期收到系统发出的一些 policy report。这些 report 的内容会非常大,涉及整个 OS 里面的很多方方面面的一些 vulnerability。我们会对这些问题进行 filtering。首先我们会根据里面一些问题的 severity 来进行 filtering,来分析哪些是 OS 级别的。我们会结合我们另外一个 branch,那是我们整个 cloud central excellence team 提供的持续的 OS 级别的 Linux AMI hardening。
|
||||
|
||||
最早我们一开始使用的是 AWS 原生的 Linux hardening。后期因为各种客户的需要,包括我们有一些 specific 的 security requirement,我们就开始 adopt CCOE 提供的标准 Linux AMI。它发布的周期差不多是每个月发布一个新的版本,包含最新的 patch。我们会在收到他们测试过的 AMI 之后来进行 project plan。 除了通过AMI的升级能够修复一些OS级别的问题之外,我们还会定义我们还会 filtering 一些是不是通过OS升级来发现的那些问题。
|
||||
|
||||
针对这些问题我们就需要去做一些额外的动作。有可能我们会要去通知 product team 升级某些 library 的版本。如果一些旧的版本包含了一些 vulnerability 的问题的话,我们需要去联系研发团队、开发团队去更新这样的 library。
|
||||
|
||||
最终我们会定义一些 plan:哪些问题我们会在下一个版本进行修复。当我们升级完 AMI,包括开发团队也提供了相应的一些产品补丁之后,我还会针对这个最新的 policy 来和之前的 review 进行一次比较。看看哪些问题可以标注为已经修复了,哪些问题其实还是没有修复呢。如果有些问题并没有完全修复,我们会 highlight 出来,然后再进一步制定一些新的计划。计划还是根据整个扫描出来的问题的 severity 来进行下一步的新的计划。
|
||||
|
||||
比如之前Qualy扫描出来的有些问题是由于 Kubernetes 的版本太低造成的。我们就需要计划在下一个版本中对 AWS 上面的 EKS 的版本进行升级来解决这个问题。
|
||||
|
||||
像类似的这种问题我们都会进行检查,并把所有的发现、filtering、review,包括后续 fix、testing、sign off,这些所有的内容整个过程全部都记录下来,以便以后进行后续的 software audit。
|
||||
|
||||
## Availbility
|
||||
- **Backups**
|
||||
针对我们 SaaS application 的 backup,我们基本上是 90% 依赖于 AWS 原生的一些 cloud-native backup feature。
|
||||
|
||||
比如说我们的 application 首先是基于 AWS EKS 整个进行 Kubernetes 容器化部署的。所以在备份过程当中整个 cloud application 的备份,我们并不是要备份所有这些庞大的 images,因为这些 images 通过 R&D 团队每个版本发布到 GitHub 上面去的。
|
||||
|
||||
唯一我们需要备份的其实是 Kubernetes 的一些 configuration files。这样的话我们可以通过 configuration files(那些 yaml 文件)能够快速地把整个 web application 整个容器化,包括它的 pod 分布和 work node 的分布。我们可以很快地按照这个 configuration 还原出来。
|
||||
|
||||
所以我们只要备份一些 configuration files 就可以了。数据库RDS和数据存储EFS这方面我们是依赖 AWS backup。 这个 AWS backup,我们是会制定一系列的 backup plan,包括了整个备份的频率。我们每 6 个小时会对 RDS/EFS 的整个数据库和数据存储进行一次备份。这个备份不仅仅是备份到当前的 region;同时我们会通过一些额外的脚本来实现把 RDS 能够备份到一个remote region:
|
||||
AWS Oregon -> AWS North Viginia
|
||||
AWS Frankfult -> AWS Ireland
|
||||
|
||||
然后我们整个数据保留 7 天。这个 backup 是根据我们对整个 disaster recover 的一个 commitment 来做到的。RPO/RTO 我们在 SaaS 的一个 service description 里面是提到的,我们会把数据的影响控制在大概 6 个小时之内。
|
||||
- **Processing Capacity**
|
||||
- **Replication**
|
||||
|
||||
- **数据复制**:在多个系统/地域间同步数据副本,确保即使主系统故障,数据仍可被访问
|
||||
- **实时或准实时的备份机制**,避免单点故障导致数据丢失
|
||||
- **恢复时间目标(RTO)的支撑**:通过即时可用的副本,快速切换到备用系统
|
||||
我们每 6 个小时会对 RDS 的整个数据库进行一次备份。这个备份不仅仅是备份到当前的 region;同时我们会通过一些额外的脚本来实现把 RDS 能够备份到一个remote region:
|
||||
AWS Oregon -> AWS North Viginia
|
||||
AWS Frankfult -> AWS Ireland
|
||||
|
||||
到的这个方案基本上还是属于一个 cold backup、远备份的方案,因为是考虑到一些成本的原因。其实在这个基础上我们还有一套更快速能够缩短整个 RTO 时间的预案,我们称之为一个热备份方案。
|
||||
|
||||
那个方案基本上是会把 RDS 和 EFS 这些 snapshot 在另一个 Viginia 的 AWS 环境里面定时地直接恢复到这个数据库里面,而不是以 snapshot 的形式存在。这样做虽然我们并不是说真正去启一套 runtime 的 instance,但是它可以大大缩短我们在另外一个 region 恢复整个 cloud 的状态所需的时间。这样的话这个成本会相对来说有所提高。
|
||||
|
||||
在这个方面我们会根据客户一些不同的要求来进行取舍。
|
||||
|
||||
|
||||
- **Multi-location Strategies**
|
||||
|
||||
|
||||
|
||||
- **Business Continuity Planning and Testing**
|
||||
- **Disaster Recovery Planning and Testing**
|
||||
我们一年会进行两次DR testing,我们称之为 disaster recovery integrity testing。测试一次是轻量级的,不影响生产环境的,只是验证数据完整性的测试。那另一次就相对来说更复杂一点,它是一个完全DR的测试。
|
||||
是针对我们的 backup 的数据。这里面包含了 device 的 configuration file 和我们的 RDS 数据的备份。根据这些内容在我们一个 backup 环境里面恢复备份数据。这个 backup 环境跟生产环境在同一个 region 但是不影响正常环境。
|
||||
|
||||
我们去利用备份数据恢复一套账,看整个恢复需要多长时间,但是这个恢复不切换生产环境。
|
||||
|
||||
第二次测试我们会 test 这个 remote 的 backup。因为我刚才提到了,在备份过程中我们会把数据备份到另外一个 remote region。我们利用这些 remote region 的 backup data 去 restore 整个 instance,然后这个是要配合在生产环境。
|
||||
|
||||
我们基本上是要在周末做停机切换:会把生产环境停掉,然后拿最新的 backup 的 delta 来恢复 remote region 的一个 instance。恢复好了以后我们会把整个生产环境的流量切换到新的 remote region 的环境下面,再恢复整个数据的流量,让客户能够使用。基本上是在一个星期之后,我们把数据反向恢复到我们原来的旧生产环境当中,再做一次服务压力测试,给它切换回去。
|
||||
|
||||
这个对我们的要求会比较高。
|
||||
|
||||
抵押的标准要求是按照我们在 SAS service description 里面给用户的承诺:RPO 和 RTO。 Recovery product objective 是 6 小时,recovery time objective 是 24 小时。这个策略也是根据这个标准来执行的。
|
||||
|
||||
## Confidentialty
|
||||
- **Confidential Information Classification**
|
||||
- **Confidential Information in Non-Production Environments**
|
||||
针对这一块的 software audit,我们主要是用来证明我们在测试环境方面有专门的 staging,用于测试 application 的 upgrade,包括一些 patch、hotfix 的 deployment,以及我们在调整 cloud infrastructure 的结构上面的一些部署。包括一些自动化的测试都是在 staging 环境上面进行测试的。
|
||||
|
||||
这个的目的就是要测试我们在 staging 的环境上面没有用到大数据上面的 custom、confidential 的一些数据。这个主要的证明方法其实就是我们提供相应的 staging form 上面一些 tenant 的名称,包括做一些实际的演示,来证明我们仅仅只用到了一些测试的数据,并没有用到真实的客户的数据;也绝对不会包含 PII 个人信息的一些内容。
|
||||
|
||||
- **Data Deletion and Removal Practices**
|
||||
Customer Exit Process:
|
||||
## Introduction
|
||||
|
||||
When a SaaS customer decides to leave, it's crucial to handle the transition smoothly and professionally to ensure a positive experience, which can impact future business opportunities and the company’s reputation. This document describes the main processes and actions regarding customer exits.
|
||||
|
||||
## Service Description about Service Decomission
|
||||
|
||||
Service Decommissioning
|
||||
Upon expiration or termination of the SaaS Order Term, Micro Focus may disable all Customer access to
|
||||
SaaS, and Customer shall promptly return to Micro Focus (or at Micro Focus’s request destroy) any
|
||||
Micro Focus materials.
|
||||
Micro Focus will make available to Customer any SaaS Data in Micro Focus’ possession in the format
|
||||
generally provided by Micro Focus. The target timeframe is set forth below in Termination Data
|
||||
Retrieval Period SLO. After such time, Micro Focus shall have no obligation to maintain or provide any
|
||||
such data, which will be deleted in the ordinary course.
|
||||
|
||||
### **Communication and Coordination**
|
||||
|
||||
- **Notify Relevant Teams**: Inform all relevant internal teams (support, billing, account management, cloud service etc.) about the customer's decision.
|
||||
- **Designate a Point of Contact**: Assign a single point of contact to manage the transition and ensure all queries are addressed promptly. Usually it's the CSM.
|
||||
|
||||
### **Data Management**
|
||||
|
||||
- **Data Backup and Export**: Ensure the customer can export their data easily. Provide assistance if necessary.
|
||||
- **Data Deletion**: Plan for secure deletion of the customer’s data from your servers after a certain period, in compliance with data protection regulations and your data retention policy.
|
||||
- **Data Access Period**: Provide a clear timeline for how long their data will remain accessible after service termination.
|
||||
|
||||
### **Security and Compliance**
|
||||
|
||||
- **Revoke Access**: Ensure all user accounts associated with the customer are disabled and access to the system is revoked.
|
||||
- **Compliance Check**: Ensure that the termination process complies with all relevant legal and regulatory requirements, such as GDPR or CCPA.
|
||||
|
||||
## Detailed Steps for customer exit
|
||||
|
||||
### Customer to submit service request to trigger customer exit project
|
||||
|
||||
The customer needs to submit a service request in PCS to start the customer exit process. All related communication will be still handled in PCS until all the tasks are done and close the user account in PCS.
|
||||
|
||||
- In the request, the customer needs to clarify the following specific needs:
|
||||
Whether they wish to export existing ESM/SMAX transaction data?
|
||||
What's the expected date customer want all tenant data to be emptied out completely?
|
||||
What's the exact date Opentext to commit all relevant date (including backup data) will be cleaned out completely?
|
||||
What's the exact data to close PCS support channel?
|
||||
|
||||
### Assist with data export
|
||||
|
||||
- What’s the suggestion to customer to export data?
|
||||
- SMAX
|
||||
- SMAX Offer customer to use OData export to export data
|
||||
- Cloud Ops team can help to use existing OOTB OData export script to export SMAX transaction data per tenant
|
||||
- CMS/HCMX/OO
|
||||
- Not support by now
|
||||
- PCS data
|
||||
- No Support by now
|
||||
|
||||
### Plan data deletion
|
||||
|
||||
- Notification to customer to notify when we will terminate the tenant and delete all data
|
||||
- Cloud Ops will handle such notification from PCS.
|
||||
- Scope of data deletion
|
||||
- Tenant data/user data/ account data (in BO)
|
||||
- Inactive PCS entitlement
|
||||
- Data retention- farm level data retention is only 7 days. After 7 days customer data will permanently removed from Cloud environment
|
||||
## Processor Integrity
|
||||
[[Major Incident Management Process]]
|
||||
[[Cloud Change Management Process]]
|
||||
|
||||
## Privacy
|
||||
**Data Controller**
|
||||
在这个方面我们实际的操作过程当中是的确有这样的一个需求的。我们管理的所有环境当中有两个比较特殊的环境:
|
||||
|
||||
- 在美国的一个 instance,因为它是要符合 FedRAMP 整个标准规范的。因此他对所有能够接触到数据的工程师有这样的要求:只有在美国当地的美国公民才能去 touch 这个环境里面的所有数据。因此我们在做数据控制方面会特意规避这样的要求。我们只允许美国的 operation 工程师能够操作、访问以及维护这一套环境上面所有的数据。包括在其他国家的中国、印度以及在欧洲的工程师都没有权限去 touch 这个数据。
|
||||
- 在欧洲的环境,它是要符合欧盟的一些规范,包括一些 GDPR 的要求。同样地,只允许欧洲的工程师访问;其他 region 的工程师都不能访问那几个特定的环节。
|
||||
|
||||
[[GDPR]]
|
||||
[[FedRAMP Basics Understanding Federal Cloud Security Standards]]
|
||||
|
||||
|
||||
|
||||
|
||||
@@ -1,413 +0,0 @@
|
||||
# SOC 2 Compliance Training Course - Transcription Summary
|
||||
|
||||
## 【Chapter 1】SOC 2 Compliance Fundamentals
|
||||
|
||||
### Core Information
|
||||
- SOC 2 is a third-party audit framework validating that a company has security controls in place
|
||||
- Primarily used for B2B business partnerships to establish trust and security validation
|
||||
- Governed and standardized by AICPA (American Institute of Certified Public Accountants)
|
||||
- Has become the market default requirement for SaaS and cloud platform companies
|
||||
- Understanding SOC 2 eliminates compliance confusion and prevents wasted effort
|
||||
|
||||
### Key Terms and Definitions
|
||||
- **Service Organization**: The company/customer undergoing the SOC 2 audit
|
||||
- **Service Auditor**: The CPA who performs the SOC 2 examination according to AICPA standards
|
||||
- **User Entities**: The customers receiving and reviewing the SOC 2 report
|
||||
- **Subservice Organizations**: Organizations that perform controls on behalf of the service organization (e.g., AWS, GCP, Azure)
|
||||
- **Trust Services Categories (TSCs)**: The five pillars against which companies are evaluated
|
||||
- Security (安全性)
|
||||
- Availability (可用性)
|
||||
- Confidentiality (保密性)
|
||||
- Processing Integrity (处理完整性)
|
||||
- Privacy (隐私)
|
||||
|
||||
|
||||
## 【Chapter 2】Why Companies Pursue SOC 2
|
||||
|
||||
### Core Information
|
||||
- Customer demand is the primary driver (mandatory in contracts/RFPs)
|
||||
- SOC 2 is an "audit once, use many" solution
|
||||
- Compliance with regulatory requirements and industry-specific demands
|
||||
- Serves as a market differentiation and marketing tool
|
||||
- SOC 2's flexibility allows for customized control measures tailored to your organization
|
||||
|
||||
### Detailed Breakdown
|
||||
|
||||
#### Business Drivers
|
||||
- **Customer Demand**: Common requirement in contracts and RFPs; often blocks business deals
|
||||
- **Sales Department Focus**: SOC 2 reports are used to close deals and increase revenue
|
||||
- **Audit Once, Use Many**: One report responds to multiple customers' security inquiries instead of responding to different requests from each new customer
|
||||
- **Competitive Differentiation**: Demonstrates maturity and robust cybersecurity program to stand out from competitors
|
||||
|
||||
#### Regulatory and Compliance Factors
|
||||
- Industry-specific requirements (HIPAA, PCI DSS, etc.)
|
||||
- Regulatory mandates or breach fines
|
||||
- Investor due diligence requirements
|
||||
|
||||
#### Flexibility Advantages
|
||||
- Unlike prescriptive frameworks like PCI DSS and ISO 27001, SOC 2 is not prescriptive
|
||||
- Companies are not told exactly what to do, just what criteria or control objectives to meet
|
||||
- Allows for customized controls tailored to your specific company and application
|
||||
- Makes SOC 2 reports more robust and valuable to report readers
|
||||
- This flexibility is a vital competitive advantage
|
||||
|
||||
|
||||
## 【Chapter 3】SOC 2 Report Distribution and Use
|
||||
|
||||
### Core Information
|
||||
- SOC 2 reports are restricted use reports (50+ pages with sensitive control details)
|
||||
- Must be shared via Non-Disclosure Agreements (NDA) with specific parties
|
||||
- SOC 3 is the public version of SOC 2 Type 2 reports
|
||||
- AICPA logo can be publicly displayed on websites (no NDA required)
|
||||
- SOC 2 Plus allows integration of other compliance frameworks (HIPAA, ISO 27001, etc.)
|
||||
|
||||
### Detailed Breakdown
|
||||
|
||||
#### SOC 2 Report Sharing
|
||||
- **Restricted Use Report**: Contains sensitive control and audit details
|
||||
- **Specific Purpose Distribution**: Used for vendor due diligence or investor due diligence
|
||||
- **NDA Protection**: Requires non-disclosure agreements between companies
|
||||
- **Administrative Burden**: NDA process is tedious, but automated tools exist to simplify it
|
||||
|
||||
#### Public Promotion Options
|
||||
- **AICPA SOC 2 Logo**: Complete simple application (<10 minutes) to publicly display the official logo
|
||||
- **SOC 3 Reports**: Trimmed-down versions of SOC 2 Type 2 reports with sensitive details removed (Sections 3 and 4)
|
||||
- **SOC 3 can be publicly posted on websites for marketing purposes**
|
||||
|
||||
#### SOC 2 Plus Reports
|
||||
- Combines SOC 2 + HIPAA, ISO 27001, PCI DSS, or other frameworks
|
||||
- Satisfies multiple compliance framework requirements with one audit
|
||||
- Service auditor performs tests for all controls in additional frameworks
|
||||
- Increases effort but saves time and cost compared to separate audits
|
||||
|
||||
### Action Items
|
||||
- Establish NDA process or use automated tools for report sharing
|
||||
- Apply for AICPA logo to display on company website
|
||||
- Evaluate if SOC 2 Plus or SOC 3 reports are needed
|
||||
|
||||
|
||||
## 【Chapter 4】SOC 2 Report Types
|
||||
|
||||
### Core Information
|
||||
- SOC 2 Type 1: Point-in-time design assessment (fast, low cost)
|
||||
- SOC 2 Type 2: 12-month operational effectiveness assessment (target type, most commonly required)
|
||||
- 90% of contracts require Type 2 reports
|
||||
- Type 1 is a stepping stone to Type 2 (crawl, walk, run principle)
|
||||
- Sample testing is used in Type 2 to verify control consistency throughout the period
|
||||
|
||||
### Detailed Breakdown
|
||||
|
||||
#### SOC 2 Type 1
|
||||
- **Point-in-Time Assessment**: Controls evaluated as of a specific date
|
||||
- **Design and Suitability Focus**: Verifies controls are in place but NOT that they operate effectively
|
||||
- **Low Evidence Requirements**: Only one example needed (e.g., one employee completing training)
|
||||
- **No Test Steps or Results**: Section 4 only lists controls
|
||||
- **Significantly Lower Effort**: Much less work compared to Type 2
|
||||
- **Fast Path**: Quickest way to achieve SOC 2 compliance
|
||||
|
||||
#### SOC 2 Type 2
|
||||
- **Time Span**: Typically 12 months, but can range from 3-12 months
|
||||
- **Operational Effectiveness Focus**: Verifies controls operated effectively throughout the entire period
|
||||
- **Backward-Looking Assessment**: Auditor looks at controls over a prior period
|
||||
- **Sample Testing Methodology**: Auditor uses random selection to test a representative sample
|
||||
- Example: From 100 new hires during the audit period, auditor tests a random sample rather than all 100
|
||||
- **Detailed Test Steps and Results**: Section 4 includes test steps and their results
|
||||
- **Significantly Higher Effort**: Much more work for both auditor and company being audited
|
||||
- **Annual Renewal Expected**: Customers expect Type 2 reports to be renewed each year
|
||||
|
||||
### Action Items
|
||||
- First Audit: Conduct Type 1, then move to Type 2 (follow crawl, walk, run principle)
|
||||
- Third-Party Verification: Verify controls are in place before evaluating operational effectiveness over time
|
||||
|
||||
|
||||
## 【Chapter 5】SOC 2 Report Structure Analysis
|
||||
|
||||
### Core Information
|
||||
- SOC 2 reports contain 5 main sections + 1 optional section
|
||||
- Section 1: Auditor's Opinion (pass/fail determination point)
|
||||
- Section 3: System Description (most important detailed information)
|
||||
- Section 4: Controls and Test Results
|
||||
- Section 5: Optional management response and framework mapping
|
||||
|
||||
### Detailed Breakdown
|
||||
|
||||
#### Section 1: Independent Service Auditor's Report - The Opinion
|
||||
- **Opinion Types**:
|
||||
- Unqualified Opinion: Perfect pass, no issues found
|
||||
- Qualified Opinion: One or more issues identified
|
||||
- Adverse Opinion: Significant issues found (very rare)
|
||||
- Disclaimer of Opinion: Unable to audit
|
||||
- **Most Common**: Unqualified opinions are most common, but qualified opinions are not rare
|
||||
- **Exception Definition**: When auditor finds a control not operating effectively
|
||||
- Examples: Employee didn't complete security awareness training; employee with sensitive data access lacks MFA
|
||||
- **Reader Guide**: This is where you determine if the company passed or failed SOC 2
|
||||
|
||||
#### Section 2: Management's Assertion
|
||||
- Management confirms that the description of systems and controls provided is accurate and complete
|
||||
- Management acknowledges design and operational effectiveness (for Type 2)
|
||||
- Must be signed by company leadership (CEO, CTO, etc.)
|
||||
- Demonstrates company ownership and responsibility for the audit
|
||||
|
||||
#### Section 3: System Description (Most Important Section)
|
||||
9 Description Criteria (DC1-DC9):
|
||||
|
||||
- **DC1 - Overview of Services Provided**: Brief overview of services (must be objective facts, not marketing language)
|
||||
- **DC2 - Principal Service Commitments and System Requirements**: Customer contract commitments related to in-scope TSCs
|
||||
- **DC3 - System Components**: Technical details including:
|
||||
- Hosting location (AWS/GCP/Azure, etc.)
|
||||
- Software tools used
|
||||
- Infrastructure, software, people, procedures, and data
|
||||
- ✓ Best section to quickly understand company tech stack
|
||||
- **DC4 - Events Not Aligning with Service Commitments**: Description of incidents failing to meet commitments (e.g., outages) and remediation
|
||||
- **DC5 - Control Activities**: Narrative description of controls evaluated
|
||||
- **DC6 - Complementary User Entity Controls (CUEC)**: Controls users should have in place
|
||||
- Example: Users must notify the company to remove access of terminated employees
|
||||
- **DC7 - Complementary Subservice Organization Controls**: Controls third parties should have (shared responsibility model)
|
||||
- **DC8 - Non-Applicable Criteria**: Standards not applicable to the organization
|
||||
- **DC9 - Significant Changes to the System**: Major system changes during the period (Type 2 only)
|
||||
|
||||
✓ **Critical Tip**: Read Section 3 to verify the SOC 2 covers services relevant to your organization
|
||||
|
||||
#### Section 4: Trust Services Criteria and Related Controls
|
||||
- **Type 1**: Lists controls only
|
||||
- **Type 2**: Lists controls + test steps + test results
|
||||
- **Exceptions and Deviations**: When auditor finds a control not operating effectively
|
||||
- Example: During sampling, auditor discovers one employee didn't complete required training
|
||||
- Type 2 Focus: Review any controls with exceptions and assess the risk
|
||||
|
||||
#### Section 5: Other Information Not Covered by Auditor's Report (Optional)
|
||||
|
||||
Two Common Uses:
|
||||
|
||||
1. **Management Response to Exceptions**:
|
||||
- Background information about identified issues
|
||||
- Remediation steps taken by the company
|
||||
- Explanation of how the exception is not systemic
|
||||
|
||||
2. **Mapping to Other Frameworks**:
|
||||
- Map SOC 2 controls to HIPAA, ISO 27001, PCI DSS
|
||||
- Help industry-specific customers understand framework relevance
|
||||
- Example: Healthcare companies showing HIPAA compliance alignment
|
||||
|
||||
✓ Not audited by third party; management's responsibility
|
||||
✓ Framework mapping helps sales in specific verticals
|
||||
|
||||
|
||||
## 【Chapter 6】Trust Services Categories (TSCs) - Scoping
|
||||
|
||||
### Core Information
|
||||
- TSCs are the pillars of evaluation (choose from 5 categories)
|
||||
- Scoping Decision Principle: **Base decisions on customer commitments**
|
||||
- Look for commitments in customer contracts, SLAs, and MSAs (For example: Uptime Commitment: 99.9%)
|
||||
- Don't include categories just to include them; cost and effort increase significantly
|
||||
|
||||
### Detailed Breakdown
|
||||
|
||||
#### Selection Process
|
||||
1. Review customer contracts/SLAs/MSAs for commitments
|
||||
2. Identify which TSCs align with these commitments
|
||||
3. Ensure you have documented commitments for each in-scope TSC
|
||||
|
||||
### Key Terms and Definitions
|
||||
- **Commitment**: Pledges made to customers in contracts, service level agreements, master service agreements, or terms and conditions
|
||||
- **Scoping**: The process of choosing which TSCs to include in your SOC 2 audit
|
||||
- **In-Scope**: TSCs and controls that are part of your SOC 2 audit
|
||||
|
||||
### Action Items
|
||||
- Review all customer contracts for security-related commitments
|
||||
- Identify commitments related to each potential TSC
|
||||
- Document which TSCs align with your actual business commitments
|
||||
- Avoid including TSCs without corresponding commitments
|
||||
|
||||
|
||||
## 【Chapter 7】Security TSC (Security Trust Services Category)
|
||||
|
||||
### Core Information
|
||||
- Almost every SOC 2 includes the Security category
|
||||
- Security is the foundation and minimum requirement
|
||||
- Contains 9 Common Criteria (AICPA standard baseline)
|
||||
- Typical control count: 40-50 controls
|
||||
|
||||
### Detailed Breakdown
|
||||
|
||||
#### Covered Security Topics
|
||||
- Onboarding and Offboarding
|
||||
- Risk Assessments
|
||||
- Vulnerability Management
|
||||
- Access Control
|
||||
- Information Security Policies and Procedures
|
||||
- Vendor Management
|
||||
- Other foundational security practices
|
||||
|
||||
#### Best Practices
|
||||
- Early-stage startups can achieve SOC 2 with Security category alone
|
||||
- Commonly paired with Availability and Confidentiality (3-category combination is very common)
|
||||
- 50% of SOC 2 reports include Security + Availability + Confidentiality
|
||||
|
||||
|
||||
## 【Chapter 8】Availability TSC
|
||||
|
||||
### Core Information
|
||||
- Common for cloud-hosted companies (cloud provider features support native capabilities)
|
||||
- Only include if you have availability commitments
|
||||
- Smallest control count: 8-10 controls
|
||||
- Contains 3 criteria (vs Security's 9)
|
||||
|
||||
### Detailed Breakdown
|
||||
|
||||
#### Covered Availability Topics
|
||||
- Backups
|
||||
- Processing Capacity
|
||||
- Replication
|
||||
- Multi-location Strategies
|
||||
- Business Continuity Planning and Testing
|
||||
- Disaster Recovery Planning and Testing
|
||||
|
||||
#### Advantages
|
||||
- Cloud provider default features make evidence provision easy
|
||||
- Natural choice for cloud-native companies
|
||||
|
||||
#### Common Combinations
|
||||
- 50% of SOC 2 reports include Security + Availability + Confidentiality
|
||||
- Especially common in early-stage startups
|
||||
|
||||
### Action Items
|
||||
- Don't include this category just because you're on the cloud
|
||||
- Always base decisions on actual commitments
|
||||
- Verify you have documented availability commitments before including
|
||||
|
||||
|
||||
## 【Chapter 9】Confidentiality TSC
|
||||
|
||||
### Core Information
|
||||
- Key Question: How do you handle customer data when they leave your service?
|
||||
- Focus: Data classification and secure data handling
|
||||
- Control count: 4-8 controls
|
||||
- Contains 2 criteria
|
||||
|
||||
### Detailed Breakdown
|
||||
|
||||
#### Covered Topics
|
||||
- Confidential Information Classification
|
||||
- Confidential Information in Non-Production Environments
|
||||
- Data Deletion and Removal Practices
|
||||
|
||||
#### When to Include
|
||||
- If your MSA commits to deleting all customer data within X days of contract termination
|
||||
- If you make commitments about data handling during customer termination/offboarding
|
||||
|
||||
#### Implementation Challenges
|
||||
- Data deletion and removal practices are difficult to execute correctly
|
||||
- Must establish and mature these processes before the audit
|
||||
- Don't underestimate implementation difficulty
|
||||
- This is not an easy add-on
|
||||
|
||||
### Action Items
|
||||
- Review data deletion commitments in customer contracts
|
||||
- Establish and test data deletion procedures before audit
|
||||
- Ensure processes are mature and consistent
|
||||
- Verify deletion procedures for all data types
|
||||
|
||||
|
||||
## 【Chapter 10】Processing Integrity TSC
|
||||
|
||||
### Core Information
|
||||
- Common Misconception: NOT the "I" in CIA triad (data integrity)
|
||||
- Focus: **Completeness and accuracy of information produced by your system**
|
||||
- Customers depend on your data accuracy
|
||||
- Less Common: Primarily in financial/payment industries
|
||||
|
||||
### Detailed Breakdown
|
||||
|
||||
#### Key Distinction
|
||||
- **CIA Integrity**: Prevents unauthorized deletion or modification
|
||||
- **Processing Integrity**: Ensures system-produced data is complete and accurate
|
||||
|
||||
#### Typical Use Cases
|
||||
- Payroll Processing Systems: Ensure salary calculations are accurate
|
||||
- Payment Processors: Ensure transaction processing accuracy
|
||||
- HR Tools: Ensure HR data accuracy
|
||||
- Financial Software: Ensure financial report accuracy
|
||||
|
||||
#### Implementation Characteristics
|
||||
- 5 criteria
|
||||
- Controls are often very specific and unique to the application
|
||||
- Cannot use generic control sets
|
||||
- Requires customization
|
||||
|
||||
### Action Items
|
||||
- Evaluate if your system produces data customers depend on for accuracy
|
||||
- Document accuracy commitments in customer contracts
|
||||
- Only include if you commit to providing complete and accurate information
|
||||
|
||||
|
||||
## 【Chapter 11】Privacy TSC
|
||||
|
||||
### Core Information
|
||||
- Privacy ≠ Security (often confused in industry)
|
||||
- Privacy has narrow scope; only include if relevant
|
||||
- **Don't include just because it's a fashionable term**
|
||||
- Adds significant effort and cost
|
||||
|
||||
### Detailed Breakdown
|
||||
|
||||
#### Key Question
|
||||
- Are you a **Data Controller** (directly interact with data subjects) or **Data Processor** (process data on behalf of others)?
|
||||
|
||||
#### When to Include Privacy TSC
|
||||
- Data Controllers: Directly interact with individuals; handle PII (Personally Identifiable Information)
|
||||
- Have genuine privacy commitments to customers
|
||||
|
||||
#### When Privacy TSC is NOT Needed
|
||||
- Data Processors: Only process PII on behalf of others without direct data subject interaction
|
||||
- Confidentiality TSC should suffice for report readers/customers
|
||||
|
||||
#### Implementation Complexity
|
||||
- 8 criteria (second largest after Security's 9)
|
||||
- Significant complexity increase in reporting and testing
|
||||
- Many "not applicable" criteria often result in report redundancy
|
||||
|
||||
#### Common Mistakes
|
||||
- Companies mistakenly include Privacy in scope
|
||||
- Result: Pay auditors to repeatedly mark "This criterion is not applicable"
|
||||
- Additional work and cost provides no business value
|
||||
|
||||
### Key Terms and Definitions
|
||||
- **Data Controller**: Organization that determines the purposes and means of processing personal data
|
||||
- **Data Processor**: Organization that processes personal data on behalf of the data controller
|
||||
- **PII (Personally Identifiable Information)**: Any information that can identify an individual
|
||||
|
||||
### Action Items
|
||||
- Determine if you are a data controller or processor
|
||||
- Review privacy commitments in customer contracts
|
||||
- Only include if you interact directly with data subjects
|
||||
- Evaluate if additional complexity is worth the effort
|
||||
|
||||
|
||||
## 【Chapter 12】Next Steps
|
||||
|
||||
### Further Learning Resources
|
||||
- SANS Institute SOC 2 Blog
|
||||
- ByteCheck Resource Library (ByteCheck.io)
|
||||
- LinkedIn: @Ajay Yond or Twitter: @aj_yond
|
||||
|
||||
### Key Takeaways
|
||||
- SOC 2 is third-party proof that security controls are in place
|
||||
- Choose the right TSCs based on customer commitments
|
||||
- Type 2 is the end goal, but Type 1 is a necessary stepping stone
|
||||
- Understanding the 5 report sections enables proper compliance assessment
|
||||
- Flexibility is SOC 2's greatest strength
|
||||
|
||||
### Action Items
|
||||
- Review your customer contracts for security commitments
|
||||
- Identify which TSCs align with your commitments
|
||||
- Plan your SOC 2 journey starting with Type 1
|
||||
- Engage an auditor to discuss your specific situation
|
||||
- Build internal security maturity before the audit
|
||||
|
||||
---
|
||||
|
||||
**End of Summary**
|
||||
|
||||
*Generated with Course Transcript Summarizer skill*
|
||||
*Format: Markdown | Language: English*
|
||||
Reference in New Issue
Block a user